Full Text
INTRODUCTION TO THE THEORY OF ELEMENTARY PARTICLES*
D. Ivanenko
CONTENTS
§ 11. Cosmic rays . . . . . . . . . . . . . . . . . . . . . . . . 261
§ 12. Abundance of elements and particles . . . . . . . . . . . . 265
§ 13. General questions of relativistic quantum mechanics . . . . 268
1. Transformation groups and invariance . . . . . . . . . . . 269
2. Wave functions . . . . . . . . . . . . . . . . . . . . . . 271
3. Spin . . . . . . . . . . . . . . . . . . . . . . . . . . . 273
4. Equations of relativistic quantum mechanics . . . . . . . . 274
5. Lagrangian function . . . . . . . . . . . . . . . . . . . . 286
6. Theory of interaction . . . . . . . . . . . . . . . . . . . 290
7. Secondary quantization and statistics . . . . . . . . . . . 296
§ 14. Difficulties of the theory . . . . . . . . . . . . . . . . . 301
1. The problem of proper mass. The field hypothesis . . . . . . 301
2. New hypotheses . . . . . . . . . . . . . . . . . . . . . . 306
§ 11. COSMIC RAYS
As is known, cosmic rays, possessing enormous energies—on the average several billions of electron-volts in the upper layers of the atmosphere—fall upon the Earth from all directions of world space. The primary particles, apparently consisting mainly of protons, produce mesotrons, and then—directly or indirectly—electrons, positrons, and photons. Thus, in the flux of cosmic rays we have matter in the most highly subdivided state, down to elementary particles. At sea level, on average, one particle per minute falls on \(1 \text{ cm}^2\). It is essential to emphasize
D. IVANENKO
as an insignificant secondary source of “additional” ionization and were then discovered with complete clarity by Hess in 1911 during ascents in balloons up to 5 kilometers.
Let us briefly list the main stages in the study of cosmic rays.
a) Millikan, Regener, Piccard, and Schein advanced measurements of cosmic rays to 15–35 km and, finally, recently, by using a rocket, to 150 km above sea level. At an altitude of about 20 km an increase in intensity by approximately 80 times was discovered.
b) Skobeltsyn (1929), Anderson, Kunze, and Blackett photographed cosmic rays by various methods in a Wilson chamber, which made it possible to discover many new phenomena in this field.
c) Clay and Compton (1932) discovered, in extensive expeditions, the influence of the Earth’s magnetic field on cosmic rays and thereby proved the presence of charged particles in the primary flux.
d) Anderson discovered the positron in cosmic rays (1932).
e) Blackett first discovered in cosmic rays the formation of pairs and of many pairs, i.e. showers of electrons and positrons (1933).
f) Anderson and Neddermeyer discovered positive and negative mesotrons in cosmic rays (1937).
g) For deciphering the nature of cosmic rays and of the most complex processes occurring in the cosmic-ray flux, a significant role was played by Auger’s proposed division of cosmic rays into soft rays, absorbed in 10 cm of lead and consisting chiefly of streams of electrons, positrons, and photons with an admixture of slow mesotrons and protons (the number of which rapidly increases with altitude), and hard penetrating cosmic rays, consisting chiefly of streams of mesotrons, with an admixture of nucleons (protons, neutrons), the number of which likewise rapidly increases with altitude.
As to the origin of cosmic rays, as yet nothing is known. All acceleration hypotheses have proved unsuccessful: those that attempted to explain the great energy of the particles by subsequent acceleration in hypothetical electric (Ross, Gunn) or magnetic (Alfvén, Terletsky) fields, as well as Zwicky’s astronomical hypothesis, which tried to connect cosmic rays with explo—
they awaited, until recently, a closely related hypothesis of the formation of cosmic rays as a result of the annihilation of the protons and neutrons of our world with the hypothetical antiprotons and antineutrons of Dirac’s purely hypothetical “antiworld.” To explain the energy of cosmic showers Auger (up to \(10^{15}\)—\(10^{17}\ \mathrm{eV}\)) Klein is forced to invoke processes of annihilation of whole macroscopic grains and pieces of matter! An interesting consequence of his purely hypothetical and preliminary estimates is the circumstance that in these processes there is by no means too rapid a destruction of particles, but the observed density of cosmic rays can be understood. In such a situation one cannot but try to consider cosmic rays from the cosmological point of view, linking their origin with processes in a hypothetical “special” nuclear state of matter (see § 12). It would, however, be premature to discuss this hypothesis here in greater detail.
On the other hand, the picture of the interaction of primary cosmic rays, in all probability protons with energies of about \(10^{10}\ \mathrm{eV}\), with the matter of the atmosphere is now clear in many essential respects. The most plausible Johnson–Schein–Swann hypothesis states that, in collisions with atomic nuclei in the atmosphere, the mesotron field is “torn off” from primary protons. It is still not clear whether mesotrons are produced one after another (i.e. in a cascade) (Heitler), or several particles at once, in an explosive manner (Heisenberg, Bethe, Oppenheimer). The intense production of mesotrons by nucleons through such bremsstrahlung at the expense of nuclear forces occurs chiefly at altitudes of \(25\)—\(30\ \mathrm{km}\). Some of the mesotrons (possibly precisely the vector ones) decay comparatively rapidly according to the scheme: \(\mu_{-}\to e_{-}+\nu\) or \(\mu_{+}\to e_{+}+\nu\), and give rise to the electrons and positrons of the soft component. The energy of the mesotrons is also transferred to the energy of the soft component through the knocking out of electrons from atoms (Bhabha) and partly by the emission of photons. Other mesotrons (possibly precisely the pseudoscalar ones) pass through the atmosphere, forming the main share of the flux of the hard penetrating component. Slow negative mesotrons are absorbed by nuclei (Lukirskii, Yuz) and, owing to their large rest energy (\(\mu\sim100\)—\(200\,m\)), may cause the splitting of nuclei with the emission of nucleons in all directions in the form of a “star” (Zhdanov). However, nuclear disintegrations are caused unquestionably not only by mesotrons; probably all particles of the cosmic flux take part in them, in particular neutral ones, and especially nucleons (Hazen)*). It should be emphasized that, in the cosmic flux, besides the basic mesotrons of mass \(200\,m\), there are unquestionably also heavier mesotrons with masses up to \(990\,m\) (Leprince-Ringuet, Alikhanian).
*) Let us note that recently the emission of several nucleons has been observed in the splitting of nuclei by photons with an energy of \(100\ \mathrm{MeV}\), obtained in a betatron under laboratory conditions.
Deciphering the hard component, mainly near sea level, as a flux of mesotrons possessing a finite lifetime, on average $\tau = 2 \cdot 10^{-6}$ sec., made it possible to explain the temperature and barometric dependence of the intensity of cosmic rays (Blackett).
Most successfully developed, and splendidly confirmed by experiments, is the theory of the main part of the soft, absorbing component of cosmic rays, consisting of electrons, positrons, and photons.
Applying Dirac’s theory of electrons and positrons and the quantum electrodynamics of photons up to the very highest energies, Bhabha and Heitler, as well as Oppenheimer and Carlson, showed in 1937 that at energies of hundreds of millions and billions of electron-volts electrons and positrons lose their energy chiefly by the bremsstrahlung emission of photons while passing through the electric field of nuclei. Photons of enormous energies, in turn, are most likely to lose energy also in collisions with nuclei (and not in the photoelectric ejection of electrons or the transfer of energy to electrons in the Compton effect, etc.), with an electron–positron pair being produced. It turned out that when photons pass through layers of matter containing heavy elements (in which all these effects are most vividly manifested, owing to the high charge of the atomic nucleus), even when the layers have only slight thickness, there is a high probability of producing many pairs (electrons and positrons) one after another by a cascade process. Thus were explained the showers of particles discovered by Blackett in 1933. Auger’s extensive showers, spreading over enormous areas with a radius of up to a kilometer, likewise consist mainly of electrons, positrons, and photons, undoubtedly containing an admixture of mesotrons, apparently in their densest part—the “core.”
Subsequently the theory of cascade showers was worked out in detail chiefly by Soviet authors*). Apparently, the cascade theory, and thereby the relativistic quantum mechanics of electrons, positrons, and photons, is in the main confirmed even for Auger’s extensive showers with energies of $10^{15}$—$10^{17}$ eV. Still, it is necessary to point out certain “clouds” gathering over the theory of Auger’s extensive showers; their enormous extent cannot as yet be satisfactorily explained by a theory taking into account only electromagnetic interactions. Apparently, various nuclear interactions must be included here in the consideration.
Further progress in the understanding of elementary particles is undoubtedly, in many respects, connected with the study of cosmic rays, which represe—
*) Landau, Ivanenko, and Sokolov, exact solution of the equations of cascade theory; Tamm and Belenky, allowance for ionization corrections.
...which for the time being are the only possibility available to us for studying particles at such high energies and, in particular, the possibility of observing the production of mesotrons.
§ 12. ABUNDANCE OF ELEMENTS AND PARTICLES
It is clear that the problem of the abundance of elementary particles is connected with the question of the abundance of the elements and their origin, with the question of the nature of cosmic rays, and is closely adjacent to cosmology. A detailed discussion of these problems, the development of which has for the most part advanced only in the last 10–20 years, lies beyond the scope of our article. We shall touch upon them only in order to outline the range of questions connected with the problems of elementary particles.
The study of the earth’s crust, meteorites, and stellar atmospheres has shown an approximately identical abundance of the elements throughout the known part of the universe. The results of the work of many investigators, in particular Goldschmidt, Russell, and Fersman, state: 1) elements of even atomic number, and among them of even atomic weight, occur more frequently than odd ones (Harkins’ rule); 2) the concentration of isotopes of elements decreases in a definite way and rapidly with increasing atomic number.
Data on the concentration of chemical elements give the number of protons and neutrons making up nuclei, and of electrons revolving around nuclei. Adding to this information on the mean density of matter \(\rho \sim 10^{-27}—10^{-30}\ \text{g}/\text{cm}^{3}\) and on the total mass of matter \(M \sim 10^{51}—10^{53}\ \text{g}\) in the part of the universe known to us (of dimensions \(R \sim 10^{27}\ \text{cm}\)), we obtain a rough, and necessarily so, estimate for the number of particles
\[ N_p \sim N_n \sim N_e \sim 10^{75}—10^{77}. \]
Here it should be noted that, according to Dirac, the possibility is not excluded of the existence of parts of the universe where all particles are replaced by antiparticles of the opposite sign of charge. Atomic nuclei in such an antiworld would be composed of antiprotons and antineutrons, around which positrons would revolve in atoms, while \(e_-\) would be as rare as \(e_+\) in our part of the universe. It is interesting that all spectral phenomena depend on the square of the charge. Therefore it is impossible by optical observations to distinguish an antiworld from an ordinary one. Nevertheless, a predominant flux of \(e_+\) or of antiprotons should in one way or another help to reveal the existence of an antiworld, for example by means of observations in cosmic rays.
In any case, it is important to emphasize that if the existence of antiprotons and the hypothesis of an antiworld are admitted, then our part of the universe, with its known distribution of elementary particles, nuclei, and chemical elements, will prove to be determined by cosmological circumstances and is by no means the only possible one.
Two questions now arise: A) Is the concentration of elementary particles, photons, chemical elements, and cosmic radiation essentially constant? B) Has the concentration of the various types of matter always been equal to that observed today, and what are the tendencies of its change?
As regards protons and neutrons, then, despite the fundamental possibility of their annihilation and creation, these processes have not been observed. Thus the total number of nucleons in the universe is conserved and, so far as is known, must always have been conserved. The question of the relative number of \(n\) and \(p\), however, is connected mainly with changes in the concentration of the various isotopes of the elements. The number of \(n\) in stable nuclei is equal to, or exceeds, the number of protons. Owing to \(K\)-capture reactions, the emission of \(e_-\) by nuclei, and the transformation of hydrogen into heavier nuclei in stars, in a known part of the universe there occurs some enrichment in the number of neutrons at the expense of protons. As for nuclear reactions connected with the natural fission of uranium, thorium, plutonium, and the subsequent transformations of the fragments; further, with natural \(\beta\)-radioactivity, \(K\)-capture, the splitting of nuclei, irradiation of radioactive elements, or cosmic rays, with the subsequent formation of artificially radioactive nuclei—all these processes are relatively rare and cannot substantially change the relative concentration of \(p\), \(n\), \(e_-\), \(e_+\), and chemical elements. Still, the processes just listed show that the concentration of elements and particles is by no means frozen even under terrestrial conditions. For example, there is a constant enrichment of \({}_{18}A^{40}\) at the expense of \({}_{19}K^{40}\), etc. A noticeable enrichment of individual isotopes at the expense of the age-long bombardment of nuclei by cosmic rays is not excluded.
So far as is known, at the present time the intensive formation of new elements and, at the same time, the processes \(p \rightleftarrows n\) occur only inside stars. As Bethe showed (specifically for stars belonging to the class of so-called dwarfs of the main sequence, including the Sun), at enormous temperatures of the order of tens of millions of degrees, hydrogen nuclei possess sufficient thermal energy to cause the disintegration of carbon, which in the final end leads to the formation of helium nuclei from protons:
\[ 4_1\mathrm{H}^1 \longrightarrow {}_2\mathrm{He}^4 + 2e_+ + 2\nu + E; \qquad E \sim 30\ \mathrm{MeV}. \]
Incidentally, the energy released in this process is the principal source of the radiation of the Sun and other similar stars*).
*) With regard to the abundance of neutrons in stars, it has repeatedly been suggested (Hund, Kotari) that in the interiors of stars of certain types, in particular in giants, there exist superdense nuclei with a density of the order of nuclear density, composed of nucleons, or even of neutrons alone. In another variant of this hypothesis the possibility was indicated of the transition of stars with masses greater than some critical value (of the order of the sol-
However, the temperature inside stars is insufficient for the formation of medium and heavy elements. It must be admitted that in the part of the universe known to us we do not know of processes that could lead to their formation and, consequently, to a substantial change in the relative concentration of protons and neutrons.
In exactly the same way, the enormous excess of \(e_{-}\) and the insignificant number of \(e_{+}\) are not substantially changed by the observed processes: \(\beta\)-decay, \(K\)-capture, the formation and annihilation of pairs under the action mainly of cosmic photons, and the decay of mesotrons \(\mu_{-}, \mu_{+}\).
It is evidently impossible to make progress in the questions that interest us without bringing in cosmology. If one adopts the apparently most plausible point of view of an expanding universe, first expressed by the Leningrad mathematician Friedmann (1923), then it is not absurd to suppose that approximately \(10^{10}\) years ago the matter of the part of the universe known to us was compressed to a density of about \(\rho \sim 10^{6}\ \dfrac{\text{g}}{\text{cm}^{3}}\) and had a temperature of about \(T \sim 10^{9}—10^{10}\) degrees. In this “prestellar” state, considered as a preliminary hypothesis by Weizsäcker, Chandrasekhar, and Vatagin, processes of formation of isotopes of various elements could proceed intensively, and one may calculate that a temperature \(T = 8 \cdot 10^{9}\) degrees best corresponds to the conditions for the formation of light and medium elements, approximately up to silver, in quantities corresponding to the observed abundance.
Subsequently, with the expansion of this part of the universe and the lowering of the temperature, the nuclear reactions for the most part ceased, and the equilibrium concentration of the elements that had formed turned out to be “frozen in.”
The concentration of heavy elements, however, from theoretical considerations, turns out to be apparently too low in comparison with that observed. As an entirely preliminary hypothesis we wish to point to the possibility of another “special” state. Indeed, the conceivable limit of compression of matter will be a still greater nuclear density of the order \(\rho \sim 10^{14}\ \text{g}/\text{cm}^{3}\). Then all the matter of the known part of the universe (in the form of \(n\) and \(p\)) will turn out—
—of course, in a degenerate state with a density of the order of nuclear density. It should, however, be emphasized that the presence of superdense nuclei in any stars is not confirmed. As for the formation of neutron stars by \(K\)-capture and the fall of all \(e_{-}\) onto nuclei, one must bear in mind the instability of neutrons, noted by the authors cited, which leads to the fact that \(n\) transform into \(p + e_{-} + \nu\); thus nucleons of both types will necessarily be present in the system.
would be compressed into a volume with linear dimensions of order \(R \sim 10^{13}\,\mathrm{cm} = 10^8\,\mathrm{km}\)*).
It should be emphasized that at ultrahigh \(T\) not only nuclear reactions with the formation of new elements begin to take place, but also processes of production of elementary particles. At
\[ T > \frac{2mc^2}{k} \sim 10^{10} \]
degrees, corresponding to \(10^8\ \mathrm{eV}\), pairs of light particles \((e_{-}, e_{+})\) will begin to be produced by various processes. In particular, radiation at very high temperatures will be mixed with electrons and positrons and will contain approximately equal fractions of \(e_{-}\), \(e_{+}\), and photons, if one abstracts from the production of mesotrons of intermediate mass \((\mu \sim 200\,m)\), which will begin to occur appreciably at a temperature \(T \sim 10^{12}\) degrees, corresponding to 100 MeV. (It is not uninteresting to recall that existing accelerator installations give an energy of 30 MeV—the cyclotron—and up to 100–150 MeV—the betatron, which corresponds to a temperature of \(10^{12}\) degrees. In the explosion of an atomic bomb the temperature apparently reached \(10^6\) degrees.)
To summarize, one may say that if the abundance of elements and isotopes on Earth can be understood on the basis of geological, chemical, and natural-radioactive processes, taking into account some action of cosmic rays, then in order to understand the abundance of isotopes of elements and particles in the universe it is necessary also to take into account nuclear reactions in the stars. The explanation of the formation of the elements and thereby of the concentration of particles requires the introduction of cosmological considerations, purely hypothetical at the present stage of knowledge.
§ 13. GENERAL QUESTIONS OF RELATIVISTIC QUANTUM MECHANICS
After analyzing the basic empirical data on the individual elementary particles, it is necessary, for an understanding of the principal laws governing the motions and interactions of all particles, to give the most concise exposition of the basic principles of relativistic quantum mechanics, which is the modern theory of elemen-
*) It is worth noting here that, from this point of view, the singular solutions which arise in all variants of nonstationary cosmological theories must be reconsidered, in view of the fact that matter cannot be compressed into a point. The “special” nuclear state of matter is the boundary of applicability of macroscopic theories; its further consideration must proceed by means of relativistic quantum mechanics, and, apparently, with allowance for the quantum theory of gravitation.
As was noted in our discussion with A. L. Zelmanov, it is not excluded that such a nuclear state lasted for a very short time, when the breakup into separate astronomical objects and atomic nuclei occurred. Paradoxically, from this there arises the hypothesis of the possibility of treating some stars as the initial, and not the final, stage of evolution.
elementary particles. In this connection, the following questions are considered here:
- Transformation groups and invariance. 2. Wave functions. 3. Spin. 4. Equations of relativistic quantum mechanics. 5. The Lagrangian function. 6. Theory of interaction. 7. Second quantization and statistics.
1. Transformation Groups and Invariance
The wave functions describing elementary particles must have definite transformation properties, i.e. must change in a regular way under transformations of coordinates or the corresponding motions of reference frames. Here one is dealing with transformations in the four-dimensional world: space—time. Of greatest importance are invariants (scalars), i.e. quantities that do not change upon passage to different coordinate systems. The very equations of relativistic quantum mechanics describing one or another particle must preserve an invariant form, since the laws of nature, i.e. the processes expressed by them, do not depend on the choice of the coordinate system.
The requirement of invariance belongs among the most general and profound laws of nature, expressing the most basic properties of space and time, the motion of particles, and their interactions.
Let us list all the groups of transformations that leave the relativistic quantum equations invariant:
1) The absence of an absolute center in space or of an origin for the reckoning of time, i.e. the homogeneity of space—time, leads to invariance with respect to a displacement of the coordinate origin (“the principle of relativity of the origin of reference”). Consequently, the equations must be differential in the four coordinates. If \(x'_s = x_s + a_s\) \((s = 1, 2, 3, 4)\), then \(dx'_s = dx_s\); \((x_1 = x; x_2 = y,\ x_3 = z)\).
2) The isotropy of space, or the “principle of the relativity of directions,” leads to invariance with respect to three-dimensional spatial rotations.
From elementary geometrical considerations it is obvious that in this case the distance (or the square of the distance) between two points remains invariant, for example with coordinates \((0, 0, 0)\) (the origin of coordinates) and \((x_1, x_2, x_3)\), or, respectively, \((0, 0, 0)\) and \((x'_1, x'_2, x'_3)\):
\[ r^2 = x_1^2 + x_2^2 + x_3^2 = r'^2 = x_1'^2 + x_2'^2 + x_3'^2. \]
For infinitesimal rotations: \(dr^2 = dr'^2\).
3) The “principle of relativity of uniform and rectilinear motions,” or the “special principle of relativity—”
“—st,” established on the basis of the works of Lorentz, Einstein, and Poincaré in 1905, asserts the equivalence of all so-called inertial reference frames moving rectilinearly and uniformly relative to one another. This requires the inclusion of a fourth coordinate—time—to a considerable extent on equal footing with the three spatial ones \((x_4 = ct,\ \text{or } x_4 = ict)\), and leads to the invariance of the equations with respect to the so-called Lorentz transformations, which reduce to rotations in the planes \((xt), (yt), (zt)\) and express the transition from one inertial reference frame to another. In this case the interval \(s\), or the distance between two points in the four-dimensional world, for example with coordinates \((t, r)\) and respectively \((t', r')\), remains invariant:
\[ s = c^2 t^2 - r^2 = s'^2 = c^2 t'^2 - r'^2, \]
or, for infinitesimal intervals:
\[ ds^2 = c^2 dt^2 - dr^2 = ds'^2 = c^2 dt'^2 - dr'^2. \]
4) The arbitrariness in the choice of right-handed and left-handed coordinate systems and the symmetry with respect to the past and the future lead to invariance with respect to mirror reflections (inversions)
\[ X'_1 \to -x_1;\qquad X'_2 \to -x_2;\qquad X'_3 \to x_3, \]
and also with respect to reversal of time: \(t' \to -t\).
It is very essential that the equations of relativistic mechanics, just like those of non-quantum and nonrelativistic mechanics, are reversible in time. As is known, irreversibility is interpreted by means of statistical consideration.
5) The absence in nature of any selected coordinate systems and the fundamental possibility of using any arbitrarily moving reference frames, and not only inertial ones, lead to the invariance of equations with respect to transformations to any curvilinear coordinate system (“general principle of relativity” in the presence of gravitation).
This means that the kinematics of elementary particles, as well as all other laws of nature, can be given the so-called generally covariant form, by writing them, according to Einstein, in tensor form with the aid of the components of the metric tensor \(g_{\mu\nu}\)*).
6) The equations of motion, generally speaking, are not invariant with respect to conformal transformations, under which the interval is multiplied by some function of the coordinates. However, as was shown by Bateman and Cunningham, Maxwell’s equations are conformally invariant, which is closely connected with the absence of mass in photons (although, as Pauli noted, for example, the Dirac equations for vanishing mass—which is probably true for the neutrino—will not be conformally invariant).
*) Or, as Tetrode–Fock–Ivanenko showed for spinors, with the aid of generalized Dirac matrices \(\gamma_\mu\).
7) Proceeding from the obvious possibility of choosing arbitrarily the origin for the reckoning of electromagnetic potentials, we arrive at the invariance of the equations of the electromagnetic field and of the equations of all other particles interacting with the field with respect to “gauge” transformations of the potentials (of the 2nd kind, according to Pauli’s nomenclature):
\[ A'_4=A_4+\frac{1}{c}\frac{\partial f}{\partial t}, \qquad A'_s=A_s+\frac{\partial f}{\partial x_s}, \qquad (s=1,2,3), \]
where \(f\) is a scalar function of all the coordinates. Therefore the potentials themselves cannot enter into Maxwell’s equations.
8) Although the potentials enter explicitly into the Dirac equation and other quantum equations of particles, a change of their gauge is compensated by a change of phase of the wave function. The obvious possibility of changing the phase of the \(\psi\)-function leads to the requirement of invariance of all quantum equations (including the nonrelativistic Schrödinger equation) with respect to a “gauge” transformation of the wave functions (of the 1st kind, according to Pauli’s nomenclature)
\[ \psi \to \psi e^{iF}. \]
Indeed, the directly observable quantities are bilinear combinations of the type of probability density: \(\rho=\psi^*\psi\), which are not changed under such a transformation.
Let us now emphasize that, according to Hilbert–Noether’s very general theorem, to every continuous group of transformations there corresponds some conservation law. For example, from invariance with respect to a displacement of the origin of coordinates (1) there follows the law of conservation of energy-momentum; invariance with respect to rotations in space (2) gives conservation of angular momentum, and so on.
2. Wave Functions
Within the framework of relativistic quantum transformations, particles may be described by the following wave functions:
1) Scalar function (or invariant, i.e. a tensor of rank 0). The scalar \(\psi\) does not change under all possible transformations.
2) Vector with four components (tensor of rank 1). The concept of a vector as a directed quantity characterized by three components along the coordinate axes is known to everyone. A four-dimensional vector will, obviously, be characterized by components along the axes \(x, y, z, ct\) (or \(ict\)), where \(c\) is the speed of light, added in order to preserve dimensionality. It is convenient to renumber the coordinates and write: \(x_1=x;\ x_2=y;\ x_3=z;\ x_4=ct\) (or \(ict\)). Under transformations of the coordinates, for example rotations of the coordinate system, the components of four-dimensional vectors transform like the coordinates. The \(\psi\)-function of the electromagnetic or vector meson field is a four-dimensional vector (potential): \(\psi_\nu=A_1, A_2, A_3, iA_4\).
3) More complicated entities, tensors of rank 2, can be obtained by taking products of two vectors. In relativistic quantum mechanics, a symmetric tensor of rank 2 (with respect to linear transformations) is the ψ-function of the graviton \(h_{\mu\nu}=h_{\nu\mu}\). An antisymmetric tensor of rank 2 \(F_{\mu\nu}=-F_{\nu\mu}\) is the collection of components of the magnetic and electric fields \((F_{12}=H_z,\ F_{13}=-H_y,\ F_{23}=H_x;\ F_{14}=-iE_x,\) etc.), or the collection of components of the quasimagnetic and quasielectric mesotron field, obeying the Proca equations. Proca or Maxwell fields are called vector fields, since in them a vector ψ-function (potential) is taken as fundamental: \(\Phi_\mu\), or, respectively, \(A_\mu\), while the components of the field strengths are obtained by applying the four-dimensional operation “curl”:
\[ F_{\mu\nu}=\frac{\partial A_\nu}{\partial x_\mu}-\frac{\partial A_\mu}{\partial x_\nu}. \]
In an analogous way one may form tensors of higher rank and use them as wave functions for describing various fields.
4) Quantities “dual” to those enumerated above play an important role in modern theory. To a scalar (invariant) there corresponds a pseudoscalar or, what essentially means the same thing, a tensor of rank 4, antisymmetric in all indices,
\[ \psi \to \psi_{\alpha\beta\gamma\delta}\quad (\psi_{\alpha\beta\gamma\delta}=-\psi_{\beta\alpha\gamma\delta}\ \text{etc.}). \]
Thus \(\psi_{\alpha\beta\gamma\delta}\) has, in essence, one component different from zero, as does a scalar: \(\varphi_{1234}=-\varphi_{2134}\), etc. To a vector there corresponds a pseudovector, or a tensor of rank 3, antisymmetric in all indices, i.e., possessing, like a vector, four components: \(A_\alpha \to \psi_{\alpha\beta\gamma}\) \((\psi_{\alpha\beta\gamma}=-\psi_{\beta\alpha\gamma}\), etc.). A pseudoscalar is an invariant under all transformations except mirror reflections, under which it changes sign. In exactly the same way, a pseudovector behaves like a vector always, with the exception of mirror transformations.
In Dirac’s theory, quantities of a new type were introduced for the first time, simpler, more primary than vectors—namely spinors, or tensors of rank \(1/2\), which may also be called “semivectors.” Speaking intuitively, a spinor is a kind of “square root” of a vector, so that from the product of spinors one can form a vector. A spinor has two components, which are sufficient for considering rotations and Lorentz transformations of the coordinates. However, if mirror reflections are taken into account, then alongside the spinor one must introduce its associated dual; then we obtain the complete system of two spinors, or a bispinor, containing in all 4 components. The wave function of particles of spin \(1/2\) is a bispinor and obeys the Dirac equation.
For the description of particles with other values of half-integer spin, for example \(^{3}/_{2}\), it is likewise necessary to employ quantities of the spinor type. Ordinary tensors and spinors can be treated together as spin-tensors or “undors” (Belingfante).
3. Spin
The transition to wave functions with several components is explained by the fact that the electron, the proton, and other particles have spin and that particles, speaking rather graphically, can be oriented in a definite way. This corresponds to the polarization of \(\psi\)-waves describing particles. The description of polarized waves requires the introduction of several components, some of which are equal to 0 in one or another state of polarization. Hence it also becomes clear that the transformation character of the \(\psi\)-function and, at the same time, the number of its components determine the value of the spin. Let us recall that, to describe the polarization of electromagnetic waves, it is necessary, as is well known, to pass to a four-component vector function, or four-dimensional potential, which includes, along with the scalar \(A_4\), also the vector potential \(A\). Further, to describe the gravitational field, from the one-component Newtonian potential one must pass to Einstein’s tensor wave function with ten components \(g_{\mu\nu}\) (or, for a weak field, \(h_{\mu\nu}\)).
If the spin is equal to \(^{1}/_{2}\), i.e., more precisely, if the spin component along any axis can take two values:
\[
s_z=\pm \frac{1}{2}\left(\frac{h}{2\pi}\right),
\]
then it is necessary to introduce a two-component function, or spinor, which transforms in a definite way under coordinate transformations. However, in order to ensure invariance with respect to reflections, as was indicated, it is necessary to consider a four-component bispinor coinciding with the Dirac \(\psi\)-function. The Dirac equation for the \(\psi\)-function and the value of spin \(^{1}/_{2}\) are uniquely connected with one another.
The analogy with polarization is also preserved for electrons. For example, in a beam of electrons after reflection the spins will be oriented predominantly in one direction. Experimental proof of the polarization of electrons as a result of double reflection was apparently obtained by Shull (theoretical calculations were given by Mott, and also by Sokolov).
Particles of spin 1 are described by vector functions with four components and obey the Proca equations, or, when the mass vanishes, Maxwell’s equations. In this case, in view of the fulfillment of the Lorentz condition, only three independent components remain. The value of the spin \(s=1\) corresponds directly to the two possible polarizations of the photon or electromagnetic wave, as well as of the vector mesotron.
In addition to Proca’s equation, particles of spin 1 can be described by the equation dual to it and coinciding with it in the absence of interactions—the pseudovector equation. Particles of vanishing spin, however, are described either by the scalar de Broglie equation, or by the equation dual to it and coinciding with it in the absence of interactions—the pseudoscalar equation. Generally speaking, the number of independent components of the $\psi$-function is related to the value of the spin $s$ by the formula: $n=2s+1$. For example, for $n=1$ ($\psi$—scalar) $s=0$; for $n=2$ ($\psi$—spinor) $s=\frac12$; for $n=3$ ($\psi$—vector) $s=1$; for $n=5$ ($\psi$—symmetric tensor of rank 2) $s=2$.
Particles of spin 2, described by a tensor of rank 2, in the case of vanishing mass obey the equations of Einstein’s weak gravitational field. The theory of particles of higher spin $(s>2)$ encounters a number of difficulties and is still far from complete (Pauli and Fierz, de Broglie, Bhabha).
4. Equations of relativistic quantum mechanics
Let us turn to the fundamental equations of relativistic quantum mechanics. It is very important that, given 1) specified transformation properties of the $\psi$-function, 2) restriction to the lowest derivatives, and also 3) restriction to linearity, the equations are obtained uniquely (for small spins $s \leq 2$). The simplest way to obtain them is from a variational principle, if the Lagrangian function is known. One may also directly construct the corresponding operator, proceeding from the classical relations between energy and momentum, or from geometrical considerations, or, finally, from the matrix formulation.
a) First of all let us recall the fundamental Schrödinger equation of nonrelativistic quantum mechanics. It may be regarded as the usual relation between total, kinetic, and potential energy, in which, however, the quantities of momentum and energy have been given the symbolic meaning of differentiation operators. Namely, instead of the familiar formula of nonrelativistic mechanics:
\[ E-\frac{p^2}{2m}=0; \]
replacing:
\[ E=-\frac{h}{2\pi i}\frac{\partial}{\partial t}, \qquad p_x=-\frac{h}{2\pi i}\frac{\partial}{\partial x}\ \text{and so on,} \tag{5} \]
and applying the whole operator to the function $\psi$, we obtain the Schrödinger equation
\[ S\psi \equiv -\frac{h}{2\pi i}\frac{\partial \psi}{\partial t} +\frac{h^2}{8\pi^2 m}\Delta\psi=0. \]
\[ \left(\text{The Laplace operator } \Delta= \frac{\partial^2}{\partial x^2} +\frac{\partial^2}{\partial y^2} +\frac{\partial^2}{\partial z^2}.\right) \]
It is easy to verify that: \((pq-qp)\psi=\dfrac{h}{2\pi i}\psi\), i.e. the momentum and coordinate operators are connected by the “commutation rules”:
\[ p_s q_s-q_s p_s=\frac{h}{2\pi i}\quad (q_s=x,\ y,\ z). \]
On the basis of the Schrödinger equation or the commutation rules one may derive the important Heisenberg “uncertainty relation”:
\(\Delta p_s \Delta q_s \sim h\), according to which the product of the uncertainties in a coordinate and in the canonically conjugate momentum cannot be less than the quantum of action.
b) The de Broglie equation, describing a spinless particle by means of a scalar \(\psi\)-function (complex in the case of charged particles and real in the case of neutral particles), can be obtained from the simplest invariant: the square of the four-dimensional gradient, supplemented by an invariant term with the rest mass:
\[ B\psi \equiv (\square-\varkappa_0^2)\psi=0, \tag{B} \]
\[ \square=\Delta-\frac{1}{c^2}\frac{\partial^2}{\partial t^2} =\frac{\partial^2}{\partial x^2}+\frac{\partial^2}{\partial y^2}+\frac{\partial^2}{\partial z^2} -\frac{1}{c^2}\frac{\partial^2}{\partial t^2} =\sum \nabla_\alpha^2,\qquad \varkappa_0=\frac{2\pi mc}{h}. \]
\[ \left(\nabla_x=\frac{\partial}{\partial x}\ \text{and so on}\right) \]
On the other hand, one may take as the basis the geometrical invariant: an infinitesimal interval in four-dimensional space,
\[ ds^2=c^2dt^2-dr^2; \]
dividing it by the square of the element of proper time \(d\tau\), we obtain:
\[ \frac{ds^2}{d\tau^2}=c^2=\frac{c^2dt^2}{d\tau^2}-\left(\frac{dr}{d\tau}\right)^2. \]
Substituting momenta for velocities, \(p=mv=m\dfrac{dr}{d\tau}\), and observing that \(d\tau=dt\sqrt{1-v^2/c^2}\), \(E=\dfrac{mc^2}{\sqrt{1-v^2/c^2}}\), we obtain the well-known relativistic formula expressing the relation between energy and momentum,
\[ \frac{E}{c^2}=p^2+m^2c^2. \]
To pass to quantum mechanics it is again necessary to replace energy and momentum by operators:
\[ E\longrightarrow -\frac{h}{2\pi i}\frac{\partial}{\partial t};\qquad p_x\longrightarrow \frac{h}{2\pi i}\frac{\partial}{\partial x}\ \text{and so on}. \]
Then we again obtain the de Broglie equation.
As Duffin and Kemmer observed, the scalar equation may first be written in the form of a system of 5 equations:
\[ \varkappa_0\chi_\alpha=\frac{\partial\psi}{\partial x_\alpha}\quad (\alpha=1,\ 2,\ 3,\ 4);\qquad \frac{\partial\chi_\alpha}{\partial x_\alpha}=\varkappa_0\psi. \]
Then this system can be represented in the form of a single matrix equation:
\[ \left(\beta_\nu \frac{\partial}{\partial x_\nu}+\chi_0\right)\phi=0, \quad \text{where} \]
\[ \chi_0=\frac{2\pi mc}{h}; \]
\(\phi\) is a column matrix with 5 components,
\[ \phi= \begin{vmatrix} \chi_1\\ \chi_2\\ \chi_3\\ \chi_4\\ \chi_0\phi \end{vmatrix}, \]
and \(\beta_\nu\) are definite matrices of rank 5. This form of writing the scalar equation is very interesting for comparison with the vector (Proca) and spinor (Dirac) equations.
The dual theory, in which a pseudoscalar is taken as the basis, coincides in vacuum or in the case of interaction with an electromagnetic field with the theory of de Broglie’s scalar equation, but in the case of interaction with spinor external fields—for example, of nucleons or electrons—it differs from it rather substantially. Replacing the scalar \(\psi\) by a quantity \(\psi_{\alpha\beta\gamma\delta}\), antisymmetric in all indices, and the vector \(\chi_\alpha\) by a pseudovector \(\chi_{\alpha\beta\gamma}\), likewise antisymmetric in all indices, we obtain:
\[ \chi_{\beta\gamma\delta} = -\frac{\partial \psi_{\alpha\beta\gamma\delta}}{\partial x_\alpha}; \qquad \frac{\partial \chi_{\beta\gamma\delta}}{\partial x_\alpha} - \frac{\partial \chi_{\gamma\delta\alpha}}{\partial x_\beta} + \frac{\partial \chi_{\delta\alpha\beta}}{\partial x_\gamma} - \frac{\partial \chi_{\alpha\beta\gamma}}{\partial x_\delta} = \chi_0^2\psi_{\alpha\beta\gamma\delta}. \tag{C} \]
Either of these equations, (B) or (C), may be applied to the description of mesotrons, if it turns out that their spin is equal to 0. For a number of considerations connected with the theory of nuclear forces, the pseudoscalar equation (C) proves the most suitable for this purpose, since it is capable of describing the spin and noncentral character of the forces between nucleons. The simplest variant of the theory is obtained for neutrality, when the real quantities of the field have a direct classical analogue of “waves” (this “classical mesodynamics” is what we shall have in mind in what follows, unless special reservations are made).
In order to take into account the creation of a scalar, possibly mesotronic, field by particles possessing (mesotronic) quasicharges, i.e. by nucleons with charge \(g\) or by light particles with charge \(g'\), it is necessary to add on the right-hand side of (B), as an inhomogeneity, the terms \(4\pi g\rho\) or \(4\pi g'\rho'\), where \(\rho\) and \(\rho'\) are invariant density functions of the distribution of nucleons and light particles of the form: \(\rho=\psi^*\rho_3\psi\) (where \(\psi\) is the Dirac function of the nucleons, or, respectively, of the light particles; \(\rho_3\) is a scalar matrix), i.e. instead of (B) write: \(B\psi=4\pi g\rho\).
Then, in the static case, de Broglie’s equation acquires the form of an expression generalizing the Laplace–Poisson equation:
\[ (\Delta-\chi_0^2)\phi = -4\pi g\rho\,\Xi \left(\rho_3 \to 1 \text{ when } v\to 0\right). \]
The solution of this equation (or the “potential” of the mesotron field produced by a point nucleon) is equal to: \(\psi=-\dfrac{ge^{-\chi_0 r}}{r}\).
For a point particle the density \(\rho\) is equal to the Dirac \(\delta\)-function. Sources of the pseudoscalar field in equation (C) are taken into account analogously.
c) The Dirac equation. In Dirac’s opinion, a number of shortcomings of the theory of the de Broglie equation (for example, the erroneous fine-structure formula obtained when describing the electron in the hydrogen atom by the de Broglie equation) were connected with the presence in it of the second time derivative and the indefiniteness of the expression for the probability density. Therefore Dirac decided to “linearize” the de Broglie operator, or, so to speak, to extract from it symbolically the “square root,” taking as the basis the externally definite expression for the probability density: \(\rho=\sum_{\nu}\psi_\nu^*\psi_\nu\), similar to the nonrelativistic one, but now connected not with one \(\psi\)-function, but with 4 components \(\psi_\nu(\nu=1,2,3,4)\). It later turned out that Dirac’s criticism was one-sided, since the de Broglie equation is formally irreproachable within the framework of existing relativistic quantum mechanics, but in fact is inapplicable to the electron and other particles of spin \(1/2\), since it describes only particles of spin 0. It is therefore not surprising that the theory of fine structure in the hydrogen atom, or the description of spin phenomena with the aid of the de Broglie equation, proved unsuccessful.
Be that as it may, Dirac’s argument helped him to establish, for the first time, the equation of particles of spin \(1/2\) in the following matrix form:
\[ D\psi=\left(-\frac{h}{2\pi i}\frac{\partial}{\partial t}+c\rho_1(\sigma p)+\rho_3mc^2\right)\psi=0; \tag{D} \]
here \(\rho_1,\rho_3,\sigma=\sigma_1,\sigma_2,\sigma_3\) are certain matrices of fourth rank, and the wave complex function \(\psi\) consists of four components forming a bispinor (for brevity it is often called simply a spinor), which it is convenient to write in the form of a column matrix
\[ \psi= \begin{vmatrix} \psi_1\\ \psi_2\\ \psi_3\\ \psi_4 \end{vmatrix}. \]
Of course, the Dirac equation can be rewritten in the form of a system of four ordinary nonmatrix equations connecting the components \(\psi_\nu\) with one another.
Let us note two basic formulas. 1) the squares of all the matrices are equal to 1, i.e. \(\rho_1^2=\rho_2^2=\rho_3^2=\sigma_1^2=\sigma_2^2=\sigma_3^2=1\); 2) matrices of this type anticommute with one another:
\[ \rho_1\rho_3+\rho_3\rho_1=0;\qquad \sigma_1\sigma_2+\sigma_2\sigma_1=0; \]
\[ \rho_3\rho_1=i\rho_2;\qquad \sigma_1\sigma_2=i\sigma_3. \]
and so on.
Let us note that, as we have shown, the Dirac equation can also be obtained by starting from the geometric invariant: \(ds=\gamma_\alpha dx_\alpha\) of the new matrix metric, analogously to the derivation of the de Broglie equation from \(ds^2\).
With the aid of \(\psi\) one can form the components of a four-dimensional vector according to the following rule:
\[ j_s=\psi^*\rho_1\sigma_s\psi \quad (s=1,\,2,\,3);\qquad j_4=\psi^*\psi, \]
i.e., for example,
\[ j_4=\psi_1^*\psi_1+\psi_2^*\psi_2+\psi_3^*\psi_3+\psi_4^*\psi_4. \]
In addition, in an analogous way one can construct quantities of four other types: the scalar \(\rho_3\); the antisymmetric tensor of the second rank: \(\sigma_{qr}=\rho_3\sigma_s;\ \rho_2\sigma_s\); the pseudovector \(\sigma_s;\ \rho_2\); the pseudoscalar \(\rho_1\)—in all, 16 quantities of five types. In a remarkable way, a direct physical meaning can be assigned to almost all of these 16 quantities. For example, the matrices \(\rho_1\sigma_s\) characterize the velocity vector of a Dirac particle: \(v_s=c(\rho_1\sigma_s)\). This correspondence should be understood in the sense that the mean value of the velocity in any state described by the given \(\psi\) is equal to
\[ \bar v_s=\int \psi^*\rho_1\sigma_s\psi\,d\tau. \]
The scalar corresponds to the proper mass:
\[ \bar m=m\sqrt{1-\frac{v^2}{c^2}}\,\rho_3; \]
the tensor describes the magnetic moment
\[ \bar\mu_s=\frac{eh}{4\pi mc}\rho_3\sigma_s, \]
as well as the electric moment, etc.; the three components of the pseudovector give the spin
\[ \bar s_r=\frac{1}{2}\frac{h}{2\pi}\sigma_r. \]
The Dirac equation in the general theory of relativity has the following form (Fock and Ivanenko):
\[ \left(\sum_{\alpha}^{4}\gamma_\alpha\nabla_\alpha+\chi_0\right)\psi=0, \]
where the \(\gamma_\alpha\) are now functions of all four coordinates, while the covariant derivative of the spinor is equal to:
\[ \nabla_\alpha=\frac{\partial}{\partial x_\alpha}+\Gamma_\alpha-i a\varphi_\alpha\ (a=\mathrm{const.}). \]
This formula makes it possible to take into account the influence of the gravitational field on an electron or another particle of spin \(\frac{1}{2}\), and also to write the Dirac equation in curvilinear coordinates. Here \(\Gamma_\alpha\) is the coefficient of parallel transport of the spinor, analogous to the Christoffel symbol in the case of tensors, which takes into account the action of the gravitational field. The term with \(\varphi_\alpha\), arising quite naturally (and even
in a Euclidean flat world, free of gravitation), may be identified with the potential of the electromagnetic field
\[ \frac{2\pi i e}{hc}\,\varphi_\alpha \left(a=\frac{2\pi e}{hc}\right). \]
Let us note that, substituting in the operator \(D\), in place of \(\rho_3\), the quantity
\[ \sqrt{1-\frac{v^2}{c^2}}, \]
in place of \(c\rho_1\sigma_s\) the velocity \(v_s\), and replacing the momenta by velocities
\[ p=\frac{mv}{\sqrt{1-v^2/c^2}}, \]
we obtain the well-known non-quantum relativistic formula for the energy
\[ E=\frac{mc^2}{\sqrt{1-\frac{v^2}{c^2}}}. \]
Thus, the Dirac equation, like de Broglie’s, represents a peculiarly expressed relation between energy and momentum according to the usual non-quantum formula (Breit).
c) Proca’s equations for a vector complex (in the case of charged particles) \(\Phi\)-function, describing particles of spin 1 (very possibly, mesotrons), whose components we shall denote by \(\Phi_\nu=\Phi_0,\Phi_1,\Phi_2,\Phi_3\), have, in the presence of sources, the form:
\[ \frac{1}{c}\frac{\partial \mathbf F}{\partial t}-\operatorname{rot}\mathbf G-\varkappa_0^2\boldsymbol{\Phi} =-\frac{4\pi g\mathbf v}{c}; \qquad \operatorname{div}\mathbf F+\varkappa_0^2\Phi_0=-4\pi g\rho; \]
\[ \mathbf F+\nabla\Phi_0+\frac{1}{c}\frac{\partial\boldsymbol{\Phi}}{\partial t}=4\pi f\mathbf T; \qquad \mathbf G-\operatorname{rot}\boldsymbol{\Phi}=4\pi f\mathbf S. \tag{P} \]
In the case of neutral particles all the field quantities \(\Phi,\mathbf F,\mathbf G\) are real.
Initially Proca mistakenly intended his equations for the electron. But, as was shown by us jointly with Durandin, they lead to Bose statistics and therefore cannot be suitable for fermionic electrons. On the other hand, Proca’s equations are quite suitable for describing mesotrons of spin 1. For a particle mass equal to 0, Proca’s equations in the case of real functions pass into Maxwell’s equations describing the electromagnetic field. In this case, of course, the sources of the vector mesotron field, described by the terms on the right-hand sides, must be replaced by the sources of the electromagnetic field.
Let us indicate two fundamental properties of Proca’s equations. The components of the \(\Phi\)-functions, or “potentials,” satisfy in empty space the Lorentz equation
\[ \frac{1}{c}\frac{\partial\Phi_0}{\partial t}+\operatorname{div}\boldsymbol{\Phi}=0. \]
Each component of the Proca potential \(\Phi\), and also of the quasielectric \((\mathbf F)\) and quasimagnetic \((\mathbf G)\) fields, satisfies in empty space the de Broglie equation
\[ B\Phi=BF=BG=0. \]
In the presence of particles producing the field, terms are added to the right-hand sides of Proca’s equations which describe the distribution of the densities \(g\rho\) of the quasielectric mesotronic charges and currents \(\dfrac{g\rho v}{c}\) of nucleons (or of light particles \(g'\rho'\)), as well as of the quasimagnetic \(fS\) or quasielectric \(fT\) dipole mesotronic moments of nucleons (and light particles) of absolute magnitude \(f\) \((f\sim \dfrac{g}{\varkappa_0})\):
\[ \rho=\psi^{*}\psi;\qquad \rho v_s=\psi^{*}\rho_1\sigma_s\psi; \]
\[ S=\psi^{*}\rho_3\sigma\psi;\qquad T=\psi^{*}\rho_2\sigma\psi. \]
In the static case, for the quasielectric potential of Proca theory we obtain the equation which we already had in the static case of the de Broglie equation:
\[ (\Delta-\varkappa_0^2)\Phi_0=-4\pi g\rho, \]
with the former solution for a point nucleon:
\[ \Phi_0=\frac{g e^{-\varkappa_0 r}}{r}. \]
For the vector part of the potential of the mesotronic field produced by a point nucleon, similarly, in the static case we obtain
\[ \Phi=-f\,\operatorname{rot}\left(\sigma\,\frac{e^{-\varkappa_0 r}}{r}\right). \]
From this it is easy to obtain expressions for the quasielectric \(\mathbf F\) and quasimagnetic \(\mathbf G\) mesotronic fields.
d) The Maxwell–Lorentz equations of the electromagnetic field in the presence of charges, currents, and magnetic moments of the particles producing the field are obtained directly from Proca’s equations if the mesotron mass is put equal to zero:
\[ \varkappa_0=\frac{2\pi\mu c}{h}=0, \]
and have the form:
\[ \frac{1}{c}\frac{\partial \mathbf E}{\partial t}-\operatorname{rot}\mathbf H = -4\,\frac{\pi e\rho v}{c}; \qquad \operatorname{div}\mathbf E=-4\pi e\rho; \]
\[ \mathbf E+\nabla A_0+\frac{1}{c}\frac{\partial \mathbf A}{\partial t}=4\pi\mu\mathbf N; \qquad \mathbf H-\operatorname{rot}\mathbf A=4\pi\mu\mathbf M, \tag{M} \]
where \(e\rho\) is the charge density and \(e\rho v\) the current density, \(\mu\mathbf M\) and \(\mu\mathbf N\) are the densities of intrinsic nonkinematic magnetic and electric dipoles, of absolute magnitude \(\mu\) (for example, in nucleons); \(A_0\) is the scalar and \(\mathbf A\) the vector potential; \(\mathbf E\) is the electric and \(\mathbf H\) the magnetic field (all these quantities are real in view of the neutrality of the field). In the static case, in the presence of charges, Maxwell’s equations pass into the Laplace–Poisson equation \(\Delta\varphi=-4\pi e\rho\), which for a point charge gives the well-known solution: \(\varphi=e/r\). Let us note that, from the point of view of quantum mechanics, Maxwell’s equations are satisfied for the mean values of the field. On the other hand, they are operator equations for the wave functions of photons.
In complete analogy with the Dirac and de Broglie equations, the Proca equations can be written in the form of a single matrix equation:
\[ \left(\sum_{1}^{4}\beta_\nu\,\frac{\partial}{\partial x_\nu}+\chi_0\right)\Phi=0, \]
where the column matrix \(\Phi\) contains ten components
\[ \Phi= \left| \begin{array}{c} E_x\\ E_y\\ E_z\\ H_x\\ H_y\\ H_z\\ \chi_0 A_x\\ \chi_0 A_y\\ \chi_0 A_z\\ \chi_0 \varphi \end{array} \right|, \qquad \chi_0=\frac{2\pi mc}{h}, \]
and the matrices of the tenth rank \(\beta_\nu\) satisfy certain relations. It is very interesting that the matrices \(\beta_\nu\) of the Proca equations and the matrices \(\beta_\nu\) of the de Broglie equations satisfy exactly the same relations, incidentally substantially different from the formulae for the Dirac matrices. For example, we have \(\beta_\nu^{3}=\beta_\nu\), so that the eigenvalues of \(\beta_\nu\) are \(\pm 1, 0\), instead of the eigenvalues \(\pm 1\) in the Dirac case, etc.
e) Particles of spin 1 can be described not only by the vector Proca wave function \(\Phi_\nu\), or, when the mass vanishes, by the Maxwell potentials \(A_\nu=A_0,A_s\), but also by a pseudovector wave function. In vacuum we obtain the following equations of the pseudovector theory, dual to the vector Proca theory:
\[ -\frac{\partial\Phi_{\beta\gamma\delta}}{\partial x_\alpha} -\frac{\partial\Phi_{\gamma\delta\alpha}}{\partial x_\beta} +\frac{\partial\Phi_{\delta\alpha\beta}}{\partial x_\gamma} -\frac{\partial\Phi_{\alpha\beta\gamma}}{\partial x_\delta}=0; \]
\[ F_{\alpha\beta}=\frac{1}{\chi_0}\, \frac{\partial\Phi_{\alpha\beta\gamma}}{\partial x_\gamma}; \]
\[ \frac{\partial F_{\alpha\beta}}{\partial x_\gamma} +\frac{\partial F_{\beta\gamma}}{\partial x_\alpha} +\frac{\partial F_{\gamma\alpha}}{\partial x} -\chi_0\Phi_{\gamma\alpha\beta}=0; \qquad \frac{\partial F_{\alpha\beta}}{\partial x_\alpha}=0. \tag{V} \]
When the mass of the particle is equal to zero, i.e. \(\chi_0=0\), we obtain a theory dual to the Maxwell theory, capable of describing not photons but, so to speak, pseudophotons. In vacuum the equations of the pseudovector and vector theories coincide, but the interactions of pseudovector and vector particles with other, spinor, particles are quite substantially—
differ. It is not difficult to add, on the right-hand side of (V), terms describing the generation of a pseudovector field by other particles.
If Proca’s equations, in all probability, describe at least some mesotrons, while Maxwell’s equations are, on the one hand, operator equations for the wave functions of photons and at the same time are fulfilled for mean field values, then particles or fields that would be described by pseudovector equations are still unknown. In this respect, as in a number of other points, the general formal scheme of modern relativistic quantum mechanics presents more possibilities than are required by the known experiments; in other words, it turns out to be broader than the known system of elementary particles.
ж) Einstein’s equations. The connection of geometry with matter is described by Einstein’s equations: \(R_{\alpha\beta}-\dfrac{1}{2}g_{\alpha\beta}R=-\chi_e T_{\alpha\beta}\), which in many respects are analogous to the equation of Newtonian gravitation theory, and also to the Maxwell–Lorentz equations of the electromagnetic field or to the newest equations of the mesotron field. Indeed, on the right-hand side stands a term characterizing the distribution and motion of matter (the energy-momentum density tensor \(T_{\alpha\beta}\)), which generates the gravitational field.
On the left-hand side of Einstein’s equations stand the gravitational potentials (i.e. the components of the metric tensor) \(g_{\mu\nu}\) and their derivatives with respect to coordinates and time (symbolically included in the complex terms \(R_{\alpha\beta}\) and \(R\)—Einstein’s tensor and scalar).
In contrast to electrodynamics and mesodynamics, where fields could be generated only by particles possessing electric charges, magnetic moments, or mesotronic charges and moments, the gravitational field is generated by all kinds of matter. It is essential that Einstein’s equations, unlike the ordinary equations of electro- and mesodynamics, are nonlinear with respect to the (first) derivatives of \(g_{\mu\nu}\). Connected with this circumstance is the possibility of deriving the equations of motion of masses that generate the gravitational field from the field equations, which was first done by Einstein, Hoffmann, and Infeld, and also by Fock.
In the nonrelativistic Newtonian approximation of small velocities and a weak field, on the right there remains the principal term of the tensor \(4\pi\chi m\rho\), while the complex Einstein operator on the left is transformed into the Laplace operator \(\Delta\), acting on the Newtonian potential \(\varphi\), which is a small addition to the constant value \(g_{44}=1\).
In this case, therefore, one obtains the equation of the classical theory of gravitation: \(\Delta\varphi=-4\pi\chi m\rho\) (where \(\rho\) is the density of the mass distribution). For the potential of the field generated by a point mass, we obtain the well-known centrally symmetric solution:
\[ \varphi=\frac{\chi m}{r}. \]
Hence, substituting \(\varphi\) into the expression for the interaction energy of a particle of mass \(M\) with the gravitational field: \(u=-M\varphi\), we obtain Newton’s fundamental law for the interaction energy of two masses
\[ V=-\frac{\varkappa mM}{r}\left(\text{force } F=-\frac{\partial V}{\partial r}=-\frac{\varkappa mM}{r^2}\right). \]
In the general case, the most important centrally symmetric solution of Einstein’s equations, given by Schwarzschild, has the form (see § 10):
\[ ds^2=\gamma c^2dt^2-\frac{1}{\gamma}\,dr^2-r^2d\Omega^2,\qquad \gamma=1-\frac{2\varkappa m}{c^2r}; \]
from this one directly finds the values of \(g_{\mu\nu}\), i.e. the gravitational potentials characterizing the curvature of the metric of particles of mass \(m\).
For a point electrically charged particle the gravitational field will be somewhat different. As Nordström and Jeffery showed, the solution in this case is
\[ ds^2=\gamma_1 c^2dt^2-\frac{1}{\gamma}\,dr^2-r^2d\Omega^2, \]
where now:
\[ \gamma_1=1-\frac{2\varkappa m}{c^2r}+\frac{4\pi\varkappa e^2}{c^4r^2}. \]
For the case of a particle possessing a mesotronic charge \(g\), whose mesotronic field is described, for example, by the scalar de Broglie equation, the gravitational field will also be somewhat different. In this case we have, approximately in the same notation:
\[ \gamma \simeq 1-\frac{2\varkappa m}{c^2r}+A;\qquad A=\frac{\varkappa g^2}{c^4}\frac{e^{-2\varkappa_0 r}}{r^2}. \]
Let us note here as well that, for a more accurate description of gravitation, it is evidently necessary to pass to quantized Einstein equations (see § 13.7).
Taking \(g_{\mu\nu}\) as the “coordinates” of the field, one can, with the help of the Lagrangian function, find the corresponding “momenta” and write the quantum commutation rules. From this, in particular, follows the conclusion that it is impossible to measure simultaneously all the quantities characterizing the gravitational field: \(g_{\mu\nu}\), \(R_{\mu\nu}\), etc. (see § 13.7).
However, in view of the nonlinear character of the equations, it has not been possible to go further here in establishing general relations.
It is far more interesting to consider the case of quantization of a weak gravitational field, passing to the linear approximation, similarly to what was done in the nonquantum Einstein theory. For this, we replace the components of the metric tensor by their constant Galilean (pseudo-Euclidean) values \(g^0_{\alpha\beta}\), plus small additions whose squares may be neglected:
\[ g_{\mu\nu}=g^0_{\mu\nu}+h_{\mu\nu}\qquad \left(g^0_{44}=+1;\; g^0_{rs}=0\text{ for }r\ne s;\; g^0_{rs}=-1\text{ for }r=s\right). \]
Then, in vacuum, for the additions, equations of wave type are obtained.
Consequently, gravitational waves propagate with the velocity of light. This conclusion for the gravitational field was drawn by Poincaré as early as 1905. The linear equations of gravitational waves can be quantized according to the general rules. To each monochromatic wave there corresponds a quantum of the gravitational field, or graviton, devoid of rest mass and possessing spin 2. The matrix form of the equations of a weak gravitational field has not yet been considered. Let us note that at one time Einstein adhered to the idea of the necessity of supplementing the gravitational equations by a term with the so-called cosmological constant \(\Lambda\), which gives the most general form of the equations under the condition of restriction to lower (second) derivatives and the requirement of linearity in the second derivatives:
\[ R_{\alpha\beta}-\frac{1}{2}g_{\alpha\beta}R+\Lambda g_{\alpha\beta}=\varkappa T_{\alpha\beta}. \]
Then, in passing to weak fields, instead of the wave equation we obtain the de Broglie equation
\[ \left(\Delta-\frac{1}{c^{2}}\frac{\partial^{2}}{\partial t^{2}}-\Lambda\right)h_{\rho 0}=0. \]
Obviously, from the point of view of the theory of elementary particles the cosmological constant is connected with the rest mass of the graviton:
\[ \Lambda=\varkappa_{0}^{2}=\left(\frac{2\pi mc}{h}\right)^{2}; \qquad m=\sqrt{\Lambda}\,\frac{h}{2\pi c}\sim 10^{-64} \]
(when the value \(\Lambda \simeq 10^{-54}\) is substituted).
In the static case we hence obtain for the Newtonian potential \(\varphi\) the generalized Laplace–Poisson equation
\[ (\Delta-\Lambda)\varphi=4\pi \varkappa m\rho, \]
with the solution for a point particle of mass \(m\):
\[ \varphi=\frac{\varkappa m}{r}e^{-\sqrt{\Lambda}\cdot r}. \]
A similar modification of Poisson’s equation and of Newton’s law was proposed already in the nineteenth century by Neumann and Seeliger, with the special aim of obtaining a rapid disappearance of gravitational interactions at large distances and thereby ensuring the finiteness of the potential caused by the general distribution of masses in the universe. In the case of the ordinary Newtonian law, for a mean density of matter different from zero, this potential proves to be infinite. The same difficulty occurred in the cosmology of the general theory of relativity without the introduction of the cosmological term \(\Lambda\) and under the assumption of a static universe.
It is interesting to note that the Neumann–Seeliger law generalizes Newton’s law in exactly the same way as Yukawa later generalized the Coulomb potential.
In Einstein’s original static model, the universe turned out to be spatially finite, closed, with radius \(R=\sqrt{\dfrac{1}{\Lambda}}\) and total mass \(M=\dfrac{\pi R c^2}{2\varkappa}\). Substituting for \(R\) the value of the distance to the most remote objects accessible to modern astronomical instruments, \(R\sim 10^{27}\ \text{cm}\), we obtain
\[ \Lambda \sim 10^{-54},\qquad M\sim 10^{55}. \]
However, Einstein’s model proved to be unstable, while the other possible static model, de Sitter’s, which led to an infinite universe, allowed only a state with an average density of matter equal to zero, which is, of course, unsatisfactory.
Without attempting to give a detailed exposition of the cosmological problem, let us merely point out that at present the most plausible hypothesis is considered to be that of Friedmann (to which Einstein also adhered after a sharp polemic), who was the first to propose nonstatic solutions of the equations of the general theory of relativity.
In Friedmann’s theory the cosmological term turns out to be superfluous, since even without introducing it one can obtain solutions corresponding to a finite mean density \(\rho\). The metric form, under the assumption of spatial isotropy, is given by
\[ ds^2=dx_4^2-G^2A^2dr^2, \]
where \(G\) is a function of time alone, and
\[ A=\frac{1}{1+\dfrac{z}{4}r^2}. \]
Here the values \(z=+1,-1,0\) correspond: to a closed world of finite volume (a spherical or elliptic world) with positive radius of curvature, to an infinite world of the corresponding radius of curvature, or, finally, to a flat world of Euclidean type. Substitution of the values \(g_{\mu\nu}\) following from the indicated form of the interval \(ds^2\) into Einstein’s equations leads to the fundamental relation
\[ \frac{z}{G^2}=\frac{1}{3}\varkappa\rho-\delta^2,\qquad \text{where }\delta=\frac{G'}{G}. \]
The Hubble constant \(\delta\), expressing the rate of expansion of the universe, has the value \(\delta=432\ \text{km/sec}\) at a distance of \(10^6\) parsecs, i.e. \(\delta=10^{-17}\ \text{sec}^{-1}\). The enormous success of the model of a nonstatic universe consists in its explanation of the otherwise incomprehensible “recession” of the nebulae, discovered by Slipher and Hubble and manifested in the visible displacement of spectral lines toward the red end, interpreted as the Doppler effect. We see that the type of space depends on the relation between the mean density of matter and the Hubble constant \(\delta\).
Present-day, not very precise values for \(\rho\) lead to the conclusion that
\[ \frac{1}{3}\varkappa \rho < \delta^2,\quad \text{i.e. } z < 1. \]
Consequently, the universe is spatially unbounded (Lobachevsky space).
It is very significant that the model of a nonstatic, expanding universe necessarily leads to a certain special state corresponding to the beginning of expansion; moreover, the period of time that has elapsed since then proves to be equal to
\[ \tau \sim \frac{1}{\delta} \sim 10^{27}\ \text{sec.} \sim 2\cdot 10^9\ \text{years}. \]
This value, comparable in order of magnitude with the mean lifetime of uranium, the time of the earth’s existence, and the periods of various longest-known stellar processes, nevertheless appears exceedingly small. In addition, the very existence of a special state is an undoubted difficulty for the theory. However, as we have already indicated in § 12, the treatment of the universe near the special state by means of macroscopic relativistic equations appears inadmissible, since quantum phenomena involving elementary particles must play an essential role here; taking them into account will undoubtedly make it possible to remove the mentioned difficulties and will advance us in solving the difficult cosmological problem of the structure of the part of the universe known to us as a whole.
5. Lagrangian function
As in classical mechanics, it is expedient to place at the forefront of the entire theory the Lagrangian function \(L\), which must be an invariant with respect to all the transformation groups listed above in § 13.1. Under the restriction to lower derivatives of the \(\psi\)-functions and the linearity of the equations, \(L\) can be constructed as uniquely as the equations themselves. From the Lagrangian function, by means of the variational principle \(\delta\int L\,d\tau = 0\), under the condition that the variations \(\delta\psi\) vanish on the boundaries, we obtain the field equations in the form of the Lagrange–Euler equations:
\[ \frac{\partial}{\partial x_\alpha}\frac{\partial L}{\partial \dfrac{\partial q_s}{\partial x_\alpha}}-\frac{\partial L}{\partial q_s}=0, \]
\[ q_s=\psi,\ \psi^{*},\quad x_\alpha=x,\ y,\ z,\ ct;\quad d\tau=dx,\ dy,\ dz,\ cdt. \]
Moreover, directly from \(L\) one obtains:
1) the momenta canonically conjugate to the generalized coordinates, as which the components of the \(\psi\)-functions are taken:
\[ p_s=\frac{\partial L}{\partial \dfrac{\partial q_s}{\partial t}}. \]
2) Using the invariance of \(L\) with respect to translations, we obtain the canonical energy tensor (more precisely, the tensor of the densities of energy-momentum and stresses)
\[ T_{\alpha\beta} = \frac{\partial q_s}{\partial x_\alpha} \frac{\partial L}{\partial \dfrac{\partial q_s}{\partial x_\beta}} - \delta_{\alpha\beta}L \]
and, at the same time, the energy density \(T_{44}\) and the energy flux \(T_{4s}\).
The expression for the energy density or Hamiltonian function directly generalizes the expression known from point mechanics:
\[ H_{\text{mech}}=p_s q_s-L \qquad \left( p_s=\frac{\partial L}{\partial \dot q_s} \right). \]
Indeed,
\[ T_{44}=H= \frac{\partial q_s}{\partial t} \frac{\partial L}{\partial \dfrac{\partial q_s}{\partial t}} -L = \frac{\partial q_s}{\partial t}p_s-L \]
(as always, summation over two identical indices is implied, in the present case over \(s\)). It is important to note, first, that the canonical energy tensor, generally speaking, will never be symmetric, except in the case of a scalar field. Secondly, the tensor \(T_{\alpha\beta}\) obeys the generalized conservation law:
\[ \frac{\partial T_{\alpha\beta}}{\partial x_\alpha}=0, \]
which contains, as special cases, the laws of conservation of energy and momentum of the field.
3) Along with the canonical energy tensor, one should introduce the “metrical” energy tensor, which, according to Hilbert, is defined by variation of the Lagrangian function, written in generally covariant form, with respect to the components of the metric tensor \(g_{\mu\nu}\):
\[ T'_{\alpha\beta} = \frac{2}{\sqrt{-g}} \frac{\delta L\sqrt{-g}}{\delta g_{\alpha\beta}} \]
(in the final result, in the absence of gravitation, the \(g_{\alpha\beta}\) should be set equal to their constant pseudo-Euclidean values). The metrical tensor \(T'_{\alpha\beta}\) will always be symmetric (by virtue of the symmetry of \(g_{\alpha\beta}\)) and will coincide with the canonical \(T_{\alpha\beta}\) only for a scalar field. \(T'_{\alpha\beta}\) also obeys the conservation law:
\[ \frac{\partial T'_{\alpha\beta}}{\partial x_\alpha}=0. \]
The difference between \(T_{\alpha\beta}\) and \(T'_{\alpha\beta}\) is due to the presence of spin properties in the field, or, in other words, to a nonzero spin of the particles.
4) Generalizing the usual expression for angular momentum,
\[ M=[rp], \]
while taking into account the invariance of the Lagrangian with respect to rotations and Lorentz transformations, we obtain the angular-momentum tensor
\[ M_{\alpha\beta,\gamma} = x_\alpha T_{\beta\gamma} - x_\beta T_{\alpha\gamma} \qquad \left( M_{\alpha\beta,\gamma}=-M_{\beta\alpha,\gamma} \right), \]
which for \(\gamma=4\) gives the density of angular momentum (or “orbital” moment).
5) It is very important that, again most directly with the help of the Lagrangian \(L\), one can, according to Belinfante, construct the general expression for the density of the spin moment of the field \(S_{\alpha\beta,\gamma}\). Then the complete angular-momentum tensor
\[ \Omega_{\alpha\beta,\gamma}=M_{\alpha\beta,\gamma}+S_{\alpha\beta,\gamma} \]
will satisfy the conservation equation
\[ \frac{\partial \Omega_{\alpha\beta,\gamma}}{\partial x_\gamma}=0. \]
6) Variation of the Lagrangian with respect to the electromagnetic potentials leads to expressions for the densities of electric charge and current associated with the given field and particles. Further, again most simply with the aid of the Lagrangian, one obtains expressions for the densities of mesonic charge and current and of other fundamental quantities characterizing the given field.
In close connection with the formalism of the Hilbert–Noether theorem, invariant properties of the Lagrangian are here used, as we have seen, and at the same time conservation laws are also obtained directly for various quantities. To obtain expressions for the momentum tensor, spin, current, etc., from the field equations themselves, and not from the Lagrangian, would constitute a very cumbersome procedure, devoid of internal consistency.
7) Finally, with the aid of generalized coordinates and momenta, the canonical three-dimensional quantum commutation rules of the theory of secondary quantization are obtained (see § 13.7):
\[ [p_s,q_r]=p_s(x')q_r(x)-q_r(x)p_s(x')=\frac{h}{2\pi i}\,\delta(x-x')\delta_{rs}. \]
These formulae generalize to a system with an infinite number of degrees of freedom (the \(\psi\)-field) the usual commutation rules of quantum mechanics of particles:
\[ p_r q_s-q_s p_r=\frac{h}{2\pi i}\,\delta_{rs} \]
(\(x\) denotes all three coordinates).
Thus we become convinced that the Lagrangian function is indeed the most fundamental quantity characterizing all the basic properties of the given field.
Let us illustrate what has been said by a number of examples. For the de Broglie equation the Lagrangian function, in the case of neutral particles described by real \(\psi\)-functions, has the form of the simplest invariant, bilinear with respect to \(\psi\) and depending on first derivatives:
\[ L_B=-\frac{1}{2}\left\{\left(\frac{\partial \psi}{\partial x}\right)^2+ \left(\frac{\partial \psi}{\partial y}\right)^2+ \left(\frac{\partial \psi}{\partial z}\right)^2- \frac{1}{c^2}\left(\frac{\partial \psi}{\partial t}\right)^2- \chi_0^2\psi^2\right\}. \]
Indeed, from the variational principle:
\[ \delta \int L_B\,d\tau=0 \]
we obtain
\[ B\psi=0. \]
For a complex field
\[ L_B=-\frac{\partial \psi^*}{\partial x}\frac{\partial \psi}{\partial x} -\frac{\partial \psi^*}{\partial y}\frac{\partial \psi}{\partial y} -\frac{\partial \psi^*}{\partial z}\frac{\partial \psi}{\partial z} +\frac{1}{c^2}\frac{\partial \psi^*}{\partial t}\frac{\partial \psi}{\partial t} -\chi_0^2\psi^*\psi. \]
The simplest invariant associated with derivatives of the $\psi$-functions in Dirac theory is the expression
\[ L_D = \frac{1}{2}\{\psi^*(D\psi) - (\psi D)^*\psi\}, \]
in which the second term has been added to ensure symmetry with respect to $\psi$ and $\psi^*$. Variation of the Lagrangian $L_D$ gives the Dirac equations: $D\psi=0$; $(\psi D)^*=0$ for $\psi$ and $\psi^*$. Without attempting to give any exhaustive account of the results obtained by applying the general principles of relativistic quantum mechanics to individual fields, we shall confine ourselves to indicating several of the most important quantities characterizing typical fields: the scalar field (spin $0$, Bose particles), the bispinor Dirac field (spin $\frac{1}{2}$, Fermi particles), and the vector Proca or Maxwell field (spin $1$, Bose particles).
Scalar field: the energy tensor is
\[ T_{\alpha\beta} = \frac{\partial \psi^*}{\partial x_\alpha} \frac{\partial \psi}{\partial x_\beta} + \frac{\partial \psi^*}{\partial x_\beta} \frac{\partial \psi}{\partial x_\alpha} - \delta_{\alpha\beta}L. \]
Energy density:
\[ W=T_{44} = \frac{1}{c^2}\frac{\partial \psi^*}{\partial t}\frac{\partial \psi}{\partial t} + \operatorname{grad}\psi^* \operatorname{grad}\psi + \varkappa_0^2\psi^*\psi. \]
It is significant that $W$ turns out to be a positive definite quantity (as also for other particles of integral spin or of Bose type).
Charge density:
\[ \rho = iea\left(\frac{\partial \psi^*}{\partial t}\psi-\psi^*\frac{\partial\psi}{\partial t}\right) \]
(where, under the appropriate normalization, the constant $a$ is equal to $\frac{h}{2mc^2}$) turns out to be indefinite.
The spin density is equal to zero, as was to be expected in view of the presence of only one component of the field $\psi$.
Bispinor Dirac field: the energy density
\[ W=T_{44} = \frac{1}{2ic} \left( -\psi^*\frac{\partial \psi}{\partial t} + \frac{\partial\psi^*}{\partial t}\psi \right) \]
is indefinite, i.e. it can take (as in other cases of particles of half-integral spin of Fermi type) both positive and negative values. On the other hand, the charge density turns out to be positive definite:
\[ \rho = e\psi^*\psi = e(\psi_1^*\psi_1+\psi_2^*\psi_2+\psi_3^*\psi_3+\psi_4^*\psi_4). \]
The spin density, as already indicated earlier, is equal to
\[ S'=\frac{h}{4\pi}\psi^*\sigma\psi. \]
The Proca equations are obtained by variation of the following Lagrangian function, which we shall write for the case of a real field describing neutral particles (possibly neutral mesotrons or the neutretto),
\[ L_P=\frac{F^2-G^2}{8\pi}+\frac{\varkappa_0^2}{\,}\left(\Phi_4^2-\Phi_s^2\right). \]
The electromagnetic field is characterized by two invariants:
1) the “square” of the field tensor
\[ I_1=L_M=\frac{\mathbf E^2-\mathbf H^2}{8\pi}. \]
If this invariant is taken as the Lagrangian function \(L_M\), then, as was shown already by Larmor, we obtain Maxwell’s equations.
2) the “square” of the product of the electric and magnetic fields or of the electromagnetic-field tensor with its dual:
\[ I_2=(\mathbf E\mathbf H)^2. \]
The product \((\mathbf E\mathbf H)\) itself is a pseudoscalar, not a scalar. The use of the invariant \(I_2\) as the Lagrangian function would, evidently, lead to nonlinear equations and is therefore excluded in the usual linear theory of the electromagnetic field. \(I_2\) plays an essential role in the nonlinear generalization of Maxwell’s theory.
Let us give a few more quantities characteristic of a real Proca field:
\[ \text{energy density:}\qquad W=\frac{\mathbf E^2+\mathbf H^2}{8\pi}+\chi_0^2\frac{(\Phi_4^2+\Phi_s^2)}{8\pi}. \]
For \(\chi_0=0\) we obtain the well-known Maxwellian expression. The spin angular-momentum density for a Proca (possibly mesotronic) real field is equal to
\[ \mathbf S[\mathbf F\Phi]. \]
Correspondingly, for the Maxwell field:
\[ \mathbf S=[\mathbf E\mathbf A]. \]
It is not difficult to construct Lagrangians for pseudoscalar, pseudovector, and other fields and to develop for them a complete theory. Einstein’s equations for the gravitational field are also obtained from the corresponding variational principle
\[ \delta\int G\sqrt{-g}\,d\tau=0, \]
where \(g\) is the determinant constructed from \(g_{\alpha\beta}\).
6. Theory of interaction
It is important to emphasize that, under the same conditions which determine the construction of the equations themselves for elementary particles from given \(\psi\)-functions (i.e. invariance with respect to all the transformation groups enumerated in § 13.1, restriction to the lowest derivatives, and linearity of the equations), the general form of the possible law of interaction of all particles with one another also turns out to be unique. The realization of one or another interaction within the limits of the permitted possibilities depends only on the presence, in the particles, of coupling constants, i.e. charges and moments. On this point, once again, relativistic quantum mechanics admits possibilities broader than those observed, since, in principle, it turns out to be admissible for all particles to possess both charges and intrinsic nonkinematic moments of electric and mesotronic types, which is far from always the case. (For example, electrons and positrons possess only kinematic magnetic moments, etc.)
The invariants of the theory of a free particle were constructed from the functions of the particle itself, whereas in the theory of interaction it is necessary to take a mixed invariant composed of the wave functions of both interacting particles.
As an example let us consider the interaction of a proton with an electromagnetic field. Obviously, the only permitted mixed invariant will have, speaking pictorially, the form of a sum of products: the velocity vector \(v\) of the proton multiplied by the vector-potential \(A\) of the field, plus the product of the tensor of moments of the proton \(\sigma_{\alpha\beta}\) by the tensor of the field \(F_{\alpha\beta}\), or:
\[ L_{p\gamma}=-\frac{e}{c}\sum_\mu(\chi^*\alpha_\mu\chi)A_\mu+\mu_0\sum_{\alpha,\beta}(\chi^*\sigma_{\alpha\beta}\chi)F_{\alpha\beta}. \]
Here \(\chi\) are the wave functions of the protons; \(A_\mu=A_1,A_2,A_3,iA_4\) (potential); \(F_{\alpha\beta}=H,iE\) (field); \(\tau_\mu=e\rho_1\sigma\); \(ic\) is the velocity operator of the proton; \(\sigma_{\mu\nu}=\rho_1\sigma_s;-\rho_2\sigma_s\) is the operator of the moment of the proton. For clarity the summation over identical indices is indicated by \(\sum_r\) and \(\sum_{\alpha\beta}\).
\(\mu_0\) is the absolute value of the proper intrinsic, i.e. non-kinematic, magnetic moment of the proton (see § 13.4). In the case of the interaction with the electromagnetic field of a positron or an electron, the second term is absent in the Lagrangian (\(\mu_0=0\) for the electron); for the neutron the first term in the Lagrangian vanishes (the charge is zero). The same expression for the first term \(L_0\) can be obtained if, in the Lagrangian for the free proton \(L_p\), the field is included according to the rule of replacement of derivatives:
\[ \frac{\partial}{\partial x_\alpha}\to \frac{\partial}{\partial x_\alpha}+\frac{2\pi i e}{hc}A_\alpha. \]
This is equivalent to replacing the momenta by the generalized ones
\[ p_\alpha\to p_\alpha+\frac{e}{c}A_\alpha. \]
As a result, in the Dirac equation for the proton we obtain an additional term with the interaction energy, describing the action of the electromagnetic field on the proton:
\[ U=-e\alpha_sA_s+eA_4+\mu_0\rho_3(\sigma H)-\mu_0\rho_2(\sigma E), \]
or, in the nonrelativistic approximation for small velocities:
\[ U=eA_4+\mu_0(\sigma H) \]
(for the electron or positron the second term with the moment is absent; for the neutron the first term with the charge is absent).
Here for \(\sigma\) it is sufficient to take two-row Pauli matrices, and not four-row Dirac matrices. The operator of the interaction energy thus coincides with the known nonquantum expression, with that
the difference being that the velocity and the moment are replaced by the Dirac operators for the proton, as a particle of spin \(\frac{1}{2}\).
Let us emphasize that variation of one and the same mixed Lagrangian \(L_{p\gamma}\) with respect to the wave functions of the proton \(\chi\) leads to additional terms in the “equations of motion,” i.e. the Dirac equation, while variation of \(L_{p\gamma}\) with respect to the potentials \(A_\gamma\) gives terms with the charge, current, and magnetic moment of the protons on the right-hand sides of Maxwell’s equations, describing the generation of the electromagnetic field by protons (or, analogously, by electrons and neutrons).
In a similar way one describes the interaction of all charged non-Dirac particles, for example mesotrons, with the electromagnetic field. Here the action of the electromagnetic field on the charge is again taken into account by replacing the momenta by generalized ones:
\[ p_s \to p_s + \frac{e}{c} A_s;\qquad E \to E - eA_4 . \]
For example, in the interesting case of the action of an electromagnetic (vector!) field on a vector Proca mesotron field, the “equations of motion” of the mesotron field acquire the form (instead of \((P)\) § 13.4):
\[ \left(\frac{1}{c}\frac{\partial}{\partial t}+\frac{2\pi i e}{hc}A_4\right)\mathbf{F} -\left[\left(\nabla-\frac{2\pi i e}{hc}\mathbf{A}\right)\times \mathbf{G}\right] -\chi_0^2\boldsymbol{\Phi}=0; \]
\[ \left(\nabla-\frac{2\pi i e}{hc}\mathbf{A}\right)\mathbf{F} +\chi_0^2\Phi_4=0; \]
\[ \mathbf{F} +\left(\frac{1}{c}\frac{\partial}{\partial t}+\frac{2\pi i e}{hc}A_4\right)\boldsymbol{\Phi} +\left(\nabla-\frac{2\pi i e}{hc}\mathbf{A}\right)\Phi_4=0; \]
\[ \mathbf{G} -\left[\left(\nabla-\frac{2\pi i e}{hc}\mathbf{A}\right)\times\boldsymbol{\Phi}\right]=0, \]
where the sign \(\times\) denotes the vector product, \(\mathbf{F}\), \(\mathbf{G}\) are the quasi-electric and quasi-magnetic components of the mesotron field, \(\Phi_\nu=\boldsymbol{\Phi}\), \(\Phi_4\) is its potential, \(A_\nu=\mathbf{A}\), \(A_4\) is the potential of the electromagnetic field.
The interaction of all particles with the mesotron field is described in a very similar manner. If the mesotron is described by a scalar material function \(\psi=\Phi_0\), then the mixed Lagrangian taking into account the interaction, for example, of a proton with mesotrons is equal to: \(L_{p\mu}=g\Phi_0\chi^{*}\rho_3\chi\) (where \(\chi\) are the proton functions). Hence in the Dirac equation for the proton, and in general for the nucleon, there will appear the additional interaction term:
\[ U=g\Phi_0\rho_3. \]
or at small velocities \(U=g\Phi_0\).
For light particles (electrons, positrons, neutrino) we shall have only another constant—\(g'\). In the interaction with charged mesotrons, described by complex functions, in view of the Hermiticity of the interaction energy, one must take the sum of the complex conjugate
... quantities: \(\Phi + \Phi^*\). In doing so one must also take into account the possibly complex character of the coupling constant, or charge \(g\), as well as the circumstance that, upon the emission or absorption of a charged mesotron, nucleons change sign (a proton becomes a neutron, a neutron a proton).
To take account of the latter fact, one must introduce the corresponding operators for transforming nucleonic proton functions into neutron functions and conversely, \(Q\) and \(Q^*\). Then the complete expression for the energy of interaction of nucleons with charged scalar mesotrons is written in the form:
\[ U = g\Phi_0 \rho_3 Q + \text{complex conjugate expression}. \]
The action of a vector real mesotron field on nucleons and light particles is completely analogous to the action on them of the electromagnetic field. It is sufficient merely to replace the electromagnetic quantities \(A_\nu,\ \mathbf E,\ \mathbf H\) by the corresponding mesotronic ones\(^*\) \(\Phi_\nu,\ \mathbf F,\ \mathbf G\). In the case of a charged vector field, one must introduce the operators \(Q\) and form a Hermitian expression as the sum of two complex conjugate terms.
The expression for the energy of interaction of particles with a weak gravitational field \(h_{\mu\nu}\) can be constructed by the same method (see § 10). In the general case, however, the interaction of particles with the gravitational field is taken into account by rewriting the equations, or the corresponding Lagrangians, in generally covariant form with the aid of the components of the metric tensor \(g_{\mu\nu}\), which at the same time are components of the gravitational potentials (see § 10).
It is also not difficult, by the same method, to construct the expression for the energy of interaction of all particles with pseudoscalar and pseudovector fields.
We shall confine ourselves here to indicating the expression, important for the theory of mesotrons and nuclear forces, for the energy of interaction of a nucleon with a pseudoscalar (let us say, pictorially, mesotronic) field described by the function \(\Phi\):
\[ U = g_1 \rho_1 \Phi + g_2 \sigma_1 \operatorname{grad}\Phi - g_2 \rho_2 \frac{\partial \Phi}{\partial t}, \]
where \(\rho_1,\ \sigma_1,\ \rho_2\) are the pseudoscalar and pseudovector Dirac matrices; \(g_1\) and \(g_2\) are coupling constants playing the role of charge and of a dipole quasi-magnetic moment.
The enumerated types of interaction referred above all to the action on any particles (Bose and Fermi) of Bose fields which, in the case of the reality of the field, have a classical analogy: electromagnetic, neutral mesotronic, gravitational—that is, to the interaction of any particles with photons, neutrettos, and gravitons. In this case the particles have integral spin: photons, meso-
trons, gravitons, can be emitted or absorbed one particle at a time*).
Alongside such a connection of particles (fields) with other fields (particles), which makes possible the emission and absorption of field quanta or particles (for example, the emission of photons by charged particles, or of pairs \(e,\nu\) by nucleons, etc.), the interaction of particles with one another through some field plays a fundamental role, which also corresponds to the concept of “interaction” in the usual visual sense. This is the transfer of interaction by electromagnetic, neutral mesotron and gravitational fields and, moreover, by fields having no classical analogy: by a charged mesotron field, by the field of pairs of light particles \((e_-,\nu)\), \((e_+,\nu)\), and by any other conceivable combinations, for example pairs \((e_-, e_+)\), or even possibly pairs of nucleons \((p,n)\).
Let us consider, in particular, the interaction of a positron with a proton (or of any two charged particles) in the static case. It arises because \(p\) produces an electric field, while \(e_+\) absorbs it, and conversely. The binding energy \(U\) of the positron with the electrostatic field is equal to: \(U=+eA_0\). The potential of the field produced by a point proton, according to Laplace’s equation (see §13.4) \(\Delta A_0=-4\pi e\rho\), is \(A_0=e/r\); hence the interaction energy of two charged particles is \(V=-\dfrac{e_1e_2}{r}\) (therefore the force, in accordance with Coulomb’s law, is \(F=-\dfrac{dV}{dr}=\dfrac{e_1e_2}{r^2}\)).
Quantum mechanics again obtains the same expression by calculation according to perturbation theory, in the second approximation of the following two-step process: 1) a positron or electron, etc., emits a quantum of the field (a pseudophoton), 2) the proton absorbs it, and conversely.
In a completely analogous way, in classical mechanics, and likewise exactly in quantum mechanics, Newton’s law of gravitation is obtained, owing to the coupling of particles with the weak gravitational field, or through the transfer of interaction by gravitons (see §10).
The greatest interest has recently been acquired by the calculation of the interaction energy between two nucleons due to their coupling with the mesotron field in the static case. Substituting into the interaction energy of the first nucleon with the mesotron, in particular, scalar material field \(U=-g_1\Phi_0\), the value of the field potential,
*) In addition, all particles (for example, nucleons) may be coupled with the fields of pairs of other Fermi particles (for example, electrons and neutrinos, or positrons and neutrinos), which has no precise classical analogy. In particular, the interaction energy of nucleons with the field of pairs of light particles has, in the nonrelativistic approximation, the form:
\[ U=g_F\psi_\nu^{+}\psi_e+\text{complex conjugate expression} \]
(here \(g_F\) is the Fermi constant from the theory of \(\beta\)-decay).
generated by the second nucleon
\[ \Phi_0=\frac{\xi_2 e^{-\chi_0 r}}{r}, \]
we obtain the desired interaction energy in the form:
\[ V=-\frac{g_1 g_2}{r} e^{-\chi_0 r}. \]
Exactly the same expression is obtained also in quantum theory. \(V\) decreases rapidly at large distances, \(r \gg \frac{1}{\chi_0}\), and thereby correctly conveys the short-range character of nuclear forces. When the interaction is transmitted by a material mesotron vector field, the interaction energy of two nucleons, calculated according to classical or quantum theory, includes important noncentral and spin terms and has a somewhat cumbersome form:
\[ V=\frac{g_1 g_2}{r} e^{-\chi_0 r} + f_1 f_2 \frac{e^{-\chi_0 r}}{r} \left\{ (\sigma_1,\sigma_2)\left(\chi_0^2+\frac{\chi_0}{r}+\frac{1}{r^2}\right) - \frac{(\sigma_1,r)(\sigma_2,r)}{r^2} \left(\chi_0^2+\frac{3\chi_0}{r}+\frac{3}{r^2}\right) \right\}; \]
\[ \left(f_1 \sim \frac{g_1}{\chi_0};\quad f_2 \sim \frac{g_2}{\chi_0}\right). \]
At small distances the principal terms \(V \sim \frac{\mathrm{const}}{r^3}\) have precisely the form of the interaction energy of two magnetic dipoles. In view of this excessively rapid increase of \(V\) as \(r \to 0\), stable orbits in the field of such forces are impossible, which is a very substantial difficulty for the theory of nuclear mesotron vectors (and also, as can be shown, pseudoscalar) forces. In the case of an interaction carried by charged mesotrons, it is necessary to multiply \(U\) by the Heisenberg operator of charge exchange (or of the coordinates and spins) of the nucleons:
\[ P_H = Q_1 Q_2^* + Q_1^* Q_2 \]
(where the indices 1, 2 refer to the two nucleons). The interaction energy is calculated analogously in all other cases.
To summarize, one may say that the theory of interaction makes it possible successfully to explain all interactions of the electromagnetic and gravitational type and, qualitatively, mesotron nuclear forces, and also points to the fundamental possibility of a great variety of interactions of new types among all particles, either directly or through other fields or particles. All this is a major achievement of the theory. However, as was already indicated above, independently of the difficulties in the theory of vector and pseudovector nuclear forces, the theory of interaction in all cases leads in higher approximations to divergences that have no physical meaning. A brief analysis of the difficulties is given in § 14.
7. Secondary Quantization and Statistics
Electromagnetic waves proved to be a generalization of geometrical optics, just as $\psi$-waves are a generalization of ordinary mechanics. The further quantization of the electromagnetic field, or of any other field, therefore bears the name “secondary quantization.”
The basic idea of this powerful method consists in the “atomization” of fields, i.e., in assigning particle waves, for example, electromagnetic waves—to photons, a weak gravitational field—to gravitons, Dirac $\psi$-waves—to electrons and positrons, etc.; in other words, in the strict formulation of corpuscular-wave dualism.
Indeed, expanding in a Fourier series, with the cyclic condition of periodicity, the $\psi$-function, for example, of a scalar field, we obtain, using de Broglie’s equation,
\[ \psi=\frac{1}{L^{3/2}}\sum \sqrt{\frac{hc}{2\pi K}}\left(a_k e^{-icKt+i(kr)}+b_k e^{icKt+i(kr)}\right), \]
where $L$ is the length of the period, $a_k$ and $b_k$ are Fourier coefficients, from which the coefficient $\sqrt{\dfrac{hc}{2K}}$ has been singled out for convenience; $k=k_x,k_y,k_z$; $k_x=\dfrac{2\pi n_x}{L}$, etc., with $n_x,n_y,n_z$ being integers, $K^2=k^2+\varkappa_0^2$; the summation is over the three indices $n_x,n_y,n_z$.
Substituting this expansion of $\psi$ into the expressions for the energy density and the charge density and integrating over all space, we obtain the value of the total field energy $W$ and the total charge $e'$ in the form:
\[ W=\sum_k \varepsilon_k(a_k^*a_k+b_k b_k^*) \qquad (\varepsilon_k=hcK \text{ is the energy of the } k\text{-th state}); \]
\[ e'=e\sum_k (a_k^*a_k-b_k b_k^*) \qquad (e \text{ is the elementary charge}). \]
Hence it is immediately clear that the amplitudes of the waves $a_k$ and $b_k$, which previously were completely arbitrary, must now be restricted by the condition that they be equal to integers
\[ a_k^*a_k=n_k;\qquad b_k b_k^*=n_k'. \]
Then the energy of the field waves will be equal to the sum of the energies of particles of two kinds ($a$ and $b$) in different quantum states; moreover, in the $k$-th state there are $n_k$ particles of one kind and $n_k'$ particles of the other kind. The formula for the charge shows that the particles of the two kinds have opposite charges:
\[ W=\sum_k \varepsilon_k(n_k+n_k'), \]
\[ e'=e\sum_k(n_k-n_k') \]
(discarding an infinite constant).
The expressions for the Poynting vector, the current, etc., confirm the correspondence that has been made. Analogous results are obtained for all other fields as well.
It is clear from this that the amplitudes of the waves \(a_k\), \(b_k\), and, consequently, the \(\psi\)-waves themselves should be understood in a generalized sense as certain quantum operators obeying special “commutation rules,” leading to integral eigenvalues for the quantities \(n_k\) and \(n'_k\). Here there naturally arises the question of restricting the spectrum of \(n_k\), \(n'_k\) in the case of Fermi statistics to the values 0 and 1. As was first shown in 1927 by Dirac, who constructed the theory of secondary quantization of the electromagnetic field, and developed by Pauli and Weisskopf, who quantized the scalar field, the commutation rules
\[ [a_k,a^*_{k'}]_{-}=a_k a^*_{k'}-a^*_{k'}a_k=\delta_{kk'}, \qquad [b_k,b^*_{k'}]_{-}=b_k b^*_{k'}-b^*_{k'}b_k=\delta_{kk'}, \qquad a_r b_s=b_s a_r=0 \]
lead to Bose statistics, i.e.
\[ a^*_k a_k=n_k=0,1,2,\ldots \]
\[ a_k a^*_k=n_k+1=1,2,\ldots \quad \text{etc.} \]
In this case the Fourier coefficients split in a remarkable way into an amplitude and an operator phase, for example \(a^*_k=\sqrt{n_k}\,e^{-i\vartheta_k}\). The phase \(\vartheta_k\) turns out to be connected with \(n_k\), in essence, in the same way as the coordinate \(q_k\) is connected with the momentum:
\[ \vartheta_k=\frac{1}{i}\frac{\partial}{\partial n_k}; \qquad n\vartheta-\vartheta n=i, \]
whence it is seen that the phase \(\vartheta_k\) is the operator of change of the number of particles. As a result, the quantities \(a_k\) prove to be absorption operators, and \(a^*_k\) creation operators for particles. Thus secondary quantization proves necessary in the description of all processes of creation and annihilation of particles. The possibility of secondary quantization of all fields means the fundamental possibility of creation and destruction of all possible fields and particles.
As was subsequently shown by Jordan and Wigner, Fermi statistics is obtained with the following commutation rules:
\[ [a_k,a^*_{k'}]_{+}=a_k a^*_{k'}+a^*_{k'}a_k=\delta_{kk'}. \]
In the case of a continuous spectrum (with the corresponding expansion in a Fourier integral), the \(\delta\)-symbols in the commutation rules are replaced by Dirac \(\delta\)-functions: \(\delta(k-k')\).
Subsequently Pauli proved the following very important theorems.
A. For all particles of integral spin the energy of the field is a positive-definite quantity, but the charge and current will be indefinite-
(for example, for de Broglie or Proca mesotrons). For particles of half-integral spin, on the contrary, the field energy is indefinite, while the charge and current are definite (for example, for Dirac spinor electrons).
B. It turns out that particles of integral spin, i.e. of positive-definite energy, can be quantized only according to Bose statistics; Fermi quantization is algebraically inadmissible. On the other hand, particles of half-integral spin admit, from the algebraic point of view, both statistics. However, if one requires positivity of the energy, then for them only Fermi statistics is admissible. In this case the charge and current become indefinite and equal to the difference of the charges and currents of oppositely charged particles.
For example, for the field of Dirac spinor particles we have:
\[ W=\sum E_k(a_k^{*}a_k-b_k^{*}b_k), \]
\[ e'=e\sum_k(a_k^{*}a_k+b_k b_k^{*}). \]
Quantization according to Fermi statistics, taking into account the Pauli principle, gives: \(a_k^{*}a_k=n_k(=0,1)\) (the number of particles of the 1st kind), \(b_k^{*}b_k=1-n'_k(=1,0)\) (the number of “holes,” or free places, in the states of particles of the 2nd kind). Finally we obtain:
\[ W=\sum \varepsilon_k(n_k+n'_k) \]
\[ e'=e\sum_k(n_k-n'_k). \]
(discarding an infinite constant).
We see that the theory of secondary quantization naturally and rigorously arrives at Dirac’s hypothesis on the necessity of introducing antiparticles of the opposite sign of charge (positrons), the number of which is determined as the number of “holes” in the states of negative energy.
From the quantization of Fourier amplitudes one can pass directly to the quantization of the \(\psi\)-functions themselves. The components of the \(\psi\)-functions at different points and at different instants of time will now obey certain permutation rules, i.e., generally speaking, they will not all be simultaneously measurable. For example, for a scalar field we have the following permutation rules:
\[ \psi(\mathbf r,t)\psi^{*}(\mathbf r' t')-\psi^{*}(\mathbf r' t')\psi(\mathbf r,t) = -\frac{i\hbar c}{2\pi}D(\mathbf R,T), \]
where \(\psi(\mathbf r,t)\) is the value of \(\psi\) at the point with radius vector \(\mathbf r\) at the instant \(t\), \(\mathbf R=\mathbf r-\mathbf r'\), \(T=t-t'\). The first Pauli function \(D_1\), which is an invariant solution of the de Broglie equation and is essentially the image
...which determines all four-dimensional permutation rules, not only in the scalar case, is equal to:
\[ D_1(\mathbf{R},T)=\frac{1}{(2\pi)^3}\int \frac{d\mathbf{k}}{K}\,e^{i(\mathbf{kR})}\sin cKT. \]
It is very important that \(D\) turns into the Dirac \(\delta\)-function, i.e., has a singularity on the light cone at \(R=\pm cT\) and vanishes outside the light cone, but is not equal to zero inside it. The vanishing of the function \(D\) outside the light cone corresponds to the fact that quantities situated at distances so large that measurements of them cannot interfere with one another are always commensurable, and their operators commute (since information about the measurements cannot propagate with superluminal velocity). A priori one could have defined the four-dimensional permutation rules through the second Pauli function (another invariant solution of the de Broglie equation)
\[ D_2=\frac{1}{(2\pi)^3}\int \frac{d\mathbf{k}}{K}e^{i(\mathbf{kR})}\cos cKT, \]
which nowhere vanishes, but then we would come into conflict with the aforementioned exclusion of superluminal velocities. Nevertheless, as Blochintsev has recently noted, an additional analysis is required of the applicability of the quantity \(D_1\) in secondary quantization.
Alongside the quantization of Fourier amplitudes and four-dimensional quantization, one can construct the theory of three-dimensional canonical quantization. For this it is necessary either to put \(t=t'\) in the four-dimensional permutation rules, i.e., to consider all quantities at one and the same instant of time, or else to write directly the three-dimensional commutation (or anticommutation) relations between the generalized “coordinates,” taken to be the components of the \(\psi\)-function, and the “momenta” canonically conjugate to them. Three-dimensional quantization, which now plays no important role, was developed by Heisenberg and Pauli even before the establishment of four-dimensional quantization. For example, for a scalar field we have:
\[ [\psi(\mathbf{r}),\psi(\mathbf{r}')]_- = 0;\qquad \left[\frac{\partial\psi(\mathbf{r})}{\partial t},\psi^*(\mathbf{r}')\right]_-=-ihc^2\hat{\delta}(\mathbf{r}-\mathbf{r}'). \]
As was already emphasized above, secondary quantization leads to the impossibility of the simultaneous measurement of all components of the \(\psi\)-functions. In particular, it is impossible simultaneously to measure exactly all components of the electric and magnetic fields, or all components of the gravitational field. It is interesting to note that taking into account the creation and annihilation of pairs of particles leads to the conclusion—one not yet analyzed to the end—that in relativistic quantum mechanics even individual components of the field (for example, of the electric or magnetic field) cannot be measured absolutely exactly. Similar “individual errors” are also obtained for the coordinates, generalizing the “paired” Heisenberg errors for coordinates and momenta,
or for pairs of field components (Ambartsumian and Ivanenko, Schrödinger, Jordan and Fock, Halpern and Johnson.)
Such, in its most general features, is modern relativistic quantum mechanics, which is, as we see, a very consistent, well-constructed theory, successfully explaining a large number of the most fundamental properties of elementary particles (spin, statistics, the ability to be created and annihilated, the ability to interact with other particles and to move in a definite manner). This general theory underlies successful calculations of an enormous number of the most varied effects. Therefore there is no doubt as to the general correctness of the basic propositions of the theory and their applicability to an enormous range of phenomena. At the same time, modern relativistic quantum mechanics is not the final theory of elementary particles.
First, relativistic quantum mechanics does not quite exactly correspond to the system of known particles, proving in part to be broader than the latter, and in part, on the contrary, leaving unexplained a number of properties of particles*).
Second, the theory is as yet unable to explain the values of the masses of particles and of the coupling constants (charges and moments).
Third, the theory leads to a number of difficulties that have repeatedly been mentioned even within the sphere of its applicability; the chief among them are in part connected with one another (the infinite energy of any field produced by point particles, the divergence of the higher approximations of perturbation theory, infinities in the theory of pairs and of the vacuum).
We shall now pass, in conclusion, to a brief analysis of these difficulties and of the newest paths proposed for their elimination.
*) In speaking of the necessity of deriving the values of the dimensionless constants of the fine structure \(\alpha, \beta, \gamma\) in a future theory, we of course mean precisely all the difficulties in this respect. The boldest hypothesis, Eddington’s, attempted to declare the value of the reciprocal Sommerfeld constant to be an integer:
\[ \frac{1}{\alpha}=137,000 \]
(the best experimental data give
\[ \frac{1}{\alpha}=137,02 \]
) and to connect it with a certain “permutation energy” of two charges, using the 16 “degrees of freedom” inherent in a Dirac particle. Then Eddington tried to decipher \(\frac{1}{\alpha}\) as
\[ \frac{n(n+1)}{2}+1=\frac{16 \times 17}{2}+1=137. \]
Subsequently Eddington tried to derive the ratio of the proton and electron masses from the equation:
\[ 10m^2-136\,mm_0+m_0^2=0, \]
for which the ratio of the roots is 1836.5 (sic!). (\(10\) is the number of components of the metric tensor,
\[ m_0=\sqrt{\frac{N}{R}}, \]
where \(N\sim10^{79}\) is the number of protons in the known part of the universe, \(R\) is its radius of curvature \(\sim10^{27}\) cm.)
The sheer unconvincingness of these hypotheses, typical of a certain “Cambridge” tendency of theoretical thought, carried away by jong-
§ 14. DIFFICULTIES OF THE THEORY
1. The problem of proper mass. The field hypothesis
The value of the proper mass, or mass in the state of rest, is without doubt the most characteristic individual attribute of an elementary particle. (In what follows, for brevity, we shall simply say mass.) Indeed, various transformation properties of wave functions, as well as the values of spin, may be inherent in the most diverse particles. The electron, positron, and also the nucleons—the proton and neutron—and, in all probability, the neutrino, are described by spinor functions and all possess spin \(1/2\).
The electric charge likewise is not characteristic of any single particle.
Two different particles of one and the same mass have not so far been found. As for a vanishing rest mass, according to modern views, besides photons gravitons also possess it, and possibly the neutrino.
by numbers, is especially clearly apparent from the impossibility of including in Eddington’s scheme the new particles: mesotrons and neutrons.
It should be noted that preliminary attempts to connect atomic and cosmological quantities were undertaken by a number of other authors. In particular, Dirac expressed the idea that all dimensionless “large” constants are connected by simple relations with coefficients of order 1. Let us take, for example, the ratio of the constants \(\alpha, \gamma\), i.e. the ratio of the electrical and gravitational forces between two electrons
\[ \frac{\alpha}{\gamma} = \frac{2\pi e^2}{hc} : \frac{2\pi \chi m^2}{hc} = \frac{e^2}{\chi m^2} = 10^{41} \]
(this ratio will not change substantially if we take the force of attraction between an electron and a proton, or two other elementary particles). On the other hand, let us take the ratio of the expansion time of the universe (according to the hypothesis of an expanding universe of Friedman and the astronomical data of Hubble), equal to \(2\cdot 10^9\) years, to a “natural” unit of time, for which one may take one of the following expressions:
\[ \tau_1=\frac{e^2}{mc^3},\qquad \tau_2=\frac{e^2}{Mc^3},\qquad \tau_3=\frac{h}{mc^2},\qquad \tau_4=\frac{h}{Mc^2}, \]
and also, evidently,
\[ \tau_5=\frac{e^2}{\mu c^3};\qquad \tau_6=\frac{h}{\mu c^2},\qquad \tau_7=\frac{g^2}{Mc^3} \quad \text{and so on.} \]
Then we again obtain a “large” number, of the same order as \(\dfrac{\alpha}{\gamma}\):
\[ \frac{\tau}{\tau_j}\simeq 7\cdot 10^{38},\qquad \text{in general}\qquad \frac{\xi}{\xi_a}\sim 10^{38}-10^{40}. \]
From this Dirac makes the interesting supposition that the world constants—namely, in the present case, the constant of gravitation—decrease proportionally to time. Other “large” constants, which turn out to be of the order of \(10^{78}\) (obviously the number of particles), must vary proportionally to the square of the time. Subsequently Dirac connects his hypotheses with Milne’s “kinematic” cosmology.
Lighter particles are emitted and absorbed by heavier ones and can, therefore, most directly transmit the interaction between them. Gravitons transmit interaction between all particles; photons—between all charged particles; mesotrons, intermediate in mass, account, in the main, for the bond between nucleons. Heavier particles can, upon decay, turn into lighter ones. Therefore the rest mass leads to the most natural classification of particles, the one indicated in the introduction. This classification corresponds to the choice of mass (or atomic weight) as the principal, so to speak, “Mendeleevian” attribute, or corresponds, if one likes, to the still earlier Newtonian definition of mass as a “measure” of matter.
At the same time, there is still no theory whatever of proper mass, and it is entirely unclear to us not only where the concrete values of particle masses and their ratios come from, but also the very nature of proper mass. In view of this, one should be very cautious in choosing the attribute for the classification of particles, since, in view of the various transformations of particles, the differences between light and heavy particles are to a considerable extent erased. Indeed, photons can turn into electron–positron pairs or into two mesotrons. Mesotrons are not only emitted by nucleons but, in principle, can themselves produce pairs of nucleons, etc. Mesotrons can transmit interactions not only between nucleons, but also determine, albeit very weakly, the nuclear interaction between light particles*).
Emphasizing the absence of a completed theory, one cannot fail to dwell on the sole hypothesis that attempts to explain the nature of proper mass and to estimate its order of magnitude. We have in mind the hypothesis of “field” mass, the basic idea of which consists in the assumption that the proper energy of a particle (or, what is the same thing, allowing for the constant coefficient \(c^2\), its “proper” mass) is equal to the energy of the field generated by the particle. Let us consider the classical theory of the electron, for which the “field” hypothesis was first developed by J. J. Thomson, Lorentz, Abraham, and Poincaré. In the static case the field of a point charge at a distance \(r\) is equal to \(E = \dfrac{e}{r^2}\), i.e., it tends to \(\infty\) as \(r\) approaches 0.
The total energy of the electrostatic field is evidently also infinite.
*) If it is not considered impossible that the basic attribute of particles is spin and the transformation properties of wave functions, then one must evidently speak, for example, of a spin particle of “half” \(s = \dfrac{1}{2}\), which may be in a positron, proton, neutron, neutrino, or electron state. On the other hand, taking rest mass as the basis, we may expect that particles of a given type, for example nucleons, can be found in excited spin and charged states of somewhat different mass.
Indeed:
$$ W=\int_0^\infty \frac{E^2\,d\tau}{8\pi}=\frac{e^2}{2}\left.\frac{1}{r}\right|_0^\infty=\infty. $$
In other words, the energy of interaction of the charge with itself is equal to $\infty$: $\dfrac{e^2}{r}$ ($\lim r\to 0$), or “self-action.”
To eliminate this difficulty with infinity, it is proposed to pass from a point particle to a little sphere of radius $r_0$ with charge distributed, say, over its surface. Then for small $r$ inside the sphere the field vanishes, and for the total energy of the field we obtain the finite expression
$$ W=\int_{r_0}^{\infty}\frac{E^2\,d\tau}{8\pi}=\frac{e^2}{2r_0}, $$
which we may equate to the particle’s own energy $W=mc^2$.
Let us emphasize that, independently of the attempt to explain mass “by the field” through the absurd conversion to infinity of the energy of the field produced by the particle, there remains a difficulty that must be eliminated.
Since, obviously, no literal meaning can be attached to the “parts” of the electron or to the “parts of the electronic charge” and to its distribution over the surface or throughout the whole volume, we shall admit the realization of the indicated result in order of magnitude, i.e., put:
$$ W=\frac{e^2}{r_0}=mc^2. $$
Thereby the classical electrical radius of the electron is determined:
$$ r_0=\frac{e^2}{mc^2}=2.8\cdot 10^{-13}\ \text{cm}. $$
The indisputable success of the field hypothesis in its application to the electron consists in the fact that the value of $r_0$ obtained does indeed correspond to certain effective “dimensions” of the electron, observed in experiments on scattering and collisions, without any assumptions about a non-point particle. For example, according to Thomson, the effective cross section for the scattering of light by a point electron at large wavelengths is equal to $\sigma=\dfrac{8\pi}{3}r_0^2$.
Thus, serious significance should be attached to the results obtained, the meaning of which we shall summarize in the words: the proper energy of the electron, according to the hypothesis of field mass, is basically electromagnetic in nature. The qualification “basically” is added here in view of the neglect made of gravitational and mesotronic
parts of the electron energy which, however, for a sphere of radius \(r\sim 10^{-13}\) cm and mass \(m=9\cdot 10^{-28}\) g are very small.
However, at this point, in essence, certain successes of the hypothesis of field mass come to an end, quite apart from the fact that this hypothesis is highly preliminary, in view of the crude and clearly inadmissible assumption of a “sphere” of radius \(r_0\), which violates relativistic invariance. Moreover, even if the derivation of the rest mass is accepted, the theory is nevertheless unable to give the correct value of the electron momentum.
Let us now turn to quantum theory, which, as is well known, again leads to Coulomb’s law and thereby to the former expression for the energy of the field, or “interaction,” of a charge with itself. Since in quantum theory the hypothesis of an electron-sphere is, of course, wholly inadmissible, the divergence of the proper electrostatic energy of a point charge, due to the emission and absorption of photons by the charge itself, i.e. to “self-action,” is here an especially serious difficulty, quite independently of any hypotheses concerning the nature of mass. Since electrostatic energy is due to the longitudinal part of the electromagnetic field, it is customary to call it “longitudinal.” In addition to the infinite “longitudinal” energy (or, according to the field hypothesis, “longitudinal” mass), quantum mechanics leads to an infinite “transverse” mass, due to the transverse part of the electromagnetic field associated with the electron. This energy, as it turns out, is connected with field fluctuations (Heitler). The transverse mass is of a purely quantum character.
For the positron, charged mesotrons, and the proton one may repeat the same reasoning. For the classical electromagnetic radius of the mesotron we obtain
\[ r_\mu=\frac{e^2}{\mu c^2}\sim 10^{-15}\ \text{cm}, \]
and for the classical electromagnetic radius of the proton
\[ r_p=\frac{e^2}{\mu c^2}\sim 10^{-16}\ \text{cm}. \]
These values in no way correspond to the empirical effective “sizes” of these particles, manifested in collisions, since the effective cross sections prove to be of the order of the same radius as that of the electron \((r_0\sim 10^{-13}\text{ cm})\). On the other hand, if the mesotron and the proton possessed “radii” \(r_0\), then their field electromagnetic masses would be equal to the mass of the electron. Therefore, from the point of view of the field hypothesis, one should conclude that the proper masses of mesotrons and protons are, in the main, not of electromagnetic origin.
Since nucleons are particles connected mainly with the mesotron field, owing to their nuclear “quasi-charges” \(g\) and dipole moments \(f\), which are of fairly considerable magnitude, it is natural to assume that not only the interaction but also the “self-action” of nucleons has, chiefly, a mesotronic character.
Let us give the most preliminary considerations on this matter within the framework of the classical theory of the mesotron neutral field. From the infinite energy of the mesotron field produced by a point nucleon, one can pass to a finite value by introducing a “classical mesotron radius” of the nucleons, \(R_0\). It then turns out that the quasi-electric part of the field is insufficient to explain the mass of the nucleons, while the quasi-magnetic part of the field gives, in order of magnitude, the energy expression
\[ W \sim \frac{f^2}{R_0^3} \sim \left(\frac{g}{\chi_0 R_0}\right)^2 \frac{1}{R_0}, \]
which may be equated to \(Mc^2\), where \(M\) is the mass of the nucleon.
In view of the absence of exact values of the quasicharge, \(g/\chi\), and of the mass of the mesotron, the estimate can have only the crudest character, but in any case it gives a reasonable order for \(R_0 \sim 10^{-13}\) cm. Therefore, with all caution, one may say that, from the standpoint of the field hypothesis, the mass of the nucleon is, in the main, of mesotron-dipole quasi-magnetic origin. As for the mass of the mesotron itself, the most that the field hypothesis can at present count on here is a qualitative reduction of the mass of mesotrons to some field. If the known aggregate of elementary particles and fields is to be closed in this respect (i.e., not require a new field for the explanation of the mass \(\mu\)), then it remains to try to explain the mass of the mesotron by combined fields, for example, by pairs of light particles. It would be premature to discuss this point in more detail. Let us also note that, from the standpoint of the field hypothesis, the neutrino, possessing a small mesotron charge \(g'\), must possess a finite mass of mesotron character.
Closely connected with the question of the mesotron nature of nucleons is the problem of their intrinsic nonkinematic magnetic moments.
According to Wick’s hypothesis (put forward by him still as applied to our model of nuclear pair \(\beta\)-forces), the magnetic moment of the neutron is due to the magnetic moment of the mesotron, which is emitted and reabsorbed by the neutron. Thus, for some time the neutron is in a “dissociated” state \((n \to p + \mu_-)\), and the magnetic moment of the system is determined mainly by the negative mesotron, which, owing to its small mass, has a more considerable kinematic magnetic moment.
An analogous argument leads to the presence in the proton of an intrinsic positive moment. Quantitative calculation again leads to infinities, since the matter obviously concerns processes of the “self-action” type. Preliminary estimates by introducing a radius of the nucleons give for the magnetic moments of the nucleons the correct order of magnitude and the correct signs.
It is curious to note that, if mesotrons prove to possess only kinematic magnetic moments, then this will mean
the full implementation of Ampère’s program of reducing magnetism to the action of currents, or, in quantum language, to effective kinematic moments, since the moments of nucleons will in the final analysis turn out to be due to the kinematic moments of mesotrons*).
2. New Hypotheses
a) Nonlinear electrodynamics.
Let us now turn to the consideration of other ways of eliminating the difficulty with infinite self-energy and to new attempts in the theory of mass. As a first item, let us consider the nonlinear equations in electrodynamics, first introduced by Mie (1912), who, however, made a number of errors, and by Born (1934) with the special aim of eliminating the infinite longitudinal self-energy. It should, however, be emphasized that, even after eliminating the longitudinal classical part, we still remain face to face with the divergence to infinity of the transverse, quantum part of the self-energy.
The starting point of Born’s nonlinear theory, developed by him together with Infeld, is the choice of a new, specially selected invariant Lagrangian function.
Born’s scheme is formally irreproachable in the sense of satisfying all requirements of invariance, and his theory does indeed lead to a finite longitudinal energy of a point electron; but the choice of the combination of the basic invariants is completely arbitrary and therefore lacks persuasiveness. Since any function of invariants is also an invariant, it is clear that the condition of invariance alone is insufficient for constructing a theory. What is involved here is the construction of a Lagrangian from both invariants of the electromagnetic field: \(I_1\) and \(I_2\) (see § 13.5). Wishing to prevent the field from becoming infinite at small distances and guided by the idea of the existence of some maximum admissible field \(b\), Born initially took as a basis the Lagrangian
\[ L=-\frac{b^2}{4\pi}\left(1-\sqrt{1-\frac{E^2-H^2}{b^2}}\right). \]
For weak—
* ) In discussing the transfer of interaction by any particles, one may ask the converse question: what interactions and between which particles is the transfer by the given particles capable of explaining? Or, if one accepts the hypothesis of field mass: the masses of which particles can be conditioned by the given particles? The question of the transfer of interactions by pairs of nucleons then remains nontrivial. Indeed, from the standpoint of the field hypothesis, which does not exclude the existence of certain “supra-particles” with masses of the order of several nucleon masses, interacting with one another and transferred by pairs of nucleons, this purely hypothetical possibility should not be forgotten; however, the effective radius of action of the forces transferred by nucleons will be
\[ r\sim \frac{h}{Mc}\sim 10^{-15}\ \text{cm}, \]
i.e. smaller than the value \(r_0\sim 10^{-13}\ \text{cm}\). If, however, all elementary particles prove to possess similar “sizes” \((\sim r_0)\), then the transfer of interactions by nucleons and the existence of supra-particles must be excluded.
in fields considerably smaller than the maximum, \(E<b,\ H<b\), expanding in a series, we obtain: \(L \simeq \dfrac{E^2-H^2}{8\pi}=I_1\), i.e. the Lagrangian of the ordinary Maxwell theory. Thus, in Born’s theory a certain correspondence principle is fulfilled. In the general case, the equations of the electromagnetic field will obviously be nonlinear. For example, instead of \(\operatorname{div}\mathbf E=0\), even in vacuum we have the equation: \(\operatorname{div}\mathbf D=0\), where \(\mathbf D=\varepsilon\mathbf E\), \(\varepsilon=1/\sqrt{1-E^2/b^2}\).
In a very curious way, the total energy of the electrostatic field produced by a point charge turns out to be finite:
\[ W=\frac{e^2}{r_0}\left(b=\frac{e}{r_0^2}\right), \]
and it can be equated to the proper energy of the electron, \(mc^2\).
As is not difficult to verify, Born’s theory also leads to the correct relation between momentum and energy. One may say that instead of the energy expression \(\dfrac{e^2}{r}\), Born takes \(\dfrac{e^2}{\varepsilon r}\), where \(\varepsilon\) is a dielectric constant different from 1 even for the vacuum in the nonlinear theory. The decrease of \(r\) is compensated by the increase of \(\varepsilon\), so that the total energy of the field turns out to be finite.
Subsequently a more general Lagrangian was considered:
\[ L=\frac{b^2}{4\pi}\left(1-\sqrt{1-\frac{E^2-r^2}{b^2}-\frac{(EH)^2}{b^4}}\right), \]
as well as many other combinations of the invariants \(I_1\) and \(I_2\), leading to a finite longitudinal energy.
With this, in essence, end all the successes of Born’s theory, which in a certain sense carries out the program of constructing a purely field-theoretic “unitary” theory of charge (even if only of a single electron). There is no point in discussing comparison with experiment, since the whole foundation of the theory rests on an arbitrary choice of the Lagrangian. At the same time, on the basis of Born’s theory it is still often convenient to demonstrate characteristic nonlinear effects.
The problem of nonlinear electrodynamics first acquired real significance after the discovery of the creation and annihilation of particles and Dirac’s theory of the positron. Indeed, according to relativistic quantum mechanics, two photons can, upon colliding with one another, produce an electron–positron pair \((e_-,e_+)\), which, in turn, can annihilate and give two or more photons. If the intermediate process proceeds virtually, then we shall be dealing with the transformation of two photons into two or more photons; moreover, the new photons may possess a different polarization and, generally speaking, other frequencies. It is clear that this whole phenomenon of the scattering of light by light has the nonlinear character of a “collision” of two photons or two electromagnetic waves, violating the superposition principle that holds in linear Maxwell electrodynamics, according to which
waves pass through one another without interacting. Symbolically, the scattering of light by light is written in the form \(\gamma+\gamma' \to (e_-+e_+) \to \gamma''+\gamma'''\).
Thus, the possibility of the creation and annihilation of particles necessarily leads to nonlinear effects. Moreover, relativistic quantum mechanics predicts the possibility of any nonlinear processes: polarization of the vacuum, scattering of light by light, reflection of light from light, etc. The probability of scattering of light by light is very small and, apparently, at the existing radiation densities this effect cannot be observed, just as other nonlinear effects lying at the limit of observational possibility cannot be observed.
One may ask the question: what change in Maxwell’s theory in the nonlinear sense will be equivalent to such an inclusion of effects associated with particle pairs? Symbolically, the theory of Dirac pairs and the vacuum plus the linear Maxwell theory will be equivalent to a certain nonlinear electrodynamics characterized by a special Lagrangian. Since the sought Lagrangian \(L\), when higher derivatives are neglected, is a function of the invariants \(I_1\) and \(I_2\), then for not too high field values it can be expanded in a series in powers of \(I_1\) and \(I_2\), and written as:
\[ L_{NL}=I_1+\alpha I_2+\beta I_2^{2}+\cdots . \]
The Dirac pair theory leads to definite values of the coefficients:
\[ \alpha=7\beta;\qquad \beta=\frac{1}{360\pi^{2}}\cdot \frac{1}{(2\pi)^{5}}\cdot \frac{e^{4}h^{5}}{m^{4}c^{7}} . \]
It is curious that various variants of the Born–Infeld theory led to values of these coefficients close to these.
The difficult problem of constructing nonlinear electrodynamics by means of the Dirac pair theory was considered by Weisskopf and Heisenberg’s group only under certain restrictions (weak fields and weak field gradients), and moreover under the condition of applying a definite “subtractive” prescription required to eliminate a number of infinities associated with the theory of Dirac pairs (or the Dirac “vacuum”). It is important to emphasize that the nonlinear electrodynamics following from quantum mechanics proved, by itself, incapable of eliminating the difficulty with the infinite longitudinal energy of the field or the proper mass*.
* According to Weisskopf’s calculations, this infinity acquires, however, a logarithmic, i.e. weaker, character, instead of a divergence of the type
\[ \frac{1}{r}\to\infty \]
\[ \lim r\to 0. \]
Here one should make an essential remark: nonlinearity enters electrodynamics not only owing to the virtual creation of pairs of elec-
Similarly in mesodynamics, where the fundamental possibility of the virtual creation of a pair of nucleons by two mesotrons and their subsequent virtual annihilation with the emission of two new mesotrons gives the typical nonlinear effect of the scattering of mesotrons by mesotrons—violating the superposition of mesotron $\psi$-functions.
Symbolically: $\mu+\mu' \to (N+N') \to \mu''+\mu'''$.
Nonlinear mesodynamics, whose construction has so far been discussed only in the form of a program, must lead to such effects as polarization of the nucleon vacuum, reflection of mesotron waves from one another, nonlinear scattering of mesotrons by mesotrons, and must apparently also bring about a considerable weakening of the interaction of nucleons at small distances, helping to eliminate the dipole difficulty in the theory of nuclear forces.
Nonlinear theories, in principle, make it possible to derive the equations of motion of particles from the equations of the fields generated by these particles, just as was done in the theory of gravitation.
Since all particles, in one way or another, can in collisions virtually create pairs of various other particles, which after virtual annihilation can again turn into pairs of particles of the original type, it follows from this that one must conclude that all quantum kinematics is universally nonlinear, including the equations of the fields of electrons, positrons, and nucleons. It would be premature to discuss these new complex problems here. We shall only point out that for nonlinear field equations the expansion into separate components of the Fourier-series type loses its meaning, and thereby secondary quantization—which consists in assigning particle numbers to Fourier amplitudes—loses its directly intuitive meaning. Thus, apparently, the very concept of particles acquires a different meaning in nonlinear theory.
b) Higher derivatives. Recently the question has been discussed of a possible generalization of the field equations of elementary particles by introducing higher derivatives. The simplest example of such a theory is the scheme of Bopp and Podolsky. By introducing higher derivatives of the fields into Maxwell’s equations, we obtain a formally irreproachable invariant theory, which passes over, for slowly varying fields, into ordinary electrodynamics. In particular, in the static case, instead of Laplace’s equation, we obtain an equation of the 4th order:
\[ \Delta(\Delta-\varkappa_0^2)\varphi = 4\pi e_1 \rho, \]
which, for a point charge producing the field, has the solution:
\[ \varphi = \frac{e_1}{r}\left(1-e^{-\varkappa_0 r}\right). \]
The interaction energy $V=e_2\varphi$ of two charges, due to the new
trons and positrons, but also through pairs of charged mesotrons. Investigations of the nonlinearity of this type have hardly yet begun. (See Ivanchko-Sokolov-Razman.) All possible pairs of charged particles must contribute their share to the final nonlinear Lagrangian.
field, will be a combination of potentials of the Coulomb and Yukawa type. The self-energy of a point particle turns out to be finite, and it may be equated to the proper energy, i.e., one may put
\[ \varepsilon_0 = mc^2 = V_{(r=0)} = e^2 \varkappa_0,\quad \text{hence,}\quad \varkappa_0 = \frac{1}{r_0}. \]
The difficulty arising from the quasi-magnetic term of the form \(r^{-3}\) in the expression for the interaction energy of two nucleons under the transfer of forces by vector or pseudoscalar mesotrons is completely removed, as was shown by us (Ivanenko and Sokolov). It should be noted that, for the entire field theory with higher derivatives (which, incidentally, cannot be said to be among the widely discussed hypotheses), the most essential point at present is the substantiation of the very necessity of introducing higher derivatives, and not the consideration of particular concrete variants. This substantiation must be carried out with the same persuasiveness with which, for example, the general necessity of nonlinear generalization is proved. Recently, the first steps in this direction have been taken by Kramer, Belinfante, and Lubanski, who showed that equations with higher derivatives are obtained naturally within the framework of general undor (spin-tensor) equations.
Thus, in view of the general arbitrariness of the theory, which is moreover restricted to the nearest degrees of higher derivatives, and in view of a number of difficulties with negative energy, this attempt also is not ultimately convincing and should be regarded rather only as a very preliminary indication of new possibilities.
c) Damping. As the next point we shall consider the very important, now widely used theory of damping, or of the back-reaction of the field on particles. The allowance for the back-reaction of the field on the particles that have emitted it is, obviously, a part of the problem of self-action, and therefore the complete solution of this problem is very difficult. In itself, however, the allowance for damping contains nothing hypothetical; the only surprise was the significant role of damping for fields possessing an effective dipole character (for example, vector mesotrons), proved recently (Bhabha, Ivanenko and Sokolov, Heitler, Wilson). On the other hand, as is known, for electromagnetic effects the role of damping is very small. For effectively dipole fields, similar to the vector or pseudoscalar mesotron field, in problems of scattering and particle production the allowance for damping is very substantial and makes it possible to remove the difficulties connected with the unlimited growth of the effective cross sections or probabilities of these effects at high energies. (For example, the probabilities of scattering of light by a vector mesotron, of scattering or production of vector mesotrons by nucleons, etc.) Unfortunately, up to now it has not been possible to take into account an effect analogous to damping in the static problem of the interaction of two nucleons.
As an example of taking damping into account, let us consider the scattering of mesotrons by nucleons. The simplest course is to start from the classical mesody-
notes. Then the nonrelativistic equation of motion of a nucleon of mass \(M\) and quasi-electric charge \(g\) in the field of a plane mesotron wave will have the form:
\[ M\ddot{x}=gF. \]
Taking into account the values of the energy of the mesotron field radiated by the nucleons in 1 sec. in all directions,
\[ J=\frac{2}{3vc^2}\dot{p}^{\,2}\left(1+\frac{1}{2}\frac{\varkappa_0^2}{K^2}\right), \]
where \(v\) is the phase velocity, and dividing the radiated energy \(J\) by the mesotron energy \(I\) incident in 1 sec. through \(\mathrm{cm}^2\), we obtain the required effective scattering cross section
\[ \sigma_0=\frac{J}{I}; \]
here \(I\) is equal to the energy density
\[ u=\frac{1}{8\pi}\left[F^2+G^2+\varkappa_0^2\left(\Phi_0^2+\Phi^2\right)\right], \]
multiplied by the group velocity of the mesotron waves
\[ I=uv_g;\qquad v_g=c^2/v. \]
Finally we have:
\[ \sigma_0=\frac{8\pi r_n^2}{3}\left(1+\frac{\varkappa_0^2}{2K^2}\right), \]
where \(r_n\) is the quasi-electric radius of the nucleon,
\[ r_n=\frac{g^2}{Mc^2}. \]
For \(\varkappa_0=0\) and replacing \(r_n\) by the electromagnetic radius, we obtain Thomson’s well-known formula for the scattering of light by a charge. Taking into account the force of radiative mesotron friction
\[ F_e=\frac{2}{3}\frac{g^2}{c^3}\dddot{x} \]
does not introduce any substantial change, and the effective cross section assumes the form:
\[ \sigma_e=\frac{\sigma_0}{1+\left(\dfrac{2}{3}\dfrac{r_n\nu}{c}\right)^2} \]
(where \(\nu\) is the frequency of the mesotron wave).
A completely different state of affairs is obtained for the scattering of transverse mesotrons by a quasi-magnetic nucleon dipole (and not a charge), when the effective dipole character of the mesotron vector field is substantially manifested. Taking as a basis the equation of motion of the dipole
\[ \dot{m}=\eta[mG], \]
where \(\eta\) is the ratio of the quasi-magnetic moment of the nucleon to the mechanical one, and repeating the preceding arguments, we obtain, in the case of high energies most interesting to us, the following effective cross section for the scattering of mesotrons by nucleons
\[ \sigma_m=\frac{32\pi}{3}\left(\frac{f^2}{\hbar c}\right)^2\frac{\nu^2}{c^2\varkappa_0^2} =\mathrm{const.}\ \nu^2. \]
Here the magnitude of the radiation is equal to \(J'=\dfrac{2}{3}\dfrac{\ddot m^{\,2}}{c^3}\), and \(f=g/x_0\) is the quasi-magnetic moment of the nucleon. Thus, the effective cross section increases without bound with the frequency of the mesotron field, which is physically meaningless and formally absurd. Now damping must be taken into account. As is not hard to show (Bhabha, Ivanenko), the force of the radiation quasi-magnetic friction is equal to
\[ F_m=\frac{2}{3}\eta \frac{\ddot m}{v^3}. \]
At high energies, taking radiation friction into account, we obtain an effective cross section which does not increase, but even decreases with frequency,
\[ \sigma'_m=\frac{\sigma_m}{1+\dfrac{4}{9}(f^2/hc)^2\cdot \lambda^4/c^4 x_0^4}. \]
A quantum calculation for neutral mesotrons leads exactly to the same expression (Sokolov). Thus, allowing for damping removes the difficulties connected with the growth of effective cross sections.
In the further development of the theory of damping, Heitler and Peng indicated a preliminary, invariant method of discarding the infinite self-action terms and separating the pure damping effect. However, this theory still lacks final demonstrative force and is not capable, for example, of eliminating the difficulty with the emission of low-frequency photons pointed out by Bloch and Nordsieck.
d) Excited states. Heitler–Ma and Bhabha were the first to point to the hypothesis of higher excited states of nucleons of somewhat different higher mass, as a way of eliminating the difficulties connected with the unlimited growth of effective cross sections. Subsequently Wentzel and Pauli showed that excited states of nucleons are obtained of necessity in the so-called theory of “strong coupling,” in which the interaction of nucleons with the mesotron field is computed not by the usual perturbation theory by expansion in the constant
\[ \beta=\frac{2\pi g^2}{hc}, \]
but, on the contrary, the considerable magnitude of \(\beta\) in comparison with Sommerfeld’s fine-structure constant
\[ \alpha=\frac{2\pi e^2}{hc}, \]
is essentially taken into account, and the interaction is assumed from the very beginning to be large. Later, however, it became clear that the theory of “strong coupling,” which from the very beginning suffered from the defect of a nonrelativistic description of point nucleons, leads to incorrect consequences in its application to nuclear problems, and recently this theory was abandoned by its authors themselves. In addition, Heitler also inclined in favor of damping as a physically justified means of eliminating the inadmissible growth of effective cross sections. However, the hypothesis itself of excited states of nucleons and, obviously, of other elementary particles is physically rather interesting independently of its insufficient justification, and recently it has again and again been advanced in various preliminary theo-
...series (see, for example, Baba’s recent theory and the works of Tamm–Ginzburg). This concerns higher spin states of nucleons, hypothetical protons and neutrons of spin $^{3}/_{2}$, $^{5}/_{2}$, etc., and also, we add, mesotrons of spin 1, 2, etc. Further, in another version of the hypothesis, nucleons (and mesotrons, we add) with charges $+2e$, $+3e$, $-e$, $-2e$, etc. are allowed, i.e., in particular, antiprotons, antiprotons with double negative charge, etc. In all this, a somewhat greater mass is assigned to these nucleon “isobars,” with an excitation energy lying, according to various authors, within the range from several million electron-volts up to $10^8$ electron-volts and even up to values equal to the nucleons’ own energy. It is needless to emphasize that the final argument in favor of the hypothesis of excited states can only be the experimental discovery of such particles.
e) The limiting process. Alongside the theories considered above (of field mass, or theories with higher derivatives), which sought to obtain a finite value of the proper mass, one should point out other tendencies in contemporary physics, characterized by a temporary abandonment of the understanding of mass and directed only toward the elimination of the difficulties connected with the infinite proper energy of the field. The most interesting is the Wentzel–Dirac theory, which artificially introduces a certain vector $\lambda$, playing, figuratively speaking, the role of effective “dimensions,” or rather “extent,” of the particle in time.
Instead of calculating the force acting on an electron, for example, at 12 h. 00 min. 00 sec., Wentzel calculates the force as half the sum of the values at an instant preceding 12 h. 00 min. 00 sec. by some small amount $\lambda_0$, and at an instant following 12 h. 00 min. 00 sec. by $\lambda_0$. In the final result $\lambda_0$ tends to zero. It then turns out that the electrostatic energy corresponding to the proper mass of the field produced by a point charge, and at the same time the field proper mass, vanishes exactly, both in the classical and in the quantum theory; the energy of the mesotronic longitudinal field produced by a nucleon at rest, although it does not vanish, proves to be finite and, in order of magnitude, equal to the proper energy of the mesotron. Thus the hypothesis of the limiting “$\lambda$-process” is fundamentally at variance with the hypothesis of the field mass of particles, since the latter quantity is here simply, in fact, eliminated. The question of the nature of mass is not posed, so that the method of the $\lambda$-process plays, so to speak, not a creative but a “surgical” role.
For the elimination of the infinities caused by the transverse part of the field, the $\lambda$-process proves insufficient; therefore Dirac proposed admitting the existence of negative energies and negative probabilities, which, as can be shown, eliminates the infinite transverse mass. We shall not dwell on the formalism
of Dirac’s theory, since his hypotheses have so far led to no concrete results. Moreover, as Pauli emphasized, the cumbersome and not very transparent method of negative probabilities in any case does not remove all the infinities encountered in relativistic quantum mechanics.
It is important to point out that both the field hypothesis and the $\lambda$-limiting process attempt to remove the difficulty of the infinite energy of the field of a point charge by ascribing to the particle one or another “form” (i.e., charge distribution). Along with the $\lambda$-limiting process, a very large number of other formally irreproachable ways of introducing relativistically invariant “form factors” were proposed, cutting off the unbounded growth of the field at the smallest distances, whereby the energy becomes finite. The quantum commutation rules both for the amplitudes of wave functions and for the $\psi$ themselves must then also be modified. However, all such proposals by Marx, Wataghin, Landé, Sherser, Markov had no success and, as was recently shown by the last-named author, in essence lead to additional difficulties.
f) Characteristic matrix. A new fundamental program, also widely discussed at the present time, was advanced in 1942–1944 by Heisenberg. The latter proposes placing at the head of the whole theory not the Hamiltonian function (energy) and the wave equations of relativistic quantum mechanics, which in one way or another, along with successes, led to the indicated difficulties, or in any case failed to remove the difficulties with the infinite energy of the field produced by a point charge, and others, but rather a certain characteristic “scattering matrix” $S$, which is to transform directly the observable incident waves into scattered waves, observable at large distances. A more exact description of collision processes at the smallest distances, as unobservable, should have no place in the theory. As was shown by Heisenberg, Pauli, Møller, Stueckelberg, and Blokhintsev, in simple cases the matrix $S$ indeed replaces the Hamiltonian and makes it possible to solve not only the scattering problem, but also to determine stationary energy states. For a weak interaction, expanding the unitary matrix $S=e^{i\eta}$ in a series, we obtain $S \simeq 1+i\eta$, where $\eta$, in essence, coincides with the interaction energy. However, for an unambiguous construction of $S$ in the general case the requirements of invariance and unitarity alone are insufficient, so that Heisenberg’s entire theory, proceeding under the broad slogans of renouncing the space-time description of phenomena at small distances and expelling unobservable quantities, so far remains a fairly arbitrary scheme. What is interesting and promising is the presence of certain common features in the theory of damping and in the hypothesis of the characteristic matrix.
g) “Quantum” geometry. Finally, let us note the fairly widespread conviction, expressed in many theories, that it is necessary to introduce into the theory some universal length, probably of the order of the electron radius. It is therefore not excluded that one will have to pass to some new “quantum” geometry, taking into account the known “discontinuity” of space-time. Ordinary geometry may then prove to be only a rough averaging of this quantum metric. Preliminary attempts in the indicated direction have been undertaken from time to time, but so far have not led to completed results (Ambartsumian and Ivanenko, Heisenberg, March). An essential step in the same direction was recently taken by Snyder, who showed that the requirements of Lorentz invariance can be satisfied by introducing new coordinate operators with integer-valued eigenvalues. Snyder’s result is equivalent to the introduction of a curvilinear metric in momentum space. In this connection it can be shown that the difficulties with the infinite energy of the field do indeed, apparently, disappear. Independently of the further development of the theory, these results compel attention.
In summary, it may be said that, in the struggle with the difficulties of the theory of nuclear forces and of infinite self-mass, the difficulties of the vacuum, and also the divergences of higher approximations, new paths are emerging toward the construction of a more general theory, which must constitute a fully consistent, non-contradictory relativistic quantum mechanics of elementary particles and fields and will be able not only to remove all the indicated difficulties, but also to derive the values of the masses and charges of all particles*).
) In keeping with the purpose of the present review article, which aims only to provide, as accessibly as possible, an introduction to the theory of elementary particles, it does not seem advisable to give precise references to the literature. We therefore refer readers to the reviews in Uspekhi fizicheskikh nauk and Reviews of Modern Physics, as well as to the books soon to appear in Russian translation: Wentzel’s Quantum Theory of Fields and Pauli’s Relativistic Theory of Elementary Particles and his Meson Theory of Nuclear Forces*.