Full Text
THE LAW OF CONSERVATION OF MOTION AND THE MEASURE OF MOTION IN PHYSICS
V. S. Sorokin
INTRODUCTION
The discussion that has taken place in recent years concerning the meaning of Einstein’s law (also called the law of equivalence of mass and energy, or the law of the interrelation of mass and energy), despite the fact that a large number of physicists and philosophers have taken part in it, nevertheless leaves an impression of dissatisfaction. Einstein’s law connects two concepts that, at first glance, are entirely different—the concepts of energy and mass, of which only the concept of energy seems sufficiently clear. The concept of mass, when Einstein’s law is discussed, is usually not analyzed—perhaps because it is regarded as familiar. At best it is mentioned in passing that “mass is a measure of inertia.” In reality, however, the concept of mass is not so simple.
To say that mass is a measure of inertia is to say nothing, since it is quite unclear what inertia is, and still less clear how its measure is determined.
In this lack of clarity in the concept of mass lies, in our view, the reason why the discussion of Einstein’s law has not led to clarification of its meaning. In fact, mass, by its very concept, is inseparably connected with energy, and the meaning of both these quantities cannot be understood without examining the law of conservation of motion. This law is not equivalent to the law of conservation of energy, which is only one of its aspects.
Einstein’s law, on the other hand, is only one of the consequences of the law of conservation of motion, and any discussion of it must proceed from this law and from the concept of the measure of motion, which are by no means characteristic only of relativistic physics but were established already in classical theory.
That is why we consider it useful to set forth the law of conservation of motion and the doctrine, inseparably connected with this law, of the measure of motion as they are understood in physics. It seems to us that
the facts set forth in this paper, firmly established by modern theory, must serve as the starting point for any philosophizing about mass, energy, and the law of conservation of motion.
In order to make the exposition more accessible and to reveal more clearly the roots of the law of equivalence, the entire theory is first set forth on the basis of the classical conceptions of space and time, and then all the reasoning is repeated with allowance for the existence of a critical velocity. A reader not familiar with the mathematical side of the theory of relativity may merely follow the course of thought in the second part (§§ 3–4), omitting all calculations.
§ 1. THE MEASURE OF MECHANICAL MOTION IN CLASSICAL PHYSICS
Different forms of motion of matter are capable of transforming into one another without restriction. Such transformations may occur either within the limits of a single physical system—for example, when mechanical motion is transformed into heat—or motion in one system may excite motion in others. However, in all transformations motion is neither annihilated nor created.
Usually this assertion is regarded as almost obvious and completely clear in content. In fact, however, without the concept of a measure of motion it would be devoid of any meaning, since one can speak of the conservation of motion in its transformations only when one can speak of the “quantity” of motion of each kind. The concept of measure is inseparably connected with the law of conservation of motion, and if there were no measure of motion, one could not speak of any indestructibility of it.
Thus, we must assume and prove experimentally that every physical system in which motions of various kinds occur possesses a measure of all the motion taking place in it. If this system interacts with other systems and transmits motion to them (or receives motion from them), then its measure of motion changes in the same way as do the measures of motion of the systems with which it is connected; however, the total measure of motion of all interacting systems must remain unchanged. This assertion is usually called the principle of conservation of energy. Until now it has proved true in all cases, and it may be regarded as one of the fundamental principles of modern science.
The term “energy” in the name of this principle means the same as the term “motion.” But this same word “energy” is used in physics in another sense—namely, in the sense of “measure of motion.” Such usage is quite understandable and once again confirms that the conservation of motion is inseparably connected with the concept of its
measure. But one must not confuse these two entirely different concepts, and below we shall speak of the law of conservation and transformation of motion and of the measure of motion. The term “energy,” however, is inappropriate to introduce before clarifying the properties of the measure of motion.
In the study of any motion, the first question that arises is that of the numerical value of its measure. But qualitatively different kinds of motion are not directly comparable quantitatively. Only by transforming them into motion of one and the same kind can one compare their measures and, consequently, determine their numerical values. Therefore, in order for the concept of a measure of motion to be introducible at all, it is necessary that every motion be capable, at least in principle, of being transformed into motion of one definite kind; the assertion that any kind of motion is capable of being transformed into any other kind of motion is part of the content of the law of conservation of motion.
Some definite kind of motion must be adopted as the basic one and must play the role of the universal equivalent of motion, embodying the concept of measure. Which particular kind of motion should be chosen is, in principle, immaterial, but in modern physics mechanical motion is usually regarded as basic. Motion is, above all, change, and mechanical motion is the change in the position of macroscopic particles of matter relative to one another. The choice of this kind of motion as the universal equivalent is not accidental. For the study of any physical systems and any kinds of motion, one has to construct “frames of reference,” whose reference points are particles of matter. Any system (or any phenomenon) can change its position with respect to the reference particles. In other words, space and time are studied in modern physics in close connection with mechanical motion. The problem of the measure of this kind of motion acquires fundamental significance.
By its very concept, mechanical motion is relative in the sense that one particle of matter can move only relative to others, and it is meaningless to speak of the mechanical motion of a particle in the absence of other particles. Of course, the actual motion of particles is absolute in another sense—both the particle itself and its motion exist independently of whether we study them or not. But mechanical motion is a change in the mutual position of particles, and therefore a single particle considered in isolation does not possess mechanical motion, since there is no “mutual position” that could change. Only in this sense will the term “relative” be used below.
Since the state of a particle, insofar as it determines its mechanical motion, must be determined relative to other particles, it is first necessary to choose a frame of reference,
i.e., reference particles, with respect to which the motion will be studied. Reference particles may, of course, be chosen arbitrarily; but since, generally speaking, they will transmit motion to one another, the motion of the particle under study will be intertwined with the motion of other particles, and it will be very difficult to discern its laws. Therefore we shall require, first of all, that the reference particles be connected neither with one another, nor with the particle under investigation, nor with any external systems. In other words, we shall require that frames of reference be constructed on free particles. For free particles a law has long since been established according to which such particles move with constant velocities relative to one another (the law of inertia).
This law makes it possible to impose the second condition: the reference particles must be at rest relative to one another. Frames of reference satisfying these two conditions are called inertial.
In more detail, the study of the motion of a particle relative to an inertial frame of reference could be described as follows: the basis of the frame of reference will be any four free particles, at rest relative to one another and not lying in one plane. One of them may be taken as the origin of a rectilinear coordinate system, and the axes of this system may be drawn through the three remaining particles. The position of the particle under study can now be specified by its coordinates in this rectilinear system.
If we were to construct a second inertial frame of reference, then, by the law of inertia, each of its reference particles would move with one and the same constant velocity with respect to the reference particles of the first inertial frame of reference. This velocity may be regarded as the velocity of the second frame of reference relative to the first. Thus, any inertial frame of reference moves relative to any other with constant velocity. Conversely, we obtain all inertial frames of reference if, taking any such frame, we construct for every conceivable velocity an inertial frame of reference moving relative to the first with this velocity. Consequently, there exists an innumerable multitude of inertial frames of reference, and from their definition it is clear that none of them is in any way distinguished among the others. It is important to emphasize that this is not a purely logical proposition. At its foundation lies the experimental fact—the law of inertia. But from the law of inertia and the definition of inertial frames of reference there already follows their equal status.
Every phenomenon may be considered in connection with some inertial frame of reference which, of course, will have no substantial relation to it (for the reference particles of the frame of reference are assumed to be free). It is obvious that exactly the same phenomenon may occur in exactly the same
same way with respect to any other inertial frame of reference. Everything that can happen in one inertial frame of reference can, in exactly the same way, happen in any other. Logically, this proposition seems an undoubted consequence of the law of inertia, and in modern physics it is also regarded as a firmly established experimental fact, with the proviso that the phenomena in question must take place in not very large regions of space and over not very long intervals of time. This proposition evidently expresses a certain property of space and time, reminiscent of isotropy. All phenomena occurring in space and time must, so to speak, adapt themselves to these forms of existence of matter. As is well known, Galileo’s discovery of this principle for mechanical phenomena marked the beginning of modern mechanics.
This principle is called the principle of relativity. From it it follows, in particular, that every physical law must be formulated in such a way that its formulation contains no reference to any individual inertial frame of reference. This circumstance gives the principle of relativity exceptional heuristic power. The requirement that the formulation of the laws of nature be the same with respect to all inertial frames of reference often proves so strong that it makes it possible to derive exact functional relations between various physical quantities on the basis of the simplest qualitative regularities. Such an application of the principle of relativity is often considered characteristic of so-called relativistic physics. It is, however, perfectly clear that this principle plays exactly the same role in both classical and relativistic theory. It was not applied in classical physics only because all the essential results in it had been obtained before its significance was understood. As we shall see, for this reason certain essential regularities remained undiscovered in the old physics.
Let us now consider a particle of matter in some inertial frame of reference. The state of its motion at any moment is determined by the distribution of velocities existing in the particle at that moment, since different points of the particle may move with different velocities. However, one can always mentally divide the particle into smaller ones so that within each of the smaller particles the velocity is everywhere almost the same. Then one may speak simply of the velocity of the particle, and in what follows it will be assumed that the particle under consideration is sufficiently small in this sense. Let its velocity in the chosen inertial frame of reference (frame of reference 1) be $\mathbf{v}$. This velocity completely characterizes the state of motion of the particle, for the position of the particle is immaterial:
at the same velocity in another place the motion will be exactly the same, while the internal motions are regarded either as absent or as having no significance. Consequently, the measure of the mechanical motion of a particle in a given frame of reference must in a definite way depend on its velocity and be completely determined by it.
Most readers will probably at once imagine this measure as a certain number giving the “quantity of motion” of the particle and, confusing indestructibility with absoluteness, will be inclined to demand of this number invariance, i.e., will be inclined to think that this number must be one and the same in all inertial frames of reference. But this is obviously impossible: we have just seen that the measure must be determined by the velocity, while the velocity of one and the same particle at one and the same instant is entirely different in different frames of reference. An invariant measure would have to have one and the same value at all velocities, i.e., it could not measure motion.
Thus, the measure of the mechanical motion of a particle must be relative, like mechanical motion itself, and one should speak of the measure of motion relative to a given inertial frame of reference, or of the measure in a given frame of reference. In another inertial frame of reference the measure of the mechanical motion of the very same particle at the same instant of time will be different, and this does not contradict the indestructibility of motion.
What, then, is the measure of motion at a given velocity of a particle in a definite frame of reference? It is clear that there must exist a function, characteristic of the given particle, of its velocity,
\[ \varepsilon(v), \tag{1.1} \]
which determines the measure of motion. The form of this function, i.e., the character of the dependence of the measure of motion on the velocity, is of course not known in advance and may be different for different particles. But whatever this dependence may be, it expresses a certain dynamical law—the law determining the “quantity of motion” of a particle at a given velocity—and, like any law, this dependence must not contain any indication of any individual inertial frame of reference. If a particle moves with velocity \(v\) relative to frame of reference I, it must possess in this frame of reference exactly the same measure of motion as would be possessed relative to another inertial frame of reference II by exactly the same particle moving in this frame II with the same velocity \(v\). The form of the function \(\varepsilon(v)\) must be, for the given particle, one and the same in all inertial frames of reference.
Let us express this requirement mathematically. Let inertial frame of reference II move relative to the same frame of reference I ...
account I with velocity $\mathbf{u}$. Suppose that in reference system I the velocity of the particle was $\mathbf{v}$ and, consequently, the measure of motion was $\varepsilon(\mathbf{v})$. In reference system II the velocity of this same particle will no longer be $\mathbf{v}$, but, by the rule for addition of velocities,
\[ (\mathbf{v}-\mathbf{u})^{*}). \tag{1.2} \]
The measure of motion in reference system II is computed from this velocity in exactly the same way as in reference system I, i.e. it will be
\[ \varepsilon(\mathbf{v}-\mathbf{u}) \tag{1.3} \]
with the very same function $\varepsilon$ as before.
Everything that follows is already contained in these formulas and in the concept of a measure. It is the task of mathematics to clarify what follows from the assumption of invariance of the connection between velocity and the measure of motion, or, more precisely, what we assert in assuming this invariance.
The concept of a measure contains, first, the property of its conservation: if motion is not transferred to other systems, but is only transformed within the same system, then its measure must not change. Secondly, the concept of a measure contains its additivity: the measure of motion of a system consisting of parts not connected with one another is equal to the sum of the measures of these parts. The functions $\varepsilon(\mathbf{v})$ (for different particles they may be different) must be such that both of these properties are satisfied in systems consisting of particles.
Let us consider the case where two particles, initially not connected with one another, enter into interaction, transfer motion to one another, and again separate, so that every connection between them ceases. Suppose, moreover, that the motion remains mechanical and is not transformed into any other forms. We first examine this process in reference system I. Let the velocities of the particles in this reference system before the encounter be $\mathbf{v}_1$ and $\mathbf{v}_2$, and after the encounter $\mathbf{v}'_1$ and $\mathbf{v}'_2$.
The measures of motion of both particles are connected with their velocities by certain dependences, generally speaking different for the two particles, but the same for one and the same particle before and after the encounter (since their structure does not change). Since before and after the encounter the particles are not connected, by the property of the measure the measure of the whole system is equal to the sum of the measures of both particles. Conservation of motion then requires that the condition
\[ \varepsilon_1(\mathbf{v}_1)+\varepsilon_2(\mathbf{v}_2) = \varepsilon_1(\mathbf{v}'_1)+\varepsilon_2(\mathbf{v}'_2). \tag{1.4} \]
be satisfied.
*) The validity of this rule, and consequently also the validity of everything that follows in this paragraph, is restricted to velocities small in comparison with the velocity of light; classical theory ignores this circumstance.
Now let us consider this very same process with respect to reference frame II (whose velocity in I, as already stated, is equal to \(\mathbf{u}\)).
In this reference frame the velocities of the particles before approach will be
\[ (\mathbf{v}_1-\mathbf{u}),\quad (\mathbf{v}_2-\mathbf{u}), \]
and after approach
\[ (\mathbf{v}'_1-\mathbf{u}),\quad (\mathbf{v}'_2-\mathbf{u}). \]
From these velocities one can compute the measures of motion of both particles in reference frame II; moreover, by the principle of relativity the dependence of the measure of motion of each particle on its velocity will be, in the new reference frame, exactly the same as in the old one, so that this dependence will be expressed for the first particle by the same function \(\varepsilon_1\), and for the second by \(\varepsilon_2\). Therefore the conservation of motion is written in the new reference frame as follows:
\[ \varepsilon_1(\mathbf{v}_1-\mathbf{u})+\varepsilon_2(\mathbf{v}_2-\mathbf{u}) = \varepsilon_1(\mathbf{v}'_1-\mathbf{u})+\varepsilon_2(\mathbf{v}'_2-\mathbf{u}). \tag{1.5} \]
The velocity \(\mathbf{u}\), however, is completely arbitrary, since for any \(\mathbf{u}\) one can construct the corresponding inertial system.
It is seen from this that the character of the functions \(\varepsilon_1\) and \(\varepsilon_2\) is quite special: when all the arguments in equation (1.4) are decreased by one and the same arbitrary quantity \(\mathbf{u}\), the equality is not violated.
From this fact there follows one consequence which may seem unexpected, although it is already contained in the concept of a relative measure invariantly connected with velocity. Let us denote explicitly the dependence of the measure of motion on the three components of velocity,
\[ \varepsilon(\mathbf{v})=\varepsilon(v_x,\ v_y,\ v_z) \tag{1.6} \]
and choose the velocity \(\mathbf{u}\) so that it is parallel to the \(x\)-axis,
\[ u_x=u,\quad u_y=u_z=0 \tag{1.7} \]
and, moreover, infinitesimally small.
To accuracy up to small quantities of second order one may write
\[ \varepsilon_1(\mathbf{v}_1-\mathbf{u}) = \varepsilon_1(v_{1x}-u,\ v_{1y},\ v_{1z}) = \]
\[ = \varepsilon_1(v_{1x},\ v_{1y},\ v_{1z}) - u\frac{\partial \varepsilon_1(v_{1x},\ v_{1y},\ v_{1z})}{\partial v_{1x}} \tag{1.8} \]
and similarly for the other terms in formula (1.5). Then the condition
conservation of the measure in reference system II (1.5) will be:
\[ \varepsilon_1(\mathbf v_1)+\varepsilon_2(\mathbf v_2)-u\left[ \frac{\partial\varepsilon_1(\mathbf v_1)}{\partial v_{1x}}+ \frac{\partial\varepsilon_2(\mathbf v_2)}{\partial v_{2x}} \right] = \]
\[ =\varepsilon_1(\mathbf v'_1)+\varepsilon_2(\mathbf v'_2)-u\left[ \frac{\partial\varepsilon_1(\mathbf v'_1)}{\partial v'_{1x}}+ \frac{\partial\varepsilon_2(\mathbf v'_2)}{\partial v'_{2x}} \right]. \tag{1.9} \]
By virtue of conservation of the measure in reference system I, i.e. equation (1.4), canceling also by \(u\) (which can be done, since \(u\) is arbitrary), we obtain from this:
\[ \frac{\partial\varepsilon_1(\mathbf v_1)}{\partial v_{1x}}+ \frac{\partial\varepsilon_2(\mathbf v_2)}{\partial v_{2x}} = \frac{\partial\varepsilon_1(\mathbf v'_1)}{\partial v'_{1x}}+ \frac{\partial\varepsilon_2(\mathbf v'_2)}{\partial v'_{2x}}. \tag{1.10} \]
Two more similar equations we obtain by choosing \(\mathbf u\) parallel to the axes \(y\) and \(z\):
\[ \frac{\partial\varepsilon_1(\mathbf v_1)}{\partial v_{1y}}+ \frac{\partial\varepsilon_2(\mathbf v_2)}{\partial v_{2y}} = \frac{\partial\varepsilon_1(\mathbf v'_1)}{\partial v'_{1y}}+ \frac{\partial\varepsilon_2(\mathbf v'_2)}{\partial v'_{2y}}, \tag{1.11} \]
\[ \frac{\partial\varepsilon_1(\mathbf v_1)}{\partial v_{1z}}+ \frac{\partial\varepsilon_2(\mathbf v_2)}{\partial v_{2z}} = \frac{\partial\varepsilon_1(\mathbf v'_1)}{\partial v'_{1z}}+ \frac{\partial\varepsilon_2(\mathbf v'_2)}{\partial v'_{2z}}. \tag{1.12} \]
All these three equations can be written more briefly if one notes that the three quantities
\[ \frac{\partial\varepsilon(\mathbf v)}{\partial v_x},\quad \frac{\partial\varepsilon(\mathbf v)}{\partial v_y},\quad \frac{\partial\varepsilon(\mathbf v)}{\partial v_z} \tag{1.13} \]
form the three components of a certain vector (this means that, under a rotation of the coordinate axes, these same three derivatives, computed in the new axes, will be related to the old derivatives by exactly the same formulas as the components of a segment along the new axes are related to its components along the old ones). We shall denote this vector by
\[ \frac{\partial\varepsilon(\mathbf v)}{\partial\mathbf v}. \]
Then the three equations (1.10—1.12) can be written in the form of a single vector equation
\[ \frac{\partial\varepsilon_1(\mathbf v_1)}{\partial\mathbf v_1}+ \frac{\partial\varepsilon_2(\mathbf v_2)}{\partial\mathbf v_2} = \frac{\partial\varepsilon_1(\mathbf v'_1)}{\partial\mathbf v'_1}+ \frac{\partial\varepsilon_2(\mathbf v'_2)}{\partial\mathbf v'_2}, \tag{1.14} \]
which has the form of a certain conservation law,
namely, the law of conservation of the vector quantity
\[ \frac{\partial \varepsilon(\mathbf v)}{\partial \mathbf v}. \]
This quantity therefore has the same right to be called a measure of (mechanical) motion as does the scalar measure \(\varepsilon(\mathbf v)\), only it will be a vector measure.
Having admitted the existence of a scalar measure, we are, so far as mechanical motion alone is concerned, compelled also to admit the existence of a vector measure. The properties of space and time are such that the measure of motion, in any case for mechanical motion, must be dual; otherwise it, so to speak, does not enter into space and time.
The scalar measure of the motion of a particle is customarily called kinetic energy
\[ \varepsilon=\varepsilon(\mathbf v), \tag{1.15} \]
and the vector measure—momentum
\[ \mathbf p=\frac{\partial \varepsilon(\mathbf v)}{\partial \mathbf v}. \tag{1.16} \]
We must now dwell on two objections that might be raised against the legitimacy of our conclusions.
First, one might say that in reality no vector measure exists, since the derivative \(\dfrac{\partial \varepsilon}{\partial \mathbf v}\) may simply be constant, and the vector conservation law is then satisfied identically—the same thing always stands on the left and on the right. This objection is easily refuted. If
\[ \frac{\partial \varepsilon}{\partial \mathbf v}=\mathbf a=\mathrm{const}, \]
then
\[ \varepsilon(\mathbf v)=\mathbf a\mathbf v+b. \]
But this would mean that in space there exists some direction (namely, the direction of the vector \(\mathbf a\)) such that, for a given rapidity of motion, the kinetic energy depends on the angle formed by the velocity with this direction. Experience, however, shows that energy depends only on speed and not on the direction of the velocity. Consequently, momentum depends on velocity and is a true measure of motion.
The second objection is as follows. Conservation of momentum, just like conservation of energy, must hold in all inertial frames of reference. Hence equation (1.14) must also remain valid when all arguments are decreased by one and the same deriv-
voluntary quantity \(\mathbf{u}\). Repeating the previous reasoning, we would arrive at new conservation laws for nine new quantities, which it is convenient to write in the form of a table:
\[ \left. \begin{array}{ccc} \dfrac{\partial^{2}\varepsilon}{\partial v_x^{2}}, & \dfrac{\partial^{2}\varepsilon}{\partial v_y \partial v_x}, & \dfrac{\partial^{2}\varepsilon}{\partial v_z \partial v_x}, \\[1.2em] \dfrac{\partial^{2}\varepsilon}{\partial v_x \partial v_y}, & \dfrac{\partial^{2}\varepsilon}{\partial v_y^{2}}, & \dfrac{\partial^{2}\varepsilon}{\partial v_z \partial v_y}, \\[1.2em] \dfrac{\partial^{2}\varepsilon}{\partial v_x \partial v_z}, & \dfrac{\partial^{2}\varepsilon}{\partial v_y \partial v_z}, & \dfrac{\partial^{2}\varepsilon}{\partial v_z^{2}}. \end{array} \right\} \tag{1.17} \]
Of these nine quantities only six are distinct, since the order of differentiation with respect to independent variables is immaterial. All these derivatives are components of a symmetric tensor of the second rank, and it would seem that, besides the scalar and vector laws, there should also exist a tensor conservation law and, consequently, also a tensor measure of motion. Then, starting from this tensor conservation law and reasoning as before, we would arrive at an infinite set of tensor measures of arbitrary rank.
The situation seems catastrophic. After all, each conservation law imposes certain conditions that must be satisfied by the motion of the particles after their approach. If there are too many of these conditions, it will in general be impossible to satisfy them, and any transfer of motion will become impossible. In our case the motion of the particles after approach is characterized by two vectors—the velocities of both particles, i.e. by only six quantities:
\[ v'_{1x}, \quad v'_{1y}, \quad v'_{1z}; \quad v'_{2x}, \quad v'_{2y}, \quad v'_{2z}, \]
and, consequently, at most six conditions can be satisfied, i.e. six conservation laws. One of these conditions is the conservation of energy. Three more are supplied by the conservation of momentum. If the tensor quantity (1.17) must also be conserved, we obtain already \(1+3+6=10\) conditions, which it will be impossible to satisfy, and particles will in general be unable to collide. However, experience shows that they transfer motion to one another very well.
This means that all tensor conservation laws, beginning with the second rank, must either be fulfilled in a trivial way, i.e. identically satisfied for arbitrary velocities, or they must be consequences of the scalar and vector laws.
This requirement makes it possible to determine completely the character of the dependence of the measure of motion on velocity.
Indeed, first, we have already seen that, owing to the isotropy of space, the function \(\varepsilon(\mathbf v)\) depends only on the magnitude, and not on the direction, of the velocity. Therefore, changing the notation, we write:
\[ \varepsilon=\varepsilon(\mathbf v^2) \tag{1.18} \]
and compute the first and second derivatives. Obviously:
\[ \left. \begin{aligned} \frac{\partial \varepsilon}{\partial v_x} &=\varepsilon'(\mathbf v^2)\cdot 2v_x,\\ \frac{\partial^2 \varepsilon}{\partial v_y \partial v_x} &=\frac{\partial}{\partial v_y}\left[\varepsilon'(\mathbf v^2)\cdot 2v_x\right] =\varepsilon''(\mathbf v^2)\cdot 4v_xv_y,\\ \frac{\partial^2 \varepsilon}{\partial v_x^2} &=\frac{\partial}{\partial v_x}\left[\varepsilon'(\mathbf v^2)\cdot 2v_x\right] =\varepsilon''(\mathbf v^2)\cdot 4v_x^2+2\varepsilon'(\mathbf v^2) \end{aligned} \right\} \tag{1.19} \]
(here the prime denotes differentiation with respect to \(\mathbf v^2\)) and similarly for the remaining derivatives.
Secondly, conservation of the tensor (1.17) must follow either automatically, or as a consequence of conservation of the scalar \(\varepsilon(\mathbf v^2)\) and the vector \(\dfrac{\partial \varepsilon}{\partial \mathbf v}\). In order that, from the additive conservation laws (1.4) and (1.14), there should result an additive conservation law for the tensor, it is necessary that the components of the tensor be expressed linearly in terms of the components of the vector and the scalar, i.e. that one have \((i,\ k=x,\ y,\ z)\)
\[ \frac{\partial^2\varepsilon}{\partial v_i\partial v_k} =a_{ik}+b_{ik}\varepsilon+\sum_l c_{ikl}\frac{\partial\varepsilon}{\partial v_l}, \tag{1.20} \]
where \(a_{ik}\), \(b_{ik}\), and \(c_{ikl}\) are constant tensors, of which the first may be different for different particles, while the others are universal, i.e. the same for all particles. If \(i\ne k\), then, as is clear from (1.19), the left-hand side of formula (1.20) will contain the product \(v_iv_k\), while the right-hand side will not contain it. Hence it follows that
\[ \varepsilon''(\mathbf v^2)=0, \]
i.e. that
\[ \varepsilon(\mathbf v^2)=\frac{m\mathbf v^2}{2}+\varepsilon_0, \]
where \(m\) and \(\varepsilon_0\) are constants, whose values may be different for different particles.
Thus, every particle of matter has a scalar measure of its mechanical motion (kinetic energy), related to the velocity by the law
\[ \varepsilon=\frac{m\mathbf v^2}{2}+\varepsilon_0 \tag{1.21} \]
and, in addition, a vector measure—the momentum, whose dependence on
velocity is determined by the formula
\[ \mathbf{p}=\frac{\partial \varepsilon\left(v^{2}\right)}{\partial \mathbf{v}}=m\mathbf{v}. \tag{1.22} \]
The constant \(\varepsilon_0\) obviously gives a scalar measure of the internal motions of an immobile particle. So long as the particle is not destroyed and does not change its structure, this energy is unchanged. Therefore in mechanics it is simply not considered. The constant \(m\), however, which enters into the expression of both measures, is called the inertial mass of the particle. It characterizes the ability of the particle, at a given velocity, to possess a definite “quantity of motion,” measured in two ways—by energy and by momentum—and is, consequently, its property. That this property is invariant (in particular, does not depend on the state of motion) follows here of itself as a consequence of the principle of relativity and the rule for the addition of velocities (1.2).
To summarize, one may say: in our space and time there can exist only a double measure of mechanical motion—a scalar-vector measure, both constituent parts of which, kinetic energy and momentum, are determined by the velocity of the particle. The character of their dependence on velocity is determined uniquely by means of the principle of relativity: momentum is simply proportional to velocity, while energy depends quadratically on velocity. Mass, i.e., the quantity connecting both measures with velocity, proves to be a property of matter independent of velocity, i.e. constant; it depends only on the nature of the particle.
§ 2. GENERAL THEORY OF THE MEASURE OF MOTION IN CLASSICAL PHYSICS
Now the main question arises before us: are the properties of the measure established for mechanical motion universal, or are they specific to mechanical motion? In particular, does every kind of motion possess a vector measure—momentum? The existence of a scalar measure (energy) for any kind of motion, so far as I know, has never been disputed. But momentum is often regarded as a specific measure of mechanical motion. Although this question is entirely clear both theoretically (Einstein’s works are already fifty years old) and experimentally (P. N. Lebedev’s experiments are also about fifty years old), the notion of momentum as a specially mechanical quantity is still very widespread.
In reality, however, it is easy to prove that any physical system in which motions of any kind can occur has, just like a simple particle, two measures of motion—energy and momentum. Momentum is just as universal a measure of motion as energy.
For the proof, let us first clarify how the measures of the mechanical motion of one and the same particle, calculated with respect to two different inertial frames of reference, are related to one another. Suppose that in system I, moving with velocity $\mathbf v$, the particle has the scalar measure
\[ \varepsilon_{\mathrm I}=\varepsilon(\mathbf v^2)=\frac{m v^2}{2}+\varepsilon_0 \tag{2.1} \]
and the vector measure
\[ \mathbf p_{\mathrm I}=\frac{\partial \varepsilon}{\partial \mathbf v}=m\mathbf v. \tag{2.2} \]
In the inertial frame of reference II, whose velocity relative to I is $\mathbf u$, the velocity of the same particle will be $(\mathbf v-\mathbf u)$. At this velocity the scalar measure will be
\[ \begin{aligned} \varepsilon_{\mathrm{II}} &=\varepsilon(\mathbf v-\mathbf u) =\frac{m}{2}(\mathbf v-\mathbf u)^2+\varepsilon_0 \\ &=\left(\frac{m v^2}{2}+\varepsilon_0\right)-m\mathbf v\mathbf u+\frac{m u^2}{2} =\varepsilon_{\mathrm I}-\mathbf u\mathbf p_{\mathrm I}+\frac{m u^2}{2}, \end{aligned} \tag{2.3} \]
and the vector one
\[ \mathbf p_{\mathrm{II}}=m(\mathbf v-\mathbf u)=\mathbf p_{\mathrm I}-m\mathbf u. \tag{2.4} \]
These formulas show that both measures should rather be regarded as two components of one composite measure (just as, for example, the components of a vector are components of one more complex quantity), since in a new frame of reference the energy, for example, is expressed not only through the energy but also through the momentum in the old frame of reference.
Having established the “transformation law” for the components of the measure of mechanical motion, let us now consider the transfer of motion from a particle $A$ to some nonmechanical (or not purely mechanical) system $\Sigma$. Suppose that these two systems are initially not connected with one another, and then, having entered into interaction, transfer motion to one another and separate again. Let us first consider this process in the inertial frame of reference I. Suppose that in it the energy of the particle changes as a result of the interaction from $\varepsilon_{\mathrm I}$ to $\varepsilon'_{\mathrm I}$, while the energy of the system $\Sigma$ changes from $\mathcal E_{\mathrm I}$ to $\mathcal E'_{\mathrm I}$. Conservation of energy requires that the equality
\[ \varepsilon_{\mathrm I}+\mathcal E_{\mathrm I} = \varepsilon'_{\mathrm I}+\mathcal E'_{\mathrm I}. \tag{2.5} \]
hold.
The energy of the nonmechanical system $\Sigma$ depends on quantities characterizing the state of its motion, among which the velocity may also not be included—for the simple reason that such a quantity may not exist for a complex nonmechanical system.
Let us now pass to an inertial frame of reference II, moving with velocity $\mathbf{u}$ in the first frame of reference. Now the energy of particle $A$ before interaction with the system $\Sigma$ will be $\varepsilon_{\mathrm{II}}$, and after the interaction $\varepsilon'_{\mathrm{II}}$. The energies of the system $\Sigma$ also need not be the same as in the first frame of reference (although we, of course, cannot know in advance whether this is so or not). Conservation of energy now gives the equality
\[ \varepsilon_{\mathrm{II}}+\mathcal{E}_{\mathrm{II}}=\varepsilon'_{\mathrm{II}}+\mathcal{E}'_{\mathrm{II}} . \tag{2.6} \]
Let us express in this formula the energies of particle $A$ before and after the interaction in terms of its energies and momenta in frame of reference I, using formulas (2.3) and (2.4), and, transferring all quantities referring to particle $A$ to one side and those referring to the system $\Sigma$ to the other, write (2.6) in the form
\[ \mathcal{E}'_{\mathrm{II}}-\mathcal{E}_{\mathrm{II}} =(\varepsilon_{\mathrm{I}}-\varepsilon'_{\mathrm{I}}) -\mathbf{u}\,(\mathbf{p}_{\mathrm{I}}-\mathbf{p}'_{\mathrm{I}}). \tag{2.7} \]
Finally, using conservation of energy in frame of reference I, we replace the change in the energy of particle $A$ by the change in the energy of the system $\Sigma$, equal to it in magnitude and opposite in sign (by equation (2.5)). This gives:
\[ \mathcal{E}'_{\mathrm{II}}-\mathcal{E}_{\mathrm{II}} =(\mathcal{E}'_{\mathrm{I}}-\mathcal{E}_{\mathrm{I}}) -\mathbf{u}\,(\mathbf{p}_{\mathrm{I}}-\mathbf{p}'_{\mathrm{I}}). \tag{2.8} \]
Considering this equation, we see that the change in the energy of the nonmechanical system $\Sigma$ will be different in different frames of reference, provided only that the momentum of the particle interacting with it changes. But the momentum of the particle must necessarily change—since if the momentum does not change, then its velocity will not change, and consequently its energy will not change either, i.e. in general there will be no transfer of motion. And we must conclude that the energy of a nonmechanical system, for one and the same state of its motion, will be different in different frames of reference.
Having chosen some one inertial frame of reference as the fundamental one, we can then find the energy of the system $\Sigma$ in any frame of reference having velocity $\mathbf{u}$ relative to the fundamental frame. We shall then obtain the energy of our nonmechanical system in a certain state of it as a function of the velocity $\mathbf{u}$,
\[ \mathcal{E}=\mathcal{E}(\mathbf{u}). \tag{2.9} \]
Let us emphasize that $\mathbf{u}$ is not the velocity of the nonmechanical system $\Sigma$ itself, but the velocity of that inertial frame of reference in which $\Sigma$ is considered, relative to the arbitrarily chosen fundamental frame of reference. The character of the dependence of the energy on this velocity must be determined by the nature of the system.
Let us now repeat verbatim the arguments that earlier led us to the law of conservation of momentum for mechanical motion. From the conservation of energy in reference frame II, taking reference frame I as the principal one and using formula (2.3), we obtain:
\[ \left[\varepsilon+\frac{m u^2}{2}-\mathbf{u}\mathbf{p}\right]+\mathcal{E}(\mathbf{u}) = \left[\varepsilon' + \frac{m u^2}{2}-\mathbf{u}\mathbf{p}'\right]+\mathcal{E}'(\mathbf{u}), \]
or
\[ [\varepsilon-\mathbf{u}\mathbf{p}+\mathcal{E}(\mathbf{u})] = [\varepsilon'-\mathbf{u}\mathbf{p}'+\mathcal{E}'(\mathbf{u})]. \tag{2.10} \]
Taking \(\mathbf{u}\) to be infinitesimal, we further obtain:
\[ \varepsilon-\mathbf{u}\mathbf{p}+\mathcal{E}(0) +\mathbf{u}\left(\frac{\partial\mathcal{E}}{\partial\mathbf{u}}\right)_{u=0} = \varepsilon'-\mathbf{u}\mathbf{p}'+\mathcal{E}'(0) +\mathbf{u}\left(\frac{\partial\mathcal{E}'}{\partial\mathbf{u}}\right)_{u=0}. \]
If here one takes into account the conservation of energy in the principal reference frame and divides by \(\mathbf{u}\) (which is arbitrary), then the equation obtained is
\[ \mathbf{p}+\left[-\frac{\partial\mathcal{E}}{\partial\mathbf{u}}\right]_{u=0} = \mathbf{p}'+\left[-\frac{\partial\mathcal{E}'}{\partial\mathbf{u}}\right]_{u=0}. \tag{2.11} \]
This equation has the form of a law of conservation of a vector measure of motion, and the momentum of the nonmechanical system \(\Sigma\) should be regarded as the vector
\[ \mathbf{P} = \left[-\frac{\partial\mathcal{E}(\mathbf{u})}{\partial\mathbf{u}}\right]_{u=0}, \tag{2.12} \]
which, for a given principal reference frame (entirely arbitrary), depends only on the state of the nonmechanical system itself. Indeed, it is precisely this vector that, together with the momentum of the particle, remains unchanged during the interaction. And the fact that it is fully determined by the state of the system (and, of course, by the reference frame in which it is calculated) is evident from its definition: for a given state of the nonmechanical system, which is something having no relation to any reference frames, the dependence of its energy on the inertial reference frame in which this energy is determined is established, i.e. the form of the function \(\mathcal{E}(\mathbf{u})\); then this function is differentiated with respect to its argument, and as a result this argument (i.e. the velocity \(\mathbf{u}\)) is set equal to zero. The last operation removes the dependence on \(\mathbf{u}\).
Thus, motion of any kind has both a scalar and a vector measure, so that the measure of motion is, by its very concept, a dual, scalar-vector measure. Momentum is in no degree specific to mechanical motion; but, of course, momentum measures motion with allowance for its directionality. Not only the same kinds of particles of matter can move relative to particles of substance...
particles, but also states of non-mechanical systems; therefore any system can move in space relative to some reference frame, and its energy will depend on its “velocity.”
The concept of a measure of motion undoubtedly contains a contradiction: on the one hand, the measure of motion must be absolute, just as motion itself is absolute; on the other hand, we have seen that it cannot but be relative. This contradiction is resolved by the fact that the measure of motion is not a number, but a more complex quantity having several components. It itself has an absolute character, while its division into components—the scalar and vector measures—is relative, since it is different in different inertial frames of reference.
If we wish to emphasize the absoluteness of motion, we must regard its measure as something whole, without dividing it into components. Such a point of view should not seem strange. We are accustomed, for example, to regard a vector as a single complex quantity, and this concept is just as abstract as the concept of a complex measure of motion. In the case of a vector this abstractness is usually not noticed, since a vector can be represented visually by a segment. But even if this could not be done, the concept of a vector as a single quantity would nevertheless exist. Namely, a vector would be defined by its components in some one coordinate system and by the law of transformation of its components, i.e., by the law determining the relation between the components of the vector in different coordinate systems.
The matter is exactly the same with the measure of motion. This complex quantity cannot be represented visually, but one can specify its components in some inertial frame of reference and indicate the law of transformation of these components, relating the components of the measure in different frames of reference to one another. It is precisely the law of transformation of the components of the measure of motion that determines the mathematical character of this complex quantity.
To find this law, let us again consider two inertial frames of reference, I and II, and let II again move in I with velocity $\mathbf{u}$. Take some physical system in some state of it. In this state the system possesses a measure of motion, which in each of our two frames of reference decomposes into scalar and vector components. Let these components be, in the first frame of reference, $\mathcal{E}_{\mathrm{I}}$ and $\mathbf{P}_{\mathrm{I}}$, and in the second, $\mathcal{E}_{\mathrm{II}}$ and $\mathbf{P}_{\mathrm{II}}$. We must find their relation to one another.
For this purpose let us return to equation (2.8), which we obtained by considering the interaction of any system with a particle of matter. With the aid of this equation we proved the existence of momentum for any physical system. Now, however, when we already know that momentum always exists and is conserved, we can
replace in equation (2.8) the change in momentum of the particle by the change of momentum of the physical system itself, taken with the opposite sign. Then we obtain an equation into which only quantities pertaining to the system itself enter:
\[ (\mathcal E'_{II}-\mathcal E_{II})=(\mathcal E'_I-\mathcal E_I)-\mathbf u(\mathbf P'_I-\mathbf P_I). \tag{2.13} \]
This equation expresses the transformation law for the change of the energy of any physical system. The initial and final states of the system may be arbitrary, since one can imagine such a collision of the system with a particle in which its state changes in an arbitrary way.
If equation (2.13) is rewritten in the form
\[ [\mathcal E'_{II}-\mathcal E'_I+\mathbf u\mathbf P'_I]=[\mathcal E_{II}-\mathcal E_I+\mathbf u\mathbf P_I], \tag{2.14} \]
then it becomes clear that the quantity
\[ [\mathcal E_{II}-\mathcal E_I+\mathbf u\mathbf P_I] \]
must be one and the same for all states of the system and can depend only on the nature of the system and on the velocity \(\mathbf u\). In other words, we obtain the following transformation law for the energy of any system:
\[ \mathcal E_{II}=\mathcal E_I-\mathbf u\mathbf P_I+F(\mathbf u), \tag{2.15} \]
where \(F(\mathbf u)\) is some function of the velocity \(\mathbf u\), depending only on the nature of the system (for a particle, as is seen from (2.3), this function was equal to \(mu^2/2\)).
Here it is appropriate to note that a comparison of the energies of one and the same system in one and the same state, but determined with respect to different frames of reference, is possible experimentally. Namely, in reference frame II the energy \(\mathcal E_{II}\) will be the energy of some state of the system, and one can create another state which, in reference frame I, will look exactly as the first one looked in II. The energy of this state in reference frame I will be equal to the energy of the original state in reference frame II. The difference of the energies of two states of one and the same system in one and the same reference frame can be determined experimentally, so that the function \(F(\mathbf u)\) can indeed be found for every physical system.
The transformation law for momentum follows from the transformation law for energy. In fact, in order to compute the momentum in the II reference frame, it is necessary, as described above, to compute the energy in a third reference frame whose velocity relative to the second
let it be \(\mathbf w\). Then, by (2.12), this energy, taken with the opposite sign, must be differentiated with respect to \(\mathbf w\) and \(\mathbf w=0\) put. The velocity of the third reference system in the first will be:
\[ (\mathbf u+\mathbf w), \]
and, consequently, the energy of the system as a function of \(\mathbf w\) will be obtained from (2.15), if there \(\mathbf u\) is replaced by \((\mathbf u+\mathbf w)\):
\[ \mathscr E(\mathbf u+\mathbf w)=\mathscr E(0)-(\mathbf u+\mathbf w)\mathbf P(0)+F(\mathbf u+\mathbf w) \tag{2.16} \]
(we have written \(\mathscr E(0)\) instead of \(\mathscr E_1\), etc.). Consequently,
\[ \mathbf P_{\mathrm{II}}=\mathbf P(\mathbf u)=\left[-\,\frac{\partial \mathscr E(\mathbf u+\mathbf w)}{\partial \mathbf w}\right]_{w=0} =\mathbf P(0)-\frac{\partial F(\mathbf u)}{\partial \mathbf u} \tag{2.17} \]
—this is the law of transformation of momentum.
The remaining undetermined function \(F(\mathbf u)\) can be found as follows. Consider three inertial reference systems I, II, and III. Let the velocity of II relative to I be \(\mathbf u\), the velocity of III relative to II be \(\mathbf v\), and the velocity of III relative to I, obviously, \(\mathbf w=(\mathbf u+\mathbf v)\). Passing from I to II, from II to III, and from I to III, we obtain for the momentum of any system, by (2.17):
\[ \begin{aligned} \mathbf P_{\mathrm{II}}&=\mathbf P_{\mathrm{I}}-\frac{\partial F(\mathbf u)}{\partial \mathbf u},\\ \mathbf P_{\mathrm{III}}&=\mathbf P_{\mathrm{II}}-\frac{\partial F(\mathbf v)}{\partial \mathbf v},\\ \mathbf P_{\mathrm{III}}&=\mathbf P_{\mathrm{I}}-\frac{\partial F(\mathbf u+\mathbf v)}{\partial(\mathbf u+\mathbf v)}. \end{aligned} \tag{2.18} \]
Substituting \(\mathbf P_{\mathrm{II}}\) from the first of these formulas into the second and comparing with the third, we obtain:
\[ \frac{\partial F(\mathbf u+\mathbf v)}{\partial(\mathbf u+\mathbf v)} = \frac{\partial F(\mathbf u)}{\partial \mathbf u} + \frac{\partial F(\mathbf v)}{\partial \mathbf v} \tag{2.19} \]
—a functional equation which shows that \(\dfrac{\partial F}{\partial \mathbf u}\) is a linear and homogeneous function of its argument:
\[ \frac{\partial F}{\partial \mathbf u}=M\mathbf u, \tag{2.20} \]
where \(M\) is a constant, characteristic of the system but independent of its state. For the function \(F\) itself, obviously, one obtains:
\[ F(\mathbf u)=\frac{M u^2}{2}. \tag{2.21} \]
Finally, the transformation law for the components of the measure of motion has the form
\[ \left. \begin{aligned} \mathcal{E}_{\mathrm{II}}&=\mathcal{E}_{\mathrm{I}}-\mathbf{u}\mathbf{P}_{\mathrm{I}}+\frac{M u^{2}}{2},\\ \mathbf{P}_{\mathrm{II}}&=\mathbf{P}_{\mathrm{I}}-M\mathbf{u}. \end{aligned} \right\} \tag{2.22} \]
The meaning of the constant \(M\) can be clarified in the following way. If, in the reference frame I, which has been chosen quite arbitrarily, the momentum of the system is \(\mathbf{P}_{\mathrm{I}}\), then one can always find such an inertial reference frame II in which the momentum will be zero. For this it is sufficient to take the velocity of II with respect to I equal to
\[ \mathbf{u}=\frac{\mathbf{P}_{\mathrm{I}}}{M}. \tag{2.23} \]
By definition, we shall regard our physical system as having zero velocity in this new reference frame (as being at rest), i.e., we shall regard the velocity as zero if the momentum is equal to zero. The reference frame in which our physical system is at rest will henceforth be regarded as the fundamental one. Then the energy and momentum in it will be
\[ \mathcal{E}(0),\quad \mathbf{P}(0)=0. \tag{2.24} \]
If we now pass to a reference frame II moving in I with velocity \(\mathbf{u}\), then it may be regarded, again by definition, that the velocity of our physical system in reference frame II will be
\[ \mathbf{v}=-\mathbf{u}. \tag{2.25} \]
This definition is quite natural, but it is precisely a definition, since it is by no means clear what should be called the velocity of a system consisting, for example, of several particles, not to mention systems containing no matter at all. For the energy and momentum of the system at velocity \(\mathbf{v}\) (now this is already the velocity of the system itself) we obtain from the transformation law (2.22):
\[ \left. \begin{aligned} \mathcal{E}(\mathbf{v})&=\mathcal{E}(0)+\mathbf{v}\mathbf{P}(0)+\frac{M v^{2}}{2} =\mathcal{E}(0)+\frac{M v^{2}}{2},\\ \mathbf{P}(\mathbf{v})&=\mathbf{P}(0)+M\mathbf{v}=M\mathbf{v}. \end{aligned} \right\} \tag{2.26} \]
These formulas have exactly the same form as formulas (1.21) and (1.22) for a particle, and \(M\) can, obviously, be called the inertial mass of the physical system. Thus, for any physical system the momentum is proportional to its “velocity,” and the excess of the energy over the energy of “rest” (i.e., the state with zero momentum) is pro—
proportional to the square of the velocity. Energy and momentum are connected with velocity by one and the same quantity—the inertial mass, which does not depend on the state of the system (in particular, on its “velocity”) and is determined only by the nature of the system. The existence of the law of transformation of the components of the measure of motion, i.e., of the relation between these components in different inertial frames of reference, makes it possible to regard the measure of motion as a single complex quantity reflecting the absoluteness of motion.
§ 3. THE MEASURE OF MOTION IN THE THEORY OF RELATIVITY
The method applied in the preceding sections to the investigation of the concept of the measure of motion is usually considered characteristic of the new, so-called relativistic physics. However, the principle of relativity, which lies at the basis of this method, was already known in “classical” physics, and it is not this principle that distinguishes the new physics from the old. As often happens, the discovery of new phenomena makes it possible to understand more deeply regularities that have long been known, and features of new phenomena that seem entirely new and incomprehensible are found in embryonic form in phenomena long known and familiar. The arguments set forth in the preceding sections were intended to show that the method of the theory of relativity is, in essence, not characteristic of it. What is characteristic of the theory of relativity are the new properties of space and time connected with the existence of a critical (or maximum) velocity. The conceptions of space and time that underlay classical physics proved to be only approximately correct, and only approximately correct are the formulae for the addition of velocities that we used in the investigation of the measure of motion. Therefore, although all the qualitative conclusions obtained in the classical theory remain unchanged in the theory of relativity—and these qualitative conclusions are the most important—the quantitative relations must be replaced by more exact ones.
In this and the following sections the theory of the measure of motion will be considered anew, taking into account the new properties of space and time revealed by the theory of relativity, but the entire investigation will be conducted by exactly the same method as in the classical theory. Thus it will be clear that many results of the theory of relativity have in essence long been known, and if these results are disputed, then the classical laws too, long known and seeming even obvious, should also be disputed.
It is first necessary to clarify what new elements the theory of relativity has introduced into our conceptions of space and time. In classical physics it was assumed that, although the change in the position of some phenomenon relative to the reference particles of inertial frames of reference proceeds differently in different frames of reference, nevertheless time
has an absolute character, so that one may speak of the place of the phenomenon in different frames of reference at one and the same moment of time. The basis of this conviction is the certainty that there exists a single instant separating, throughout all space, the past from the future. If one tries to determine what is meant by this, the following emerges.
With respect to some event \(A\), occurring at some moment in a definite place, we divide all other events (occurring both in this same place and throughout all space) into past and future.
The past is that which can (or in principle could) influence event \(A\), but which \(A\) in principle cannot influence. The future, on the other hand, is that which \(A\) can influence (or could influence), but which itself in principle cannot influence \(A\). Everything past for event \(A\) is separated from the future for this same event by one instant, and everything that occurs at this instant throughout all space occurs “now.” Since this division is based on the possibility of a causal connection between events, it has an absolute character, i.e. it does not depend on any frames of reference.
In these conceptions it is correct that the basis for dividing all events into past and future relative to event \(A\) is the possibility of a direct causal connection between these events and \(A\). But the idea that the events separating the past from the future last only one instant is incorrect. The point is that in our space and time all actions are transmitted with velocities that are never greater than a certain limiting (or critical) velocity \(c\). (Light in a vacuum propagates precisely with this velocity, and therefore it is often called the “speed of light.” It is approximately equal to 300,000 km/sec.) This means that if the division of events into past and future is connected with the possibility of a causal connection between them, then at some point of space, distant from the place where event \(A\) occurs by a distance \(l\), the past for \(A\) will be separated from its future not by a single instant, but by an entire interval of duration \(2l/c\). Everything that occurs during this interval of time is detached from \(A\), since between these phenomena and \(A\) there can be no direct causal connection in either direction: actions issuing from these events will simply reach the place where \(A\) occurs only after \(A\) has already occurred, just as actions issuing from \(A\) will reach the point of space under consideration only after all the events of this interval have ended. Therefore “now” relative to \(A\), understood in the usual sense, is in reality not an instant, but a more or less extended interval of time, very short in places close to \(A\), but very long in places...
distant from \(A\). Any instant of this interval of time at any place is in no way distinguished, by the character of its connection with \(A\), from any other instant of this same interval, and therefore there is no instantaneous “now” in nature. The idea of a single instantaneous “now” for all of space simply does not reflect the actual properties of space and time.
The details are not important to us here. The problem of measuring time reduces, evidently, to how, from all the instants of the “intervals of simultaneity,” to single out one instant that would be “isochronous” with respect to \(A\). In this connection the concept of “isochrony” will be a new concept, differing from the ordinary concept of simultaneity. It turns out that the concept of “isochrony,” and along with it the measurement of time, can be established in each inertial frame of reference, with the definition of this concept being the same for all inertial frames of reference; but the measurement of time based on this concept leads, in different frames of reference, to different times for the same events. Let us repeat: the details are not important to us here. We shall need only to establish the relation between the spatial coordinates and the time of some event in one inertial frame of reference and the corresponding quantities in another. If we take rectangular coordinate axes in both frames of reference parallel to one another, count time in both frames of reference from the moment when both origins of coordinates coincided, and choose the \(x\)-axes so that the second frame moves relative to the first along the \(x\)-axis with velocity \(u\), then the relation between the coordinates and times of any event in both frames of reference will be expressed by the so-called Lorentz formulas. If, in addition, the critical velocity \(c\) is taken to be equal to unity (so that our velocity \(\mathbf{u}\) is the ratio of the velocity, measured in ordinary units, to the critical velocity), then these formulas will have the form:
\[ \left. \begin{aligned} t_{\mathrm{II}} &= \frac{t_{\mathrm{I}} - u x_{\mathrm{I}}}{\sqrt{1-u^2}},\\ x_{\mathrm{II}} &= \frac{x_{\mathrm{I}} - u t_{\mathrm{I}}}{\sqrt{1-u^2}},\\ y_{\mathrm{II}} &= y_{\mathrm{I}}, \qquad z_{\mathrm{II}} = z_{\mathrm{I}}. \end{aligned} \right\} \tag{3.1} \]
Our task now will be to find the measure of motion of a particle as a function of its velocity in some inertial frame of reference. This problem is solved, in essence, in exactly the same way as in § 1. But since the relation between the velocities of the particle in different frames of reference (the “rule of addition of velocities”) will now be more complicated than in the classical theory, all the calculations also become ...
more cumbersome*). Nevertheless, all the principal results—the existence, alongside the scalar measure of motion, of a vector measure as well; their inseparable connection, expressed in a definite relation of the two measures in different frames of reference to one another, etc.—remain unchanged.
The calculations may be carried out as follows. First of all, it turns out to be inconvenient to characterize the state of motion of a particle in a given frame of reference by its velocity relative to this frame. Velocity is the displacement of a particle in space per unit time of the frame of reference; and since, in passing to a new frame of reference, not only the displacement but also the time changes, the “law of addition of velocities” turns out to be complicated. It is much more convenient to characterize the state of motion of a particle by a velocity calculated with respect to the particle’s proper time, i.e. the displacement of the particle in space per unit time measured by clocks moving together with the particle (or by the time of that frame of reference in which the particle is at rest at the given instant). We shall denote the ordinary velocity of the particle by \(\mathbf v\), this new “velocity” by \(\mathbf w\), and the “proper time of the particle” by \(\tau\). Then
\[ \mathbf w=\frac{d\mathbf r}{d\tau}=\frac{d\mathbf r}{dt}\frac{dt}{d\tau}=\mathbf v\frac{dt}{d\tau}. \tag{3.2} \]
But from (3.1) it follows that
\[ \frac{dt}{d\tau}=\frac{1}{\sqrt{1-v^2}} \tag{3.3} \]
(the proper time flows more slowly than the time of the frame of reference). Therefore
\[ \mathbf w=\frac{\mathbf v}{\sqrt{1-v^2}}, \tag{3.4} \]
and, conversely, if \(\mathbf v\) is expressed through \(\mathbf w\),
\[ \mathbf v=\frac{\mathbf w}{\sqrt{1+w^2}}. \tag{3.5} \]
Thus the new “velocity” is in one-to-one correspondence with the ordinary velocity and can therefore characterize the state of motion of the particle just as well as the latter.
*) They can be simplified by introducing certain purely technical changes into the discussion. Since these changes do not at all affect the essence of the matter, the reader not interested in technical details may simply skip the calculations and proceed directly to the result—the formulas (3.20), which give a more exact relation between the velocity of the particle and the measure of its motion.
If in reference system I a particle moved with “velocity” \(\mathbf w_{\mathrm I}\), then in the new reference system II, moving along the \(x\)-axis of the old reference system with (ordinary) velocity \(u\), the particle’s “velocity” \(\mathbf w_{\mathrm{II}}\) will be different. It can easily be calculated from the Lorentz formulas, since \(d\tau\), by its definition, will be the same in both reference systems, while \(dx, dy, dz\) are directly calculated from (3.1). We then obtain:
\[ \left. \begin{aligned} \mathfrak w_{\mathrm{II}x} &= \frac{\mathfrak w_{\mathrm I x}-u\sqrt{1+\mathfrak w_{\mathrm I}^{2}}} {\sqrt{1-u^{2}}},\\ \mathfrak w_{\mathrm{II}y}&=\mathfrak w_{\mathrm I y},\qquad \mathfrak w_{\mathrm{II}z}=\mathfrak w_{\mathrm I z}, \end{aligned} \right\} \tag{3.6} \]
and since we shall need these formulas only for velocities \(u\) that are very small in comparison with the critical velocity, i.e. for \(u \ll 1\), we shall neglect \(u^{2}\) in them and obtain:
\[ \left. \begin{aligned} \mathfrak w_{\mathrm{II}x} &= \mathfrak w_{\mathrm I x}-u\sqrt{1+\mathfrak w_{\mathrm I}^{2}}+\ldots,\\ \mathfrak w_{\mathrm{II}y}&=\mathfrak w_{\mathrm I y},\qquad \mathfrak w_{\mathrm{II}z}=\mathfrak w_{\mathrm I z}. \end{aligned} \right\} \tag{3.7} \]
It is this “rule of addition” that we shall have to use instead of the former rule (1.2).
Let us now denote the scalar measure of motion of a particle (energy), regarding it as a function of the “velocity” \(\mathbf w\), by
\[ \varepsilon(\mathbf w)=\varepsilon(w^{2}) \tag{3.8} \]
(the scalar measure of motion depends, of course, only on the magnitude of the “velocity,” and not on its direction). The law of conservation of the scalar measure of motion in the collision of two particles will have the form:
\[ \varepsilon_{1}(\mathbf w_{1})+\varepsilon_{2}(\mathbf w_{2}) = \varepsilon_{1}(\mathbf w'_{1})+\varepsilon_{2}(\mathbf w'_{2}), \tag{3.9} \]
where the indices 1 and 2 denote the first and second particles, and the prime refers to quantities after the collision. If we now consider this collision in a new reference system, moving along the \(x\)-axis of the old system with a very small velocity \(u\), then from formulas (3.7) we obtain:
\[ \begin{aligned} \varepsilon_{1}(\mathbf w_{\mathrm{III}}) &= \varepsilon_{1}\!\left( \mathfrak w_{1x}-u\sqrt{1+\mathbf w_{1}^{2}}, \,\mathfrak w_{1y},\,\mathfrak w_{1z} \right) \\ &= \varepsilon_{1}(\mathbf w_{1}) - u\sqrt{1+\mathbf w_{1}^{2}}\, \frac{\partial \varepsilon_{1}(\mathbf w_{1})}{\partial \mathfrak w_{1x}} +\ldots \end{aligned} \tag{3.10} \]
and exactly the same for the other energies.
Then the law of conservation of energy in the new reference system will take the form
\[ \sqrt{1+w_1^2}\,\frac{\partial e_1}{\partial w_{1x}} + \sqrt{1+w_2^2}\,\frac{\partial e_2}{\partial w_{2x}} = \]
\[ = \sqrt{1+w_1^{\prime 2}}\,\frac{\partial e'_1}{\partial w'_{1x}} + \sqrt{1+w_2^{\prime 2}}\,\frac{\partial e'_2}{\partial w'_{2x}} . \tag{3.11} \]
Two more analogous relations are obtained if one passes to reference systems moving slowly relative to the old reference system along the axes \(y\) and \(z\). These relations, evidently, have the form of a conservation law for the vector measure of motion—the momentum, for which the formula is obtained
\[ \mathbf p=\sqrt{1+w^2}\,\frac{\partial e}{\partial \mathbf w}, \tag{3.12} \]
which now replaces formula (1.17).
Further one can reason exactly as in the classical theory. In each of the three conservation laws for the components of momentum one can again pass to new inertial reference systems and obtain new conservation laws—the laws of conservation of a tensor measure of motion whose component along the axes \(i\) and \(k\) is obtained if, in the conservation law for the \(i\)-th component of momentum, one passes to a reference system moving along the axis \(k\). Since each passage to a new inertial system gives an operation of the form
\[ \sqrt{1+w^2}\,\frac{\partial}{\partial w_k}, \]
the component \((ik)\) of the tensor measure will have the form
\[ s_{ik}=\sqrt{1+w^2}\,\frac{\partial}{\partial w_k} \left[ \sqrt{1+w^2}\,\frac{\partial e}{\partial w_i} \right]. \tag{3.13} \]
These new conservation laws must, as in the classical theory, either be satisfied in a trivial way or be a consequence of the conservation of scalar and vector measures. In addition, the tensor \(s_{ik}\) must be isotropic (for all directions of space are equivalent, and the relation of \(s\) to \(w\) must not change under a rotation of the coordinate axes), i.e. it must have the form
\[ s_{ik}=\lambda\delta_{ik}=\lambda \begin{cases} 1, & i=k,\\ 0, & i\ne k. \end{cases} \tag{3.14} \]
The quantity \(\lambda\) must be a linear function of the energy:
\[ s_{ik}=[\chi e(w)+\beta]\delta_{ik}. \tag{3.15} \]
We now carry out the differentiation in formula (3.13). Since \(\varepsilon=\varepsilon(w^2)\), we obtain:
\[ s_{ik}=\sqrt{1+w^2}\,\frac{\partial}{\partial w_k} \left[\sqrt{1+w^2}\,\varepsilon'(w^2)\,2w_i\right]= \]
\[ =2(1+w^2)\varepsilon'(w^2)\delta_{ik} +\left[4(1+w^2)\varepsilon''(w^2)+2\varepsilon'(w^2)\right]w_iw_k, \tag{3.16} \]
where the prime denotes the derivative of \(\varepsilon\) with respect to its argument, i.e. with respect to \(w^2\). Comparing this with (3.15), we obtain:
\[ \begin{aligned} 2(1+w^2)\varepsilon'(w^2)&=\alpha\varepsilon+\beta,\\ 2(1+w^2)\varepsilon''(w^2)+\varepsilon'(w^2)&=0. \end{aligned} \tag{3.17} \]
The second of these equations is easily integrated and gives:
\[ \varepsilon=m_0\sqrt{1+w^2}+\varepsilon_0, \tag{3.18} \]
where \(m_0\) and \(\varepsilon_0\) are constants of integration, while the first equation (3.17) is satisfied if one takes \(\alpha=1\) and \(\beta=-\varepsilon_0\).
The momentum is now calculated from (3.12):
\[ \mathbf{p}=m_0\mathbf{w}, \tag{3.19} \]
and if we return again to the ordinary velocity by formula (3.4), then for the energy and momentum of the particle we obtain the formulas
\[ \left. \begin{aligned} \varepsilon&=\frac{m_0}{\sqrt{1-v^2}}+\varepsilon_0,\\ \mathbf{p}&=\frac{m_0\mathbf{v}}{\sqrt{1-v^2}}\;{}^{*}). \end{aligned} \right\} \tag{3.20} \]
Thus, in the relativistic theory, exactly as in the classical one, from the existence of a scalar measure of the motion of a particle there follows the existence of a vector measure, and in exactly the same way it proves possible to determine uniquely the character of the dependence of both these measures on the velocity of the particle. Only this dependence turns out to be different from that in the classical theory. But if, as is done in the classical theory, the critical velocity is regarded as very large and, consequently, \(v\) as very small, and \(v^2\) is neglected, then we obtain the old formulas (1.21) and (1.22). The relation of momentum to velocity is determined, as in classical theory, by a single quantity—the mass \(m_0\), only this relation is now different. In order not to confuse different things, we
\({}^{*}\) To pass to ordinary units, in these formulas one must make the substitutions
\[ v\to\frac{v}{c},\qquad \mathbf{p}\to\frac{\mathbf{p}}{c},\qquad \varepsilon\to\frac{\varepsilon}{c^2}. \]
we shall call the mass \(m_0\) the invariant inertial mass. It depends only on the nature of the particle, but not on the state of its motion, i.e., on its velocity, and is the same in all inertial frames of reference. One could define mass differently, writing formula (3.20) in the form
\[ \begin{aligned} \varepsilon &= m+\varepsilon_0,\\ \mathbf p &= m\mathbf v, \end{aligned} \qquad\} \tag{3.21} \]
where
\[ m=\frac{m_0}{\sqrt{1-v^2}}. \tag{3.22} \]
This (“variant”) mass would depend on the velocity of the particle.
The constant \(\varepsilon_0\) can no longer be regarded as the energy of a particle at rest, since for \(v=0\) one obtains:
\[ \varepsilon=m_0+\varepsilon_0. \]
On the other hand, without further investigation the constant \(\varepsilon_0\) cannot be taken equal to zero. This could be done if we were considering only the given particle and if it could not transform into other particles or disappear, since in the final analysis only the change in energy is significant. But if we set \(\varepsilon_0=0\) for one particle, then for other particles—for example, for a particle that could result from the fusion of two particles with \(\varepsilon_0=0\)—\(\varepsilon_0\) could no longer be chosen arbitrarily. It is clear that the question of the magnitude of \(\varepsilon_0\) is connected with the question of the energy and momentum of arbitrary physical systems. Therefore we shall postpone the discussion of this question; but if we consider some particle separately, we shall take \(\varepsilon_0=0\) for it.
In this case the formulas for energy and momentum may be rewritten so that the “transformation law” for these quantities becomes evident, i.e., the relation between the components of the measure of motion computed with respect to different frames of reference. To do this, we introduce into formula (3.20) the element of the particle’s proper time according to formula (3.3), after which the energy and momentum are written as follows:
\[ \begin{aligned} \varepsilon &= m_0\,\frac{dt}{d\tau},\\ \mathbf p &= m_0\,\frac{d\mathbf r}{d\tau}. \end{aligned} \qquad\} \tag{3.23} \]
These formulas are remarkable, first, because from them the transformation law is clear: when passing to a new inertial frame of reference, energy and momentum are recalculated exactly as the time and radius vector of the particle are, i.e., simply according to the Lorен-
since (3.1), since \(d\tau\) is invariant.
\[ \left. \begin{aligned} \varepsilon_{\mathrm{II}}&=\frac{\varepsilon_{\mathrm{I}}-u p_{\mathrm{I}x}}{\sqrt{1-u^2}},\\ p_{\mathrm{II}x}&=\frac{p_{\mathrm{I}x}-u\varepsilon_{\mathrm{I}}}{\sqrt{1-u^2}},\\ p_{\mathrm{II}y}&=p_{\mathrm{I}y},\quad p_{\mathrm{II}z}=p_{\mathrm{I}z}. \end{aligned} \right\} \tag{3.24} \]
In geometry, quantities whose components transform, upon transition to a new coordinate system, like the differentials of the coordinates are called (contravariant) vectors, for they can obviously be represented by line segments. By analogy one may say that the momentum of a particle, having components \((\varepsilon, p_x, p_y, p_z)\), is a (contravariant) vector in the four-dimensional space-time manifold of the world. This geometrical analogy makes the idea of the momentum of a particle as an absolute quantity very vivid.
Secondly, formulas (3.23) show that the momentum of a particle is simply proportional to the displacement in space-time of the reference system per unit of its proper time. This remarkable fact, which is essentially quite obvious (for what else can measure the motion of a particle, if not its displacement in the world manifold?), could have been taken as the basis of the entire theory of momentum.
§ 4. GENERAL THEORY OF MOMENTUM IN THE THEORY OF RELATIVITY AND THE LAW OF EQUIVALENCE OF MASS AND ENERGY
Let us now turn to the investigation of the momentum of any physical system. As in classical theory, let us assume (here the arguments of § 2 are repeated literally) that any physical system in any of its states has a scalar measure of motion—energy—depending on the state of the system but, of course, not determined solely by the velocity of the system, which may or may not exist. However, on whatever quantities the energy may depend, it will be different with respect to different inertial reference systems. For the proof it suffices, as before, to consider the transfer of motion by the system to some particle of matter. The decrease in the energy of the system as a result of this process will be equal to the increase in the energy of the particle, and this latter will be different with respect to different inertial reference systems. Therefore the decrease in the energy of any system upon its collision with a particle will also be different in different reference systems. This could not be the case if the energy of the system itself was
in a given state would be the same in different frames of reference. The energy of any physical system is just as relative a quantity as the energy of a particle. Only by specifying with respect to which inertial frame of reference the motion is measured can one determine the magnitude of its scalar measure.
Taking some arbitrarily chosen inertial frame of reference as the basic one and characterizing all other inertial frames of reference by their velocities \(\mathbf u\) in the basic frame, one can find the energy of any physical system in any of its states with respect to any frame of reference. This energy will then be a function of its velocity \(\mathbf u\) (and not of the velocity of the physical system itself!), and we shall denote it by
\[ \mathcal E(\mathbf u). \tag{4.1} \]
Further one may reason exactly as in the classical theory. Conservation of energy in the collision of the system under consideration with a particle gives, in any inertial frame of reference,
\[ \mathcal E(\mathbf u)+\varepsilon(\mathbf u) = \mathcal E'(\mathbf u)+\varepsilon'(\mathbf u), \tag{4.2} \]
where a prime denotes the states after the collision, and \(\varepsilon(\mathbf u)\) is the energy of the particle, also as a function of the velocity of the frame of reference \(\mathbf u\). Since \(\mathbf u\) is arbitrary, differentiating with respect to \(\mathbf u\) and then putting \(\mathbf u=0\), we obtain:
\[ \left[\frac{\partial \mathcal E(\mathbf u)}{\partial \mathbf u}\right]_{\mathbf u=0} + \left[\frac{\partial \varepsilon(\mathbf u)}{\partial \mathbf u}\right]_{\mathbf u=0} = \left[\frac{\partial \mathcal E'(\mathbf u)}{\partial \mathbf u}\right]_{\mathbf u=0} + \left[\frac{\partial \varepsilon'(\mathbf u)}{\partial \mathbf u}\right]_{\mathbf u=0}. \tag{4.3} \]
It is easy to verify that for a particle
\[ \left[-\,\frac{\partial \varepsilon(\mathbf u)}{\partial \mathbf u}\right]_{\mathbf u=0} = \mathbf p \tag{4.4} \]
(this follows immediately from (3.24), where \(\varepsilon_{II}=\varepsilon(\mathbf u)\)), and it is clear that (4.3) gives the law of conservation of momentum, which for any physical system is determined, exactly as in the classical theory, by the formula
\[ \mathbf P = \left[-\,\frac{\partial \mathcal E(\mathbf u)}{\partial \mathbf u}\right]_{\mathbf u=0}. \tag{4.5} \]
Thus, in the relativistic theory too, the measure of motion has a scalar and a vector component, different with respect to different inertial frames of reference. The unity of these two measures is expressed in the law of their transformation, which will, of course, differ from the classical one and which we must find. It seems obvious that this law must be the same as for a particle, i.e.
must be expressed by formulas (3.24). But since the question of the transformation law is connected with the question of the additive constant in the energy (formulas (3.24) are valid only for \(\varepsilon_0=0\)), this law must be derived independently.
In the basic frame of reference, which here plays the role of the first, the laws of conservation of energy and momentum give:
\[ \begin{aligned} \mathcal{E}'(0)-\mathcal{E}(0)&=-[\varepsilon'(0)-\varepsilon(0)],\\ \mathbf{P}'(0)-\mathbf{P}(0)&=-[\mathbf{p}'(0)-\mathbf{p}(0)]. \end{aligned} \tag{4.6} \]
In some other frame of reference II, moving in the basic one with velocity \(u\), these same laws will be:
\[ \begin{aligned} \mathcal{E}'(u)-\mathcal{E}(u)&=-[\varepsilon'(u)-\varepsilon(u)],\\ \mathbf{P}'(u)-\mathbf{P}(u)&=-[\mathbf{p}'(u)-\mathbf{p}(u)]. \end{aligned} \tag{4.7} \]
Next, using formulas (3.24), which are applicable here, since \(\varepsilon_0\) does not enter into the change of the particle energy, we obtain:
\[ \mathcal{E}'(u)-\mathcal{E}(u)= \frac{-[\varepsilon'(0)-\varepsilon(0)]+u[p'_x(0)-p_x(0)]}{\sqrt{1-u^2}}, \]
and from (4.6)
\[ \mathcal{E}'(u)-\mathcal{E}(u)= \frac{[\mathcal{E}'(0)-\mathcal{E}(0)]-u[P'_x(0)-P_x(0)]}{\sqrt{1-u^2}}. \]
It follows from this that
\[ \mathcal{E}(u)= \frac{\mathcal{E}(0)-uP_x(0)}{\sqrt{1-u^2}}+F(u), \tag{4.8} \]
where \(F(u)\) is an as yet undetermined function of the velocity \(u\), the same for all states of the system under consideration. It obviously must not depend on the direction of \(\mathbf{u}\), so that one may write:
\[ F(u^2). \tag{4.9} \]
In exactly the same way, starting from the law of conservation of momentum, one could also obtain formulas for the transformation of momentum. But new arbitrary functions would enter into these formulas, whose relation to \(F\) would remain unknown. Therefore we shall proceed differently: regarding frame of reference II as the basic one, we pass to a third frame of reference, moving relative to II with arbitrary velocity \(v\). Its velocity in the basic frame of reference will be, by
according to the velocity-addition rule, easily derived from the Lorentz formulas, have the following components:
\[ \left. \begin{gathered} \frac{u+v_x}{1+uv_x},\\[6pt] \frac{v_y\sqrt{1-u^2}}{1+uv_x},\\[6pt] \frac{v_z\sqrt{1-u^2}}{1+uv_x}. \end{gathered} \right\} \tag{4.10} \]
Next we apply the definition of momentum (4.5): we differentiate the energy in reference frame III with respect to \(\mathbf v\) and put \(\mathbf v=0\). Straightforward calculations then give:
\[ P_x(\mathbf u)=\frac{P_x(0)-u\mathcal E(0)}{\sqrt{1-u^2}}-(1-u^2)\frac{\partial F}{\partial u_x}, \]
\[ P_y(\mathbf u)=P_y(0)-\sqrt{1-u^2}\,\frac{\partial F}{\partial u_y}, \]
\[ P_z(\mathbf u)=P_z(0)-\sqrt{1-u^2}\,\frac{\partial F}{\partial u_z}. \]
Since \(F=F(u^2)\), and \(u_y=u_z=0\), denoting the derivative with respect to \(u^2\) by a prime, we obtain:
\[ \frac{\partial F}{\partial u_x}=F'(u^2)2u_x=2uF'(u^2), \]
\[ \frac{\partial F}{\partial u_y}=F'(u^2)2u_y=0, \]
\[ \frac{\partial F}{\partial u_z}=F'(u^2)2u_z=0, \]
so that the transformation law for momentum takes the simpler form
\[ \left. \begin{gathered} P_x(\mathbf u)=\frac{P_x(0)-u\mathcal E(0)}{\sqrt{1-u^2}}-2u(1-u^2)F'(u^2),\\[6pt] P_y(\mathbf u)=P_y(0);\qquad P_z(\mathbf u)=P_z(0). \end{gathered} \right\} \tag{4.11} \]
Thus, the transformation laws (4.8) and (4.11) contain only one still unknown function \(F(u^2)\). But this too can be determined if we take into account that all transitions from inertial reference frames to other such frames form a group, or, more simply, that a successive transition from one reference frame to another and then to a third is equivalent to a direct transition from the first to the third. Therefore let us consider three reference frames: I—the basic one, II, moving in the basic frame along the \(x\)-axis with velocity \(u\), and III, moving in II again along the \(x\)-axis with velo-
…with velocity \(v\) and, consequently, moving in the principal one also along the \(x\)-axis with velocity
\[ w=\frac{u+v}{1+uv}. \tag{4.12} \]
Let now, in the principal frame of reference, the momentum of our physical system be equal to zero (this is needed only to simplify the calculations):
\[ \mathbf P_1=0. \]
Then, passing to reference frame II by formulas (4.8) and (4.11), we obtain:
\[ \left. \begin{aligned} \mathscr E_{\mathrm{II}}&=\frac{\mathscr E_{\mathrm I}}{\sqrt{1-u^2}}+F(u^2),\\[4pt] P_{\mathrm{II}x}&=-\frac{u\mathscr E_{\mathrm I}}{\sqrt{1-u^2}}-2u(1-u^2)F'(u^2). \end{aligned} \right\} \tag{4.13} \]
Next we pass from reference frame II to frame III and compute the energy \(\mathscr E_{\mathrm{III}}\) (we shall not need the momentum). Applying (4.13), we obtain
\[ \mathscr E_{\mathrm{III}} =\frac{\mathscr E_{\mathrm{II}}-vP_{\mathrm{II}x}}{\sqrt{1-v^2}}+F(v^2)= \]
\[ =\frac{\mathscr E_{\mathrm I}(1+uv)} {\sqrt{1-u^2}\sqrt{1-v^2}} +\frac{F(u^2)+2uv(1-u^2)F'(u^2)} {\sqrt{1-v^2}} +F(v^2). \tag{4.14} \]
But this same energy can be computed by passing directly from frame I to frame III:
\[ \mathscr E_{\mathrm{III}} =\frac{\mathscr E_{\mathrm I}(1+uv)} {\sqrt{1-u^2}\sqrt{1-v^2}} +F\left(\left[\frac{u+v}{1+uv}\right]^2\right), \tag{4.15} \]
and the same result must be obtained.
Therefore, comparing (4.14) and (4.15), we obtain for \(F\) the functional equation
\[ \frac{F(u^2)+2uv(1-u^2)F'(u^2)} {\sqrt{1-v^2}} +F(v^2) = F\left(\left[\frac{u+v}{1+uv}\right]^2\right). \tag{4.16} \]
Expanding both its sides in powers of \(u\) and comparing the coefficients of the expansions gives a series of equations:
\[ \left. \begin{aligned} F(0)&=0,\\ F'(0)&=(1-v^2)^{3/2}F'(v^2),\\ \ldots \end{aligned} \right\} \tag{4.17} \]
Integrating the second of them and determining the constant of integration from the first, we obtain, if we denote
\[ 2F'(0)=N, \tag{4.18} \]
for \(F\) the expression
\[ F(v^2)=N\left[\frac{1}{\sqrt{1-v^2}}-1\right]. \tag{4.19} \]
It is easy to verify that this function satisfies the complete equation (4.16).
We shall substitute this expression into formulas (4.8) and (4.11), and after elementary calculations write the transformation law for the components of the measure of motion of any physical system in the form
\[ \left. \begin{aligned} \mathcal{E}_{\parallel}+N &=\frac{(\mathcal{E}_{1}+N)-uP_{1x}}{\sqrt{1-u^2}},\\[4pt] P_{\parallel x} &=\frac{P_{1x}-u(\mathcal{E}_{1}+N)}{\sqrt{1-u^2}},\\[4pt] P_{\parallel y}&=P_{1y},\qquad P_{\parallel z}=P_{1z}. \end{aligned} \right\} \tag{4.20} \]
The constant \(N\), like the function \(F\), depends only on the nature of the system, but not on its state. Obviously, for each system one can change the definition of energy, calling \(\mathcal{E}+N\) the energy. The addition of this constant will change nothing in the conservation laws if, in the process under consideration, the system remains the same object, since then the constant \(N\) will be added both before and after the process.
If, however, for example, two systems combine into one, and for the initial systems the energy is normalized so that \(N_1=N_2=0\), then for the combined system the quantity \(N\) will automatically turn out to be zero. Indeed, in some inertial frame of reference we shall have the conservation law
\[ \mathcal{E}_1+\mathcal{E}_2=\mathcal{E}, \]
\[ \mathbf{P}_1+\mathbf{P}_2=\mathbf{P}, \]
and after transition to any other frame of reference from these one obtains (by (4.20)):
\[ \frac{\mathcal{E}_1-uP_{1x}}{\sqrt{1-u^2}} + \frac{\mathcal{E}_2-uP_{2x}}{\sqrt{1-u^2}} = \frac{(\mathcal{E}+N)-uP_x}{\sqrt{1-u^2}}-N, \]
or, by virtue of the conservation of energy and momentum in the original frame of reference,
\[ 0=\frac{N}{\sqrt{1-u^2}}-N. \]
Since \(u\) is arbitrary, \(N=0\) (in particular, for a particle in formula (1.21) one may take \(\varepsilon_0=0\)).
The transformation formulas will then become homogeneous, and the energy and momentum of any system will be related to the same quantities in another frame of reference as are time and coordinates.
Consequently, the measure of motion of any physical system is a vector in the four-dimensional space-time manifold of the world, and in this its absolute character is expressed. The separation of the measure into scalar and vector parts is different in different inertial frames of reference and is characterized by the transformation law:
\[ \begin{gathered} \mathcal{E}_{\mathrm{II}}=\frac{\mathcal{E}_{\mathrm{I}}-uP_{\mathrm{I}x}}{\sqrt{1-u^2}},\\ P_{\mathrm{II}x}=\frac{P_{\mathrm{I}x}-u\mathcal{E}_{\mathrm{I}}}{\sqrt{1-u^2}},\\ P_{\mathrm{II}y}=P_{\mathrm{I}y},\qquad P_{\mathrm{II}z}=P_{\mathrm{I}z}. \end{gathered} \tag{4.21} \]
Let us note that from these formulas there follows, as is easily verified, the relation
\[ \mathcal{E}_{\mathrm{II}}^2-P_{\mathrm{II}}^2=\mathcal{E}_{\mathrm{I}}^2-P_{\mathrm{I}}^2, \tag{4.22} \]
so that the difference of the squares of energy and momentum is the same in all frames of reference (is invariant). Later we shall discuss the meaning of this fact.
Now let us introduce the concept of the velocity of any physical system, just as was done in the classical theory. Obviously, velocity in the ordinary sense does not exist for every physical system, since different parts of one system may move in quite different ways. Consequently, we must define what we shall call velocity, and this must be done so that the new definition coincides with the usual one in those cases where the old concept has meaning, i.e., for a single particle.
Consider some physical system having energy \(\mathcal{E}_{\mathrm{I}}\) and momentum \(P_{\mathrm{I}}\) in some inertial frame of reference I. Does there exist another such frame of reference in which the momentum of the system under consideration will be equal to zero? If it exists and if its velocity in I is \(u\), then the transformation formulas (4.21) give:
\[ 0=\frac{P_{\mathrm{I}x}-u\mathcal{E}_{\mathrm{I}}}{\sqrt{1-u^2}}, \]
\[ 0=P_{\mathrm{I}y},\qquad 0=P_{\mathrm{I}z}. \]
This means that the new frame of reference must move in the old one in the direction of the momentum of our system with velocity
\[ u=\frac{P_{\mathrm{I}}}{\mathcal{E}_{\mathrm{I}}}. \tag{4.23} \]
Such a velocity exists if
\[ P_1 < \mathcal{E}_1, \tag{4.24} \]
or, every velocity must be less than the critical one. This condition is invariant, i.e., it is a property of the state of the system, since it means that the invariant \(\mathcal{E}_1^2 - P_1^2\) is positive. Is it fulfilled for all physical systems in all states? It may be noted that if this condition were not fulfilled for some system, i.e., if for some system its energy turned out to be less than its momentum, then one could indicate such an inertial frame of reference in which the energy would be equal to zero with a momentum different from zero. Indeed, for this it would be necessary to take the velocity \(u\) so that
\[ \mathcal{E}_{11}= \frac{\mathcal{E}_1-uP_{1x}}{\sqrt{1-u^2}}=0, \]
and this velocity would turn out to be greater than the critical one. Such systems are not known in modern physics.
For every “normal” system, for which always \(\mathcal{E}^2-P^2>0\) (the limiting case \(\mathcal{E}^2-P^2=0\) we shall consider later), there always exists a frame of reference in which the momentum is equal to zero. We shall call this frame of reference the principal one and, by definition, shall assume that in it the velocity of the system under consideration is equal to zero, i.e., we shall assume that for
\[ \mathbf{P}=0 \quad \mathbf{v}=0. \]
In any other inertial frame of reference, moving relative to the principal one with velocity \((-\mathbf{v})\), we shall, again by definition, ascribe to the system the velocity \(\mathbf{v}\). Then any “normal” system in any of its states will have a definite velocity in any inertial frame of reference, and its energy and momentum will depend on this velocity, so that one may write:
\[ \mathcal{E}(\mathbf{v}), \quad \mathbf{P}(\mathbf{v}). \]
The transformation law (4.21) then gives:
\[ \left. \begin{aligned} \mathcal{E}(\mathbf{v})&=\frac{\mathcal{E}(0)}{\sqrt{1-v^2}},\\ \mathbf{P}(\mathbf{v})&=\frac{E(0)\,\mathbf{v}}{\sqrt{1-v^2}}. \end{aligned} \right\} \tag{4.25} \]
From these formulas it is clear, first, that our definition of velocity coincides with the usual one if the system is simply a particle—it is enough to compare (4.25) with (3.20) and note that for a particle \(\mathcal{E}(0)=m_0\). Secondly, since the relation of energy and momentum to velocity turns out to be exactly such for all physical systems
just as for particles, one can, for any normal systems, introduce the concept of mass. Obviously, the invariant mass of any system must be defined as follows:
$$ M_0=\mathcal{E}(0)^{*}), \tag{4.26} $$
and the variant mass by the relation
$$ M=\frac{M_0}{\sqrt{1-v^2}}=\frac{\mathcal{E}(0)}{\sqrt{1-v^2}}. \tag{4.27} $$
Then we obtain, for any normal system:
$$ \left. \begin{aligned} \mathcal{E}(\mathbf{v})&=\frac{M_0}{\sqrt{1-v^2}},\\ \mathbf{P}(\mathbf{v})&=\frac{M_0\mathbf{v}}{\sqrt{1-v^2}},\\ \mathcal{E}^2(\mathbf{v})-\mathbf{P}^2(\mathbf{v})&=M_0^2, \end{aligned} \right\} \tag{4.28} $$
and the concept of mass remains the same: it is a quantity determining the momentum of the system at its given velocity. But at the same time mass also determines the energy of the system; namely, the invariant mass is simply proportional to the rest energy, and the energy is always proportional to the variant mass. In ordinary units
$$ \mathcal{E}=Mc^2,\qquad \mathbf{P}=M\mathbf{v}. \tag{4.29} $$
This is Einstein’s equivalence law. Obviously, it is simply a consequence of the more general law of the existence of a unified measure of motion or, in other words, is an expression of the inseparable connection between the scalar and vector measures of motion.
In nature there exist systems (for example, photons) which may be regarded as a limiting case of normal ones. For them
$$ \mathcal{E}^2-\mathbf{P}^2=0. $$
Our definition of velocity is not suitable for them, since for them there is no fundamental frame of reference. But for them one may write:
$$ \mathbf{P}=\mathcal{E}\mathbf{c},\qquad c=1, \tag{4.30} $$
and this formula coincides with (4.29), with \(M=\mathcal{E}\), as for normal systems. Therefore such systems should be assigned a velocity whose direction coincides with the momentum (as for normal systems) and whose magnitude is equal to the critical one. Their invariant mass, however, should be regarded as equal to zero, since for them
$$ M_0^2=\mathcal{E}^2-\mathbf{P}^2=0. \tag{4.31} $$
*) In ordinary units \(M_0c^2=\mathcal{E}(0)\).
We note that the invariant mass \(M_0 = \sqrt{\mathcal{E}^2 - P^2}\) is not additive: if a system consists of two noninteracting parts, its invariant mass is not equal to the sum of the invariant masses of the parts:
\[ \sqrt{\mathcal{E}_1^2 - P_1^2} + \sqrt{\mathcal{E}_2^2 - P_2^2} \ne \sqrt{(\mathcal{E}_1 + \mathcal{E}_2)^2 - (\mathbf{P}_1 + \mathbf{P}_2)^2}. \]
Only at very small velocities, when one may take \(\mathcal{E} \simeq M_0\), and \(P \ll \mathcal{E}\), does approximate additivity result.
CONCLUSIONS
-
The motion of matter is absolute, since it exists independently of that with respect to which it is considered, and at the same time it is relative, since physical systems move relative to other physical systems. Therefore the measure of motion must likewise be both absolute and relative at once.
-
Matter moves in space and time, and the properties of the measure of its motion must not contradict the properties of space and time. But space and time are such that the existence of a relative scalar measure of motion (energy) is inseparably connected with the existence of a likewise relative vector measure of motion (momentum). Momentum is as universal a measure of motion as energy.
-
The scalar and vector measures of motion must be regarded as the components of a single complex measure having an absolute character, but decomposed differently into components in different frames of reference. The absolute character of this complex measure of motion is expressed in the existence of a transformation law for its components, i.e., in the existence of a connection between its components in different frames of reference.
The law of transformation of the components of the measure of motion is such that this complex measure itself may be regarded as a vector in the four-dimensional space-time manifold of the world.
- The principle of relativity makes it possible to determine unambiguously the character of the connection between the components of the measure of motion of any physical system and its velocity in an arbitrary inertial frame of reference, and makes it possible to introduce, for any physical system, the concept of inertial mass as a quantity connecting velocity and the measure of motion.
By its very concept, inertial mass is connected both with momentum and with energy, as with two components of a single measure of motion. The expression of this connection is the law of equivalence of mass and energy, whose root is the existence of an absolute measure of motion.