PRINCIPLES OF QUANTUM MECHANICS. I
K. V. Nikol'skii
Submitted 1936 | SovietRxiv: ru-193601.74370 | Translated from Russian

Full Text

PRINCIPLES OF QUANTUM MECHANICS. I

K. V. Nikolsky, Moscow

1. Quantum Physics and Classical Physics

All problems of modern physics may be divided into two groups: problems of classical physics and problems of quantum physics. In studying the properties of ordinary macroscopic bodies, one almost never has to deal with quantum problems, because quantum properties become perceptible only in the microworld. Therefore the physics of the nineteenth century, which investigated only macroscopic bodies, was wholly unaware of quantum processes. This is precisely “classical” physics. What is characteristic of classical physics is that it does not take into account the atomistic structure of matter. Today, however, the development of experimental technique has so greatly expanded the limits of our acquaintance with nature that we now know, and moreover in considerable detail, the structure of individual atoms and molecules. Modern physics studies the atomic structure of matter, and therefore the principles of old nineteenth-century classical physics had to be changed in accordance with the new facts, and changed radically. This change of principles is the transition to quantum physics.

Let us recall that nowadays any substance can be broken down into separate atoms, which in turn consist of electrons and nuclei having a complex structure. Nuclei consist of protons and neutrons.

The atomism of matter is expressed in the following facts. The mass of the above-mentioned so-called “elementary” particles, as experimental measurements show, always has a quite definite value. In this is expressed the atomism of mass. Let us further recall that the most important among the properties of nuclei and electrons are their electrical—more precisely, electromagnetic—properties, which are characterized by their charges. Measurements of the charge of elementary particles show that charge likewise has an atomic structure, i.e., the charge of elementary particles always has a quite definite, one and the same value. The atomism of mass and of the charge invariably connected with it in fact leads to the concept of elementary particles. However, it would be entirely wrong to identify the atomism of the microworld with the existence of particles, so to speak “specks of dust,” of which everything consists. It is entirely true that the atomism of mass and of the charge connected with it provides a basis for introducing the concept—

of the domain, namely, by posing the problem statistically. The statistical formulation of the problem will not be the elimination of ambiguity, but will be the method which, despite the ambiguity, makes it possible to characterize quantum processes quite objectively.

Let the quantum body \(A_1\) (we shall say—a quantum particle) have quantitative attributes \(K\) and \(R\). Let the attribute \(K\) be capable of taking certain values \(k_1, k_2, k_3, \ldots k_s, \ldots\), and the attribute \(R\) the values \(r_1, r_2 \ldots, r_s \ldots\)

As a result of a reaction of the type described, carried out for the purpose of determining the value of \(K\), we obtain, say, \(K = k_1\). Suppose further that a repetition of this same reaction under the same conditions with another specimen of this same quantum particle leads to \(K = k_s\). In this the ambiguity mentioned above is expressed. Repeating this reaction many times on new specimens under the same conditions, we obtain as a result some distribution of the specimens investigated according to the various values of the attribute \(K\). Let us say that \(N_1\) of them have \(K = k_1\), \(N_2\) have \(K = k_2\), and so on.

The distribution thus obtained makes it possible to make a statistical forecast for a new, as yet uninvestigated specimen, namely, the ratios

\[ \frac{N_1}{N};\quad \frac{N_2}{N};\quad \ldots \frac{N_r}{N};\quad \ldots \frac{N_s}{N};\quad \ldots \]

will determine for us the probabilities that the attribute \(K\) has, for the \((n+1)\)-st specimen, one or another value—\(k_1, k_2, \ldots k_r, \ldots k_s\), and so on. This statistical forecast also characterizes the quantum particle quite objectively.

Using the statistical method, we shall be able to formulate rigorously the basic difference between a classical and a quantum process. For this purpose we shall formulate the classical problems statistically and compare the results obtained in the quantum case with the results obtained in the classical case.

Suppose that in the classical case the determination of the value of some quantitative attribute \(K\) has given us \(N\) specimens for which it is established that \(K = k_n\). Here let us assume that \(N\) is sufficiently large. Considering the value of some other quantity, let us call it \(L\), we shall subject precisely these \(N\) specimens, for which it is known that \(K = k_n\), to a test. As a result, we shall be able to select from them a certain number of specimens \(N_s\), for which it turned out that \(L = l_s\). In the classical case we can assert that these \(N_s\) specimens have the attributes \(K = k_n\) and, at the same time, \(L = l_s\). For the quantum case, however, this assertion, generally speaking, is false. A specimen with \(L = l_s\) from among the \(N_s\) need not have \(K = k_s\), since the reaction that led to the fixation of the value \(L = l_s\) changes the situation in such a way that there may no longer be grounds for asserting that the value \(K = k_n\) has been preserved. In this lies the basic difference between quantum and classical processes.

The role of the quantum of action was especially emphasized by N. Bohr. We shall now give a strict quantitative mathematical formulation to the quantitative consequences of the finiteness of interactions, expressed by the quantum of action, which have just been noted.

First we shall develop, independently of physics, the statistical method most rational for these purposes, and then, by means of it, formulate the quantum laws quantitatively.

3. Statistical Method of Quantum Theory

Let us recall the basic statistical concepts with which we shall operate. These will be:

  1. The mean statistical value of some quantity;
  2. The mean square error, serving as a measure of the deviation of the mean value from the possible true value, or, as it is called, the “measure of dispersion” of the given quantity for the given statistical ensemble.

Let there be independent identical specimens of one and the same object, subjected to statistical study and characterized by quantitative attributes \(K, L, R, S\). One of these attributes must be completely fixed as the attribute according to which the specimens of the objects are combined into the statistical ensemble whose trials are being performed. Let us say that

\[ L = l^*. \]

Accordingly, we must mark all quantities relating to the ensemble under study with the index \(l^*\). In carrying out a statistical trial of our ensemble, we study the distribution of specimens according to the values of various quantities \(K, L\ldots\). In other words, we determine the probabilities

\[ w(k_1, l^*),\; w(k_2, l^*),\ldots w(k_s, l^*)\ldots \]

that the \(N+1\)-st object has this or that (among the possible) value of the quantity \(K\). Further, the probabilities

\[ w(r_1, l^*),\; w(r_2, l^*),\ldots w(r_s, l^*)\ldots \]

that some as yet untested object of the ensemble has this or that value of the quantity \(R\).

Let us assume, for simplicity, that the quantities \(K\) and \(R\) have an infinite discrete set of values

\[ k_1,\; k_2,\ldots k_s,\ldots \]

\[ r_1,\; r_2,\ldots r_s,\ldots \]

(All our results are directly generalized also to the case when the values of the quantities are continuous. In this case we must

This formula determines an approximation of the function \(f(x)\) by linear expressions

\[ \sum_{n=1}^{N} c(n)\,\varphi(n,x), \]

for increasing \(N\). This approximation has the minimal quadratic error, which is what is expressed by the formula written above. It is often written in the form of the expansion

\[ f(x)\sim \sum_{n=1}^{\infty} c(n)\varphi(n,x). \]

The coefficients \(c(n)\), which determine the representation of the function \(f(x)\) by means of the given system of functions \(\varphi(n,x)\), can be expressed in terms of \(f(x)\) and \(\varphi(n,x)\); namely, they are equal to

\[ c(n)=\int \overline{\varphi(n,x)}\,f(x)\,dx. \]

The expansion formula, i.e. the definition of completeness, may be written in the form

\[ \int |f(x)|^2\,dx=\sum_{n=1}^{\infty}|c(n)|^2 \quad \text{(Bessel equality)} \]

The condition that all \(c(n)=0\) is equivalent to the condition \(f(x)=0\). We shall now make use of these well-known relations. We see, first of all, that, taking any complete system of functions, we could expand with respect to it all our functions \(w(k,l)\), \(w(r,l)\), since they contain a common independent variable \(l\), which for the given aggregate has a prescribed value \(l=t^*\). The relations between the coefficients of these expansions could then serve to establish the connections that interest us. However, this path is not expedient because the functions \(w(k,l)\), \(w(r,l)\) are not quite arbitrary. They have the meaning of probabilities, and this must be explicitly formulated in the form of conditions to which these functions must be subject. Further, we must deal with functions whose squared modulus is integrable.

These conditions, in explicit form, state:

\[ w(k,l)\geq 0, \]

i.e. probability is a real, positive number

\[ w(k_n,k_m)=0 \quad \text{for } n\ne m \]

and

\[ w(k_n,k_n)=1. \]

(Incompatibility and certainty of the events \(k=k_n\) and \(k=k_m\).)
Further,

\[ \int \rho(k,l)\,dk=1, \qquad (N) \]

if \(k\) is continuous, and

\[ \sum_n w(k_n,l)=1, \qquad (N) \]

if \(k\) is discrete.

We see that these conditions are automatically satisfied if we pass from \(w(k,l)\) to auxiliary functions—let us call them statistical—defined by the equations

\[ \rho(k,l^*)=|c(k,l^*)|^2; \quad \rho(r,l^*)=|d(r,l^*)|^2 \quad \text{etc.} \]

for continuous \(k,r\), and

\[ w(k_n,l^*)=|c(k_n,l^*)|^2 \quad \text{and} \quad w(s_n,l^*)=|d(s_n,l^*)|^2 \]

for discrete \(k\) and \(s\).

In other words, we express probabilities as squared moduli of certain auxiliary numbers (functions), which remain arbitrary. The conditions \((N)\) take the form

\[ \int |c(k,l)|^2\,dk=1 \]

or

\[ \sum_n |c(k_n,l)|^2=1, \]

i.e. we see that our functions \(c(k,l)\) are precisely functions of the required class.

That is why it is expedient to carry out the investigation not with the probabilities themselves, but with statistical functions having an integrable squared modulus.

Let us now take some closed system of functions of the variable \(n\)

\[ \varphi(n_1,l),\ \varphi(n_2,l),\ldots,\varphi(n_k,l),\ldots \]

and expand all our statistical functions in it. If the index \(n_k\) is discontinuous, then our expansions have the form

\[ c(k_s,l^*)=\sum_{\mu} c_{n_\mu}(k_s)\,\varphi_{n_\mu}(l^*) \qquad \text{etc.,} \]

where

\[ \varphi_{n_\mu}(l^*)=\varphi(n_\mu,l^*), \qquad c_{n_\mu}(k_s)=c(n_\mu,k_s). \]

If, however, \(n\) is continuous, then the expansions have the form

\[ c(k_s,l^*)=\int c(k_s,n)\varphi(n,l^*)\,dn \qquad \text{etc.,} \]

Let us now substitute these expressions into the averages. Then we shall express them through the new independent variable \(n\). For simplicity let us suppose that \(n\) is continuous, while the quantities \(K\), \(R\), etc. have a discrete infinite aggregate of values. After substitution we obtain, for example, for the mean value of the quantity \(K\):

\[ \text{av. val. } K=\sum_{\mu=1}^{\infty} k_\mu |c(k_\mu,l^*)|^2= \]

\[ =\sum_{\mu=1}^{\infty} k_\mu \int c(k_\mu,n)\varphi(n,l^*)\,dn \int \overline{c(k_\mu,n)}\,\varphi(n,l^*)\,dn = \]

\[ =\int dn\,\overline{\varphi(n,l^*)}\int dn'\sum_{\mu=1}^{\infty} k_\mu c(k_\mu,n)\overline{c(k_\mu,n')}\,\varphi(n',l)\,dn'. \]

Put

\[ K(n,n')=\sum_{\mu=1}^{\infty} k_\mu c(k_\mu,n)\overline{c(k_\mu,n')}. \]

Then we can write the mean value of \(K\) in the form

\[ \text{av. val. } K=\int dn\,\overline{\varphi(n,l^*)}\int K(n,n')\varphi(n',l^*)\,dn'. \]

PRINCIPLES OF QUANTUM MECHANICS

We note that the expression

\[ \int K(n,n')\,\varphi(n',l^*)\,dn' \equiv K\varphi(n,l^*) \]

is a linear integral operator having kernel \(k(n,n')\) and applied to the function \(\varphi(n',l^*)\). Thus we shall write our mean value in the form

\[ \text{mean val. } K=\int \overline{\varphi}(n,l^*)\,K\varphi(n,l^*)\,dn', \]

where \(K\) is the symbol of the linear operator introduced by us above.

Expressions of the type obtained are called Hermitian forms and are denoted in the following way

\[ \int \overline{\varphi}(n,l^*)\,K\varphi(n,l^*)\,dn=(K\varphi,\varphi). \]

Thus,

\[ \text{mean val. } K=(K\varphi,\varphi). \]

Carrying out the same transformation, we obtain

\[ \text{mean val. } R=(R\varphi,\varphi), \]

i.e.

\[ \text{mean val. } R=\int \overline{\varphi}(n,l^*)\,R\varphi(n,l^*)\,dn, \]

where \(\varphi(n,l^*)\) is the same function as in the mean value for \(k\), and \(R\) is a linear integral operator applied to \(\varphi(n,l^*)\)

\[ R\varphi(n,l^*)=\int R(n,n')\,\varphi(n',l^*)\,dn', \]

having the kernel

\[ R(n,n')=\sum_{\mu=1}^{\infty} l_\mu d(l_\mu,n)\,\overline{d(l_\mu,n')}, \]

where \(d(l_\mu,n)\) are the coefficients in the expansion of the statistical function determining the probability \(w(r,l^*)\) in the closed system of functions chosen by us.

In the same way we shall compute all the other mean values as well, and in all cases we shall obtain them in the form of Hermitian forms, construct-

arising from one and the same function \(\varphi(n,l^*)\) and different operators:

\[ \text{mean value } K=(K\varphi,\varphi), \]

\[ \text{mean value } R=(R\varphi,\varphi), \]

\[ \text{mean value } T=(T\varphi,\varphi) \]

and so on.

We have obtained linear integral operators \(K,\ R,\ T\), entering into Hermitian forms, for the reason that we assumed that our quantities \(K,\ R,\ T\) have a discrete infinite set of possible values. If, however, one is dealing with quantities having a continuous, or mixed, set of possible values, then it is not difficult to generalize the results obtained. For this it is only necessary to make use of the concept of the Stieltjes integral. As a result we obtain linear Hermitian (i.e., having real characteristic numbers) operators. For the linear Hermitian operator \(K\) the condition

\[ (K\varphi,\varphi)=(\varphi,K\varphi). \]

holds.

This generalization is not of a fundamental, but of a purely mathematical character, and we shall not dwell on it.

Up to now both the operators \(K,\ R,\ S\), and the statistical function \(\varphi(n,l^*)\) serving for the construction of the means, remain completely undetermined. This is only a manner of expression. However, we shall be able to establish a connection between the operators and the function \(\varphi(n,l^*)\) by considering mean square errors. We noted above that if a quantity has a definite, true value for the statistical ensemble under consideration, then for this quantity the square of the mean is equal to the mean square. Let us express this condition, using the form for means established by us. Suppose that some quantity \(K\) has for our statistical ensemble the assigned value \(K_n\). This means that it must be

\[ (\text{mean value } K)^2=\text{mean value }(K^2) \]

But

\[ (\text{mean value } K)^2=(K\varphi,\varphi), \]

where \(\varphi\) is the function characterizing the statistical ensemble under consideration.

We can easily find the mean value of \(K^2\)

\[ \text{mean value }(K^2)=\sum_{\mu=1}^{\infty} k_\mu^2 w(k_\mu,l^*), \]

namely: we obtain

\[ \text{mean value }(K^2)=(K^2\varphi,\varphi), \]

where \(K^2\) is a linear integral operator having the kernel

\[ K(n,n')=\sum_{\mu=1}^{\infty} k_\mu^2 c(k_\mu,n)c(k_\mu,n'), \]

i.e. representing the operation \(K\) repeated twice. Knowing the expressions for the mean value \(K^2\) (mean value \(k^2\)), we write the condition

\[ \Delta K=0 \]

in the form

\[ \text{mean value }(K-\text{mean value }K)^2=0, \]

i.e.

\[ [\varphi,\{K-(K\varphi,\varphi)\}^2\varphi]=0. \]

Putting \([K-(K\varphi,\varphi)]=f\) and using the Hermiticity of the operator \(\{K-(K\varphi,\varphi)\}\), we obtain

\[ (f,f)=0, \]

i.e.

\[ f=0 \]

or

\[ [K-(K\varphi,\varphi)]\varphi=0 \]

or, writing it out explicitly,

\[ \int K(n,n')\varphi(n',l^*)\,dn'=(K\varphi,\varphi)\varphi(n',l^*). \]

We obtain, consequently, an integral equation for the function \(\varphi(n,l^*)\), which characterizes our statistical ensemble. But in our case the mean coincides with the value \(k_\mu\); therefore we can rewrite the last equation in the following form:

\[ \int K(n,n')\varphi(n';l^*,k_\mu)\,dn'=k_\mu\varphi(n,l^*,k_\mu), \]

where, alongside \(l^*\), we also write \(k_\mu\), since the quantity having a prescribed value may be chosen as the one that characterizes the members of the ensemble (i.e. plays the role of the quantity \(l^*\)).

If we knew the kernel of this integral equation, we could compute \(\varphi(n,l^*,k_\mu)\), and also \(k_\mu\). In fact, \(\varphi(n,l^*,k_\mu)\) are the fundamental functions of this integral equation, and \(k_\mu\) its characteristic numbers.

Thus our method acquires a complete character.

where the function \((\varphi r_\mu, n)\) is the same as in the preceding equation. This means that the equations have common, nonzero solutions. Hence it follows that the operators \(K\) and \(R\) commute with one another, i.e.

\[ KR(f)=RK(f) \]

for arbitrary \(f\). Indeed,

\[ KR\varphi=K(r_\mu\varphi)=r_\mu(K\varphi)=RK\varphi. \]

The commutativity condition for the two operators \(R\) and \(K\) is often written in symbolic form as

\[ KR-RK=0. \]

Thus we obtain the classical case if the operators entering into the bilinear forms commute with one another. Conversely, the unavoidable scattering characteristic of quantum processes is expressed, as we see, by the noncommutativity of the operators of those quantities for which there is unavoidable scattering. Indeed, if

\[ KR-RK\ne 0, \]

then the above equations expressing the conditions \(\Delta K=0\) and \(\Delta R=0\) are incompatible, and therefore there is no such statistical ensemble for which both these conditions would be satisfied simultaneously.

We see further that, when the operators \(K\) and \(R\) of quantum quantities do not commute, Schwarz’s inequality can serve for a numerical estimate of the correlation existing between \(\Delta K\) and \(\Delta R\), since this inequality states that

\[ (\Delta K)_\psi \cdot (\Delta R)_\psi \geq \frac{1}{2}\,\mathrm{const}. \]

The constant on the right will serve as the measure of the correlation. It depends on the statistical function characterizing the statistical ensemble, and also on the operators \(K\) and \(L\). Therefore, in general it is different for different statistical ensembles and for different quantities. However, we may suppose—and this will be a fundamental supposition of quantum mechanics—that there exist such quantities whose statistical quantum correlation does not depend on the choice of the statistical ensemble, i.e. the constant appearing on the right-hand side of the inequality just written does not depend on \(\psi(x,l^*)\), which characterizes the statistical ensemble. This condition will be fulfilled only in the case where the expression

\[ \int \overline{\psi}(x,l^*)[KR-RK]\psi(x,l^*)\,dx \]

will not change when \(\psi(x,l^*)\) is replaced by other statistical functions. This is possible only in the case where the operator \((KR-RK)\) is equal to a constant, which can be taken outside the integral sign. Then

\[ \int \overline{\psi}(x,l^*)[KR-RK]\,\psi(x,l^*)\,dx = \text{const.}\int |\psi|^2\,dx = \text{const.} \]

Thus, the above “invariance condition” leads to the conclusion that the operators of quantities with invariant correlation must be connected by the relation

\[ i(KR-RK)=\hbar,\qquad i=\sqrt{-1}, \]

where \(\hbar\) is some real number. This parameter of statistical correlation is called the quantum of action. In order to distinguish the quantities connected pairwise by an invariant correlation, we introduce for them and for their operators a special notation and name. We shall call them canonical, conjugate momenta \(P_\mu\) and coordinates \(Q_\mu\), and shall write the preceding equation in the form

\[ i(P_\mu Q_\mu-Q_\mu P_\mu)=\hbar; \]

\(\mu=1,2,\ldots,f\), where \(f\) is the number of pairs of conjugate variables possible for the system under consideration (the number of degrees of freedom). This relation must be supplemented by the condition

\[ P_\mu Q_\nu-Q_\nu P_\mu=0,\qquad \nu\ne\mu \]

(independence of the degrees of freedom). For the errors \(\Delta P_\mu\) and \(\Delta Q_\mu\) we obtain

\[ \Delta P_\mu\cdot \Delta Q_\mu \ge \frac{\hbar}{2}, \]

and here we have the right to drop the index \(l^*\), since this relation, by the definition of the quantities \(P_\mu\) and \(Q_\mu\), does not depend on the choice of the statistical ensemble. The names—momenta and coordinates—are not accidental, since it can be shown that in passing to the classical limiting case we do indeed obtain from them the classical generalized momenta and coordinates.

The relation just written (the so-called Heisenberg inequality) establishes, in quantitative form, the statistical correlation between the scattering of momentum and coordinate. We see an extremely peculiar limitation on classical notions in the quantum domain. The smaller \(\Delta Q\), the larger \(\Delta P\), and conversely. However, from this same relation it is evident that the error \(\Delta Q\) can

to make arbitrarily small, i.e., the generalized coordinate can then be determined as accurately as desired. But in this case the uncertainty of the momentum increases. Conversely, we can always make \(\Delta P\) very small at the expense of increasing \(\Delta Q\).

\(P_\mu\) and \(Q_\mu\) are generalized coordinates and momenta. By specializing them so that they have the character of ordinary momentum and coordinate, we may regard their introduction as the use of the concept of the particle in the quantum domain. Indeed, the possibility of exactly determining the coordinate \(Q_\mu\) under all conditions is an expression of the possibility of localizing a particle at a definite place.

On the other hand, the possibility of exactly determining the momentum \(P_\mu\) expresses the applicability of energetic representations also in the quantum domain, i.e., the possibility of formulating the laws of conservation of momentum and energy for processes in which quantum bodies also participate. However, as we see, this concept of a “quantum particle” differs fundamentally from the old concept of a particle, since the latter is based on the joint use both of the concept of momentum and energy and of the concept of coordinate. Let us note that only all these concepts, taken together, lead to the representation of mechanical motion, which in quantum mechanics is possible only as a statistically reworked concept. If the role of the quantum of action is neglected, then there is no obstacle to carrying through the classical description completely.

Let us now consider in more detail the principal properties of quantum particles. It is most natural to use, for describing the behavior of quantum particles, generalized momenta and coordinates \(P_\mu\) and \(Q_\mu\), or more precisely their means, defined for different moments of time. Of course, various functions of \(P_\mu, Q_\mu\) must also be used.

The laws describing the behavior of quantum particles are determined by the fact that the condition of statistical correlation between \(P_\mu\) and \(Q_\mu\), expressed by the equation

\[ i(P_\mu Q_\mu - Q_\mu P_\mu)=\hbar, \]

is invariant, i.e., does not change when the state of the quantum system changes. Thus we can express the change of the state of some quantum system as a change of the statistical function \(\varphi\), or of the operators \(P_\mu, Q_\mu\), occurring while the condition of invariance of the mean is observed,

\[ [i(PQ-QP)\varphi,\varphi]=\hbar \]

(invariance of action). The change of the state of a system of quantum particles is expressed by the change of the means of various quantities \(K, L\)

\[ (K\varphi,\varphi), (L\varphi,\varphi),\ldots \]

The properties of the quantum system, however, are expressed by relations existing between averages.

Let us find the general properties of quantum systems. A change of state may be described by a change of the statistical function \(\varphi\), which passes into a new \(\varphi^*\). In this case some average \((K\varphi,\varphi)\) changes into \((K\varphi^*,\varphi^*)\). If it is invariant with respect to this transformation, as, for example, the action, then

\[ (K\varphi,\varphi)=(K\varphi^*,\varphi^*). \]

We can express \(\varphi^*\) as some function of \(\varphi\), using the representation of an operator:

\[ \varphi^* = U(\varphi) \]

In this formula \(U\) is some as yet unknown operator transforming the old function \(\varphi\) into the new \(\varphi^*\).

Alongside the transformation just written there must also exist the inverse one, i.e. one carrying \(\varphi^*\) into \(\varphi\). We shall write this inverse transformation, again using the symbol of an operator, denoting the inverse operator by the symbol \(U^{-1}\), so that

\[ \varphi = U^{-1}(\varphi^*). \]

This symbolism is chosen so that

\[ UU^{-1}=U^{-1}U=1,\quad \text{i.e. } \varphi=\varphi \text{ or } \varphi^*=\varphi^*. \]

as it should be.

Let us now write the condition of invariance of some average, using the operator \(U\)

\[ (K\varphi,\varphi)=(KU\varphi,U\varphi) \]

But this expression is invariant also with respect to the inverse transformation \(U^{-1}\). Therefore

\[ (K\varphi,\varphi)=(KU\varphi,U\varphi)=(U^{-1}KU\varphi,\varphi)=(K^*\varphi,\varphi), \]

where

\[ K^* = U^{-1}KU. \]

We see that invariant averages can be transformed either by changing the statistical function by means of the operator \(U\), or by passing to a new operator \(K^*\), constructed from \(U\) and \(K\) according to the formula just written. Then the statistical function remains unchanged. If \(K\) is a Hermitian operator, i.e. such that

\[ (K\varphi,\varphi)=(\varphi,K\varphi), \]

then the operator \(U\), which leaves this mean invariant, is called unitary. It has the following property. For a linear operator \(A\) we can choose such a new operator*—let us denote it by \(A^{+}\)—that

\[ (A\varphi,\varphi)=(\varphi,A^{+}\varphi). \]

If

\[ A^{+}=A, \]

then our linear operator \(A\) will be Hermitian. A unitary operator is one for which

\[ U^{-1}=U^{+}\quad \text{and therefore}\quad U^{+}U=UU^{+}=1. \]

The preceding formulas may be rewritten as follows:

\[ K^{*}=U^{+}KU \]

and

\[ \varphi^{*}=U(\varphi),\quad \varphi=U^{+}(\varphi). \]

Recalling that the laws of behavior of quantum systems are determined by the invariance of the action, we see that these laws are expressed by the fact that the operators \(P_{\mu}\) and \(Q_{\mu}\) are transformed into

\[ P_{\mu}^{*}=U^{+}P_{\mu}U,\quad Q_{\mu}^{*}=U^{+}Q_{\mu}U,\quad U^{+}U=UU^{+}=1, \tag{I} \]

where \(U\) is one and the same linear unitary operator. If \(P_{\mu}\) and \(Q_{\mu}\) characterize the system at some instant of time \(t_{0}\), and \(P_{\mu}^{*}\) and \(Q_{\mu}^{*}\)—at the instant \(t\), then these formulas are precisely a generalization of Newtonian laws of motion. Indeed, one can conversely show that from these quantum laws the classical ones are obtained, if one passes to the classical limiting case, when \(\hbar\to0\).

The preceding formulas for operators should be understood as a formal notation of relations holding for means, i.e. if

\[ \text{mean value } P_{\mu}=(P_{\mu}\varphi,\varphi)\quad \text{and mean value } Q_{\mu}=(Q_{\mu}\varphi,\varphi), \]

then

\[ \text{mean value } P_{\mu}^{*}=(U^{+}P_{\mu}U\varphi,\varphi)\quad \text{and mean value } Q_{\mu}^{+}=(U^{+}P_{\mu}U\varphi,\varphi) \]

\[ \text{* It is called “conjugate,” “adjoint.”} \]

[Similarly, mean val. \((U^+U)=\) mean val. \((UU^+)=1\).]

We must make these formulas somewhat more precise, namely indicate explicitly for what prescribed quantity the statistical ensemble under consideration has been composed. We must write:

\[ (\text{mean val. } P_\mu)_{\alpha=\alpha'}=(P_\mu\varphi_{\alpha'},\varphi_{\alpha'});\quad (\text{mean val. } Q_\mu)_{\alpha=\alpha'}=(Q_\mu\varphi_{\alpha'},\varphi_{\alpha'}), \]

\[ (\text{mean val. } Q_\mu^*)_{\beta=\beta'}=([U^+P_\mu U]\varphi_{\beta'},\varphi_{\beta'});\quad (\text{mean val. } Q^*)_{\beta=\beta'}= \]

\[ =(U^+QU\varphi_{\beta'},\varphi_{\beta'}), \]

where \(\alpha\) and \(\beta\) are certain quantities which, for all members of the statistical ensemble, have prescribed values \(\alpha'\) and \(\beta'\). In the first case the ensemble is characterized by the fact that \(\alpha=\alpha'\), and in the second case by the fact that \(\beta=\beta'\). Since for the present we are dealing only with four quantities—the old and the new generalized coordinates and momenta—we may choose as \(\alpha'\) and \(\beta'\) either the values of the momenta (old—\(p'\), new—\(p^{*'}\)), or the values of the coordinates (old—\(q'\), new—\(q^{*'}\)). Correspondingly, we obtain four types of transformations, effected by means of the operator \(U\), of four varieties:

\[ U(\alpha',\beta')=U(p',p^{*'}),\, U(q',q^{*'}),\, U(p',q^{*'}),\, U(q',p^{*'}). \]

In accordance with what has been said, we shall, for example, write

\[ \text{mean val. }(P_\mu)_{Q_\mu=q'}=(P_\mu\varphi_{q'},\varphi_{q'}) \]

and so on.

If the explicit form of the operators \(P_\mu\) and \(Q_\mu\) were known to us, then all the relations just noted would acquire a definite quantitative character and would express quite definite quantitative physical laws. However, this is in fact so, because the relation between \(P_\mu\) and \(Q_\mu\), which determines the invariant statistical correlation, determines the form of the operators \(P_\mu\) and \(Q_\mu\), up to a unitary transformation. To see this, we write the equation

\[ i(P_\mu Q_\mu-Q_\mu P_\mu)\varphi(\alpha,\beta)=\hbar\varphi(\alpha,\beta), \]

taking for \(\beta\) the value of the coordinate \(Q_\mu\), i.e. putting

\[ Q_\mu=q',\beta=q', \]

so that the equation will take the form

\[ i(P_\mu q' - q'P_\mu)\varphi(a,q')=\hbar \varphi(a,q'). \]

The problem is to determine the form of the operator \(P_\mu\) so that this equality is satisfied identically.

Recalling the formula for differentiation by parts

\[ d(f\cdot \varphi)=df\cdot \varphi+f\cdot d\varphi, \]

we see, rewriting it in the following way,

\[ d(f\cdot \varphi)-f\cdot d\varphi=df\cdot \varphi \]

and comparing with our equation, that, if \(f=q'\) and \(\varphi=\varphi(a,q')\), then

\[ P_\mu=\frac{\hbar}{i}\frac{\partial}{\partial Q_\mu} =\frac{\hbar}{i}\frac{\partial}{\partial q'}. \]

Indeed, then we obtain

\[ i\frac{\hbar}{i}\left[\frac{\partial}{\partial q'}(q'\varphi)-q'\frac{\partial \varphi}{\partial q'}\right]=\hbar \varphi, \]

which was what had to be proved. Knowing the explicit form of the operators \(P_\mu\) and \(Q_\mu\), we obtain the possibility of writing out also all averages in explicit form. For example,

\[ (\text{mean value }P)_{a'}=\int \overline{\varphi}(a',q')\,\frac{\hbar}{i}\frac{\partial}{\partial q'}\varphi(x',q')\,dq', \]

where the dependence of \(\mu\) on the other coordinates is not explicitly written out; integration must likewise be carried out over them.

Knowing the explicit form of the operators \(P_\mu\) and \(Q_\mu\), and being able to calculate their mean values for various statistical ensembles, we shall be able to calculate the values of the means of various functions of \(P_\mu\) and \(Q_\mu\), which we can use for the characterization of quantum systems.

Let us use the explicit form of the operators \(P_\mu\) and \(Q_\mu\) in order to write explicitly the basic quantum laws—(I). Let us note that, by multiplying them on the left by \(U\), we can give them the following form:

\[ UP_\mu^{*}=P_\mu U,\quad UQ_\mu^{*}=Q_\mu U;\quad \mu=1,2,\ldots f. \]

Noting that \(U\) is an operator taking \(\varphi'_{\alpha'}\) into \(\varphi'_{\beta'}\), let us say

\[ \varphi_{\beta'}=\int U(\beta',\alpha')\varphi_{\alpha'}\,d\alpha', \]

we can rewrite these “operator equations” in the following form:

\[ (UP_\mu^*\varphi_{\beta'},\varphi_{\alpha''})=(P_\mu U\varphi_{\beta'},\varphi_{\alpha''}), \]

and

\[ (UQ_\mu^*\varphi_{\beta'},\varphi_{\alpha''})=(Q_\mu U\varphi_{\beta'},\varphi_{\alpha''}). \]

Earlier we noted that for \(\alpha'\) and \(\beta'\) we may take either \(p^{*'}\), \(p'\), \(q^{*'}\), \(q'\), obtaining respectively four types of transformations. Making use now of the fact that we know the explicit form of the operators \(P_\mu, P_\mu^*\), we easily find the following types of transformations (we omit the indices \(\mu\) everywhere):

1)

\[ (UP^*\varphi_{q'},\varphi_{q^{*''}})=-\frac{\hbar}{i}\frac{\partial}{\partial q^{*''}}(U\varphi_{q'},\varphi_{q^{*''}}) \]

\[ (PU\varphi_{q'},\varphi_{q^{*''}})=\frac{\hbar}{i}\frac{\partial}{\partial q'}(U\varphi_{q'},\varphi_{q^{*''}});\quad U=U(q',q^{*''}). \]

2)

\[ U=U(p',q^{*}). \]

\[ (UP^*\varphi_{q'},\varphi_{q^{*''}})=-\frac{\hbar}{i}\frac{\partial}{\partial q^{*''}}(U\varphi_{p'},\varphi_{q^{*''}}) \]

Noting that

\[ Q^*=-\frac{\hbar}{i}\frac{\partial}{\partial p^{*'}} \]

and

\[ Q=\frac{\hbar}{i}\frac{\partial}{\partial p'}, \]

\[ (QU\varphi_{p'},\varphi_{q^{*''}})=-\frac{\hbar}{i}\frac{\partial}{\partial p'}(U\varphi_{p'},\varphi_{q^{*''}}), \]

3)

\[ U=U(p^{*'},q'). \]

This type is obtained from the preceding one by changing the sign and by replacing the starred quantities by unstarred ones and conversely.

4)

\[ U=U(p^{*'},p'') \]

is obtained from the first by changing the sign and replacing \(p\) by \(q\). Let us now suppose that we are considering the limiting case when

the role of the quantum of action may be neglected, i.e. when it is possible to replace the mean statistical values of all quantities by their true values. Then we can replace

\[ (UP^{*}\varphi_{q'}, \varphi_{q''}) \]

by

\[ p^{*'}(U\varphi_{q'}, \varphi_{q''}), \]

\[ (PU\varphi_{q'}, \varphi_{q''}) \]

by

\[ p'(U\varphi_{q'}, \varphi_{q''}) \]

and

\[ (Q^{*}U\varphi_{p'}, \varphi_{q^{*''}}) \]

by

\[ q^{*'}(U\varphi_{p'}, \varphi_{q^{*''}}) \]

and so on.

Accordingly, in the first case we obtain:

\[ p^{*'}\cdot (U\varphi_{q'}, \varphi_{q^{*''}}) = -\frac{\hbar}{i}\frac{\partial}{\partial q^{*''}} (U\varphi_{q'}, \varphi_{q^{*''}}), \]

\[ p'(U\varphi_{q'}, \varphi_{q^{*''}}) = \frac{\hbar}{i}\frac{\partial}{\partial q'} (U\varphi_{q'}, \varphi_{q^{*''}}), \]

these equations have the solution

\[ (U\varphi_{q'}, \varphi_{q^{*''}}) = e^{+\frac{i}{\hbar}S(q',q^{*''})} \]

and

\[ p^{*'}=-\frac{\partial S}{\partial q^{*''}},\qquad p'=\frac{\partial S}{\partial q'}, \]

which is nothing other than the classical formulas for canonical, tangent transformations of momenta and coordinates \((p',q')\) into momenta and coordinates \((p^{*'},q^{*''})\), effected by means of the action function \(S(q',q^{*''})\).

Similarly, for the second type of transformations, on passing to the classical limiting case we find

\[ (U\varphi_{p'}, \varphi_{q^{*''}}) = e^{+\frac{i}{\hbar}S(p',q^{*''})} \]

and

\[ p^{*'}=\frac{\partial S}{\partial q^{*''}},\quad -q'=\frac{\partial S}{\partial p'}, \]

i.e. again the classical formulas of canonical transformations. These formulas determine finite tangent transformations.

As is known, infinitesimal contact transformations, i.e., the change of the momenta \(P_\mu\) and coordinates \(Q_\mu\) during the time \(dt\), are determined in classical mechanics by Hamilton’s equations

\[ dP_\mu=-\frac{\partial H}{\partial q_\mu}\,dt,\qquad dq_\mu=\frac{\partial H}{\partial p_\mu}\,dt, \]

where

\[ H=H(p_1\ldots p_f;q_1\ldots q_f) \]

is the Hamiltonian function characterizing the classical system under consideration. These equations can also be obtained from quantum theory in the limiting classical case. In order to see this, we must consider infinitesimal canonical transformations of quantum variables.

Assuming, in the transformation of the operator \(A\),

\[ A^*=UAU^\dagger, \]

the operator \(U\) differs infinitesimally from unity, i.e.,

\[ U=1+iD\cdot\delta\tau, \]

where \(D\) is a Hermitian operator. Then

\[ U^\dagger=1-iD\delta\tau. \]

Substituting these expressions for \(U\) and \(U^\dagger\), we find

\[ A^*=A+i(DA-AD)\delta\tau. \]

and since

\[ \frac{dA}{dt}=\frac{A^*-A}{\delta t}, \]

then

\[ \frac{dA}{dt}=i(DA-DD). \]

This symbolic formula means that for the mean value \((A\varphi,\varphi)\) the relation

\[ \frac{d}{dt}(A\varphi,\varphi)=[i(DA-AD)\varphi,\varphi]. \]

Thus

\[ \frac{d}{dt}(P_\mu\varphi,\varphi)=[i(DP_\mu-P_\mu D)\varphi,\varphi] \]

and

\[ \frac{d}{dt}(Q_\mu \varphi,\varphi)=[\,i(DQ_\mu-Q_\mu D)\varphi,\varphi\,]. \]

Further, since \(D\) is a Hermitian operator, we have

\[ \delta(A\varphi,\varphi)=i\delta t[(DA\varphi,\varphi)-(AD\varphi,\varphi)] =i\delta t[(A\varphi,D\varphi)-(AD\varphi,\varphi)]. \]

On the other hand, we can express the change of the mean by means of the change of the statistical function \(\varphi\), namely:

\[ \delta A(t_0)=[(A\delta\varphi,\varphi)+(A\varphi,\delta\varphi)]. \]

Comparing this formula with the preceding one, we find that it must be

\[ -\delta\varphi=iD(\varphi)\,\delta t \]

or

\[ -\frac{1}{i}\frac{d\varphi}{dt}=D(\varphi), \]

where \(D\) is a linear operator defining an infinitesimal tangent transformation of the quantum canonical variables \(P_\mu\) and \(Q_\mu\). If the last equation is multiplied by \(\hbar\) and we set

\[ H=\hbar D, \]

then we obtain the equation known as the Schrödinger equation:

\[ -\frac{\hbar}{i}\frac{d\varphi}{dt}=H(\varphi). \]

In passing to the classical limiting case, it goes over into the Jacobi equation. This Schrödinger equation makes it possible to compute the statistical function \(\varphi\) characterizing the quantum system.

Returning to Hamilton’s equations, we note that, using the Schrödinger equation and the connection between the operators \(P_\mu\) and \(Q_\mu\), one can show that in the quantum domain the relations

\[ \frac{d}{dt}(P_\mu\varphi,\varphi)=-\left(\frac{\partial H}{\partial q_\mu}\varphi,\varphi\right) \]

and

\[ \frac{d}{dt}(Q_\mu\varphi,\varphi)=\left(\frac{\partial H}{\partial p_\mu}\varphi,\varphi\right);\quad(\mu=1,2\ldots f), \]

i.e.

\[ \frac{d}{dt}(\text{avg. val. } P_\mu)=-\,\text{avg. val. }\left(\frac{\partial H}{\partial q_\mu}\right) \]

and

\[ \frac{d}{dt}(\text{avg. val. } Q_\mu)=\text{avg. val. }\left(\frac{\partial H}{\partial p_\mu}\right). \]

Consequently, if we neglect the quantum of action, i.e. replace averages by values, then from these equations we obtain the classical Hamilton equations, and avg. val. \((\hbar D)\) goes over into the classical Hamiltonian function.

This result unequivocally shows that, if one ignores the scattering due to the quantum of action, one obtains the classical mechanical relations, i.e. quantum mechanics is the consistently carried-out accounting for the role of the quantum of action in mechanical processes. At the same time we see that the quantities \(P_\mu\) and \(Q_\mu\), defined by us in the quantum domain independently of any considerations by analogy with the classical theory, are indeed generalized momenta and generalized coordinates.

Even before the development of quantum mechanics, the central role of the concept of action was already clear; namely, investigations carried out at the end of the nineteenth century showed that the laws of mechanics can be formulated as laws of change of generalized momenta and coordinates that leave the action invariant. We now see the basis for this situation.

Using the method set forth, we could, step by step, develop all the mechanical quantum laws, and then, already neglecting the quantum of action, establish the classical relations as well. Here the fundamental difference between quantum laws and classical ones consists in the fact that they are fulfilled only for quite definite numerical values of various mechanical quantities. This circumstance is extremely essential and, as experiment shows, is an undoubted success of quantum theory.

In conclusion, let us say a few words about the limits of quantum mechanics. We saw at the beginning that quantum mechanics is based on an extremely peculiar formulation of all its problems. A clear conception of this peculiarity must always be kept in mind in order to see exactly what may be demanded of the quantum method. First of all, the participation of macroscopic bodies is always assumed in reactions, for which the quantum laws are formulated statistically, i.e. the laws of behavior of quantum particles. This participation of macroscopic “classical” bodies ensures the preservation of the old “classical” mechanical theories in the new domain, but formulated statistically, whereby the role of the quantum of action is taken into account. The theory, in this case, deals only with mechanical proces-

themselves, leaving electrodynamics aside, as well as the relativistic laws, i.e. the finite propagation of actions by means of the electromagnetic field. Therefore this theory, from the very beginning, could not be fully competent in all those questions in which the atomism of charge is essential and in which one cannot consider electrodynamic properties from a purely mechanical point of view. And indeed, numerous attempts to give a quantum theory of the electromagnetic field, using only quantum-mechanical principles, i.e. essentially classical mechanical principles with the successive inclusion of the quantum of action, were unsuccessful. Thus at the present time the connection between the atomism of charge (the constant \(e\)) and the atomism of action (the constant \(\hbar\)) is still completely obscure.

However, attention should be drawn to the following facts. The classical electron theory, founded by Lorentz, considers the electromagnetic field and, moreover, introduces the notion of electrically charged particles, whose properties are characterized by their charge. At the same time, if one tries to describe the properties of these particles by the same means by which the theory describes the electromagnetic field, this cannot be carried out consistently (the difficulties with the infinite self-energy and the structure of the electron). This shows the inconsistency of a theory which, alongside the electromagnetic field, operates with charge. Quite recently M. Born has succeeded in showing that a relativistically invariant modification is possible, more precisely, a generalization of the classical electron theory of Lorentz, in which the noted difficulty apparently disappears. This theory operates with nonlinear field equations, i.e. in it the sources of the field are not something external with respect to it, but the field is created by the field itself. As a consequence of this, in this theory there is, properly speaking, no constant charge \(e\), but there are constants characterizing the field, namely: the maximum values possible for the field (Lorentz’s theory assumes that infinitely large intensities of the electromagnetic field are possible).

On the other hand, we now have a theory of mechanical processes characterized by the quantum of action—\(\hbar\), which has been the subject of discussion here. This theory in its complete form is statistical, of a fundamentally nonclassical character. At the same time it does not consider electromagnetic processes and the atomism of charge.

It is very remarkable that, in the present state of theoretical physics, there are already indications of a new atomic constant—the so-called fine-structure constant

\[ \alpha=\frac{e^{2}}{\hbar c}, \]

which is an abstract number. Using this constant (let us call it the “Sommerfeld constant”), we can exclude

both the charge \(e\) and the quantum of action \(\hbar\). Thus, if such a modification of classical nonlinear electrodynamics were achieved that contained Sommerfeld’s constant, then such a theory, first, would be “classical,” and, second, would contain within itself, as two distinct limiting cases, electrodynamics (constant \(e\)) and quantum mechanics (constant \(\hbar\)).

Submission history

PRINCIPLES OF QUANTUM MECHANICS. I