CREATIVE AUTOBIOGRAPHY\*
A. Einstein
Submitted 1956 | SovietRxiv: ru-195601.77434 | Translated from Russian

Full Text

CREATIVE AUTOBIOGRAPHY*

Albert Einstein

Here I sit and write, in the 68th year of my life, something like my own obituary. I do this not only because I have been persuaded to; I myself think that it is a good thing to show my searching brethren what my own aspirations and quests look like in historical perspective. After some reflection, however, I felt how incomplete and imperfect such an attempt must be. For however brief and limited a working life may have been, however much errors and delusions may have predominated in it, it is nevertheless no easy task to select and set forth what deserves it; when a man is 67, he is not the same as he was at 50, 30, and 20. Every recollection is colored by what the person is now, and the present point of view may be misleading. This consideration might have deterred me. But, on the other hand, from one’s own experiences one can draw much that is inaccessible to the consciousness of another.

Even while still a rather precocious young man, I vividly realized the vanity of those hopes and strivings that drive most people through life, giving them no rest. Soon I also saw the cruelty of this race, which, incidentally, at that time was concealed more carefully than now by hypocrisy and fine words. Everyone was compelled by the existence of his stomach to take part in this race. This participation could lead to the satisfaction of the stomach, but by no means to the satisfaction of the whole person as a thinking and feeling being. The way out was indicated above all by religion, which is inculcated in all children by the traditional machinery of education. Thus, although I was the son of entirely nonreligious (Jewish) parents, I arrived at a deep religiosity, which, however, came to an abrupt end already at the age of 12. The reading of popular scientific books soon led me

*) A. Einstein, “Autobiographisches” (literally: “Something autobiographical”). First published in the volume Albert Einstein, Philosopher-Scientist, The Library of Living Philosophers, 1949, Illinois, USA. Translated by V. A. Fok and A. V. Lermontova.

to the conviction that much in the biblical stories could not be true. The consequence of this was downright fanatical freethinking, combined with the conclusion that youth was deliberately being deceived by the state; this was a shocking conclusion. Such experiences engendered distrust of authorities of every kind and a skeptical attitude toward the beliefs and convictions current in the social milieu that surrounded me at the time. This skepticism never left me again, although it later lost its sharpness, when I came to understand better the causal connection of phenomena.

It is clear to me that the religious paradise of youth, lost in this way, represented a first attempt to free myself from the fetters of the “merely personal,” from an existence dominated by desires, hopes, and primitive feelings.

Out there, beyond, was this great world, existing independently of us human beings and standing before us as an immense, eternal riddle, accessible, however—at least in part—to our perception and to our reason. The study of this world beckoned as a liberation, and I soon became convinced that many of those whom I had learned to esteem and respect had found their inner freedom and assurance by giving themselves wholly to this pursuit. The mental grasp, within the limits of the possibilities open to us, of this extra-personal world appeared to me—half consciously, half unconsciously—as the highest goal. Those who thought in this way, whether my contemporaries or people of the past, together with the views they had developed, were my sole and unchanging friends. The road to this paradise was not as comfortable and enticing as the road to the religious paradise, but it proved reliable, and I have never regretted taking it.

What I have just said is true only in a certain sense, just as a drawing consisting of a few strokes can convey a complex object, with its intricate small details, only in a limited sense. If a given personality especially values sharply delineated thought, then this aspect of its being may stand out more vividly than its other aspects and, to a greater degree, determine its spiritual world. It may then happen that, in retrospect, this personality will perceive systematic self-development where actual experiences followed one another in kaleidoscopic disorder. In fact, the diversity of external circumstances, combined with the fact that at any given moment one thinks only of one thing, introduces into the conscious life of every person a kind of atomic structure. In the development of a person of my cast of mind, the turning point is reached when the chief interest of life gradually breaks away from the momentary and the personal and concentrates more and more on the effort to grasp intellectually the nature of things. From this point of view, the schematic notes given above contain as much that is true as can generally be said in so few words.

CREATIVE AUTOBIOGRAPHY

What does it mean, in essence, “to think”? When, in perceiving sensations coming from the sense organs, memory-images arise in the imagination, this does not yet mean “to think.” When these images form a series, each member of which evokes the next, this too is not yet thinking. But when a definite image occurs in many such series, then, by virtue of its repetition, it begins to serve as an ordering element for such series, because it links series that in themselves are devoid of connection. Such an element becomes a tool, becomes a concept. It seems to me that the transition from free associations or “dreams” to thinking is characterized by the role—more or less dominant—which the “concept” plays in it. In itself it does not seem necessary that a concept be connected with a symbol acting upon the sense organs and reproducible (in words); but if this does take place, then the thought can be communicated to another person.

By what right, the reader will now ask, does this man operate so unceremoniously and in such an amateurish way with ideas in so problematic a field, without making even the slightest attempt to prove anything? My justification is this: all our thinking is of the same kind; it is a free play with concepts. The justification of this play lies in the possibility, attainable with its help, of surveying sensory perceptions. The concept of “truth” is not yet at all applicable to such a formation; this concept, in my opinion, can be introduced only when there is a conditional agreement concerning the elements and rules of the game.

For me there is no doubt that our thinking proceeds, for the most part, bypassing symbols (words) and, moreover, unconsciously. If it were otherwise, then why do we sometimes happen to be “surprised,” and quite spontaneously, by this or that perception (Erlebnis)? This “act of surprise” apparently occurs when a perception comes into conflict with a sufficiently established world of concepts within us. In those cases when such a conflict is experienced acutely and intensely, it in turn exerts a strong influence on our mental world. The development of this mental world is, in a certain sense, an overcoming of the feeling of surprise—an uninterrupted flight from the “surprising,” from the “miracle”*.

I experienced a miracle of this kind as a child of 4 or 5, when my father showed me a compass. The fact that this needle behaved in such a definite way did not at all fit the kind of phenomena that could find a place in my unconscious world of concepts (action through contact). I still remember—and it seems to me now that I remember—that this incident made a deep and lasting impression on me. Behind things there must be something else, deeply hidden...

* The words “miracle” and “surprise” have in German one and the same root, Wunder. (Translator’s note.)

A human being does not react in the same way to what he has seen since earliest childhood. He is not struck with wonder by the fall of bodies, by wind and rain; he is not amazed by the moon, nor by the fact that it does not fall, nor by the difference between the living and the nonliving.

At the age of 12 I experienced yet another wonder, of an entirely different kind: its source was a little book on Euclidean plane geometry that came into my hands at the beginning of the school year. There were propositions there, for example concerning the intersection of the three altitudes of a triangle at one point, which, though not self-evident in themselves, could nevertheless be proved with a certainty that seemed to exclude all doubt. This clarity and certainty made an indescribable impression on me. I was not troubled by the fact that the axioms had to be accepted without proof. In general, it was quite enough for me if, in my proofs, I could rely on propositions whose correctness seemed to me beyond dispute. I remember, for example, that the Pythagorean theorem had been shown to me by my uncle even before the holy little book on geometry came into my hands. With great difficulty I managed to “prove” this theorem by means of similar triangles; at the same time, however, it seemed to me “obvious” that the ratio of the sides of a right triangle must be completely determined by one of its acute angles. In general it seemed to me that one needed to prove only what was not “obvious” in this sense. And the objects with which geometry deals did not seem to me to be of any different nature from “visible and tangible” objects, i.e., objects perceived by the sense organs. This primitive understanding is, of course, based on the fact that the connection between geometrical concepts and observed objects (length—a rigid rod, etc.) was unconsciously taken into account. It is possible that this understanding lies at the foundation of the well-known Kantian formulation of the question concerning the possibility of a “synthetic judgment a priori.”

Although it looked as if, by means of pure reflection, one could obtain reliable information about observed objects, such a “miracle” was based on a mistake. Still, to anyone who experiences this “miracle” for the first time, the very fact that man is capable of attaining such a degree of reliability and purity in abstract thought as the Greeks first showed us in geometry seems astonishing.

Since I have allowed myself to interrupt the obituary begun with a sin halfway through, I shall no longer be embarrassed to express here, in a few phrases, my epistemological credo, although some of this has already been mentioned along the way above. These convictions of mine took shape slowly and developed much later; they do not correspond to the attitudes I had when I was younger.

I see, on the one hand, the totality of sensations coming from the sense organs; on the other hand, the totality of concepts and propositions recorded in books. The connections of concepts and propositions among themselves

A CREATIVE AUTOBIOGRAPHY

...is of a logical character; the task of logical thinking is reduced exclusively to establishing relations between concepts and propositions according to firm rules, which is what logic deals with. Concepts and propositions acquire “meaning” or “content” only thanks to their connection with sensations. The connection of the latter with the former is purely intuitive and is not in itself of a logical nature. Scientific “truth” differs from empty fantasizing only in the degree of reliability with which this connection, or intuitive correspondence, can be carried out, and in nothing else. A system of concepts is a creation of man, as are the rules of syntax that determine its structure. Although systems of concepts are, in themselves, logically entirely arbitrary, they are bound by the fact that, first, they must allow as reliable (intuitive) and complete a correspondence as possible with the totality of sensations; second, they must strive to get by with the smallest possible number of logically independent elements (basic concepts and axioms), i.e., such concepts for which no definitions are given, and such propositions for which no proofs are given.

A proposition is true if it is derived within a certain logical system according to the accepted rules. The content of truth in a system is determined by the reliability and completeness of its correspondence with the totality of sensations. More precisely, a proposition borrows its “truthfulness” from the store of truth contained in the system that contains it.

Remark on historical development. Hume clearly understood that certain concepts, for example the concept of causality, cannot be derived from experiential data by a logical path. Kant, convinced that one cannot do without certain concepts, regarded these concepts, in their accepted form, as necessary prerequisites of all thinking and distinguished them from concepts of empirical origin. I am convinced, however, that this distinction is erroneous and does not encompass the problem in a natural way. All concepts, even those closest to sensations and experiences, are from a logical point of view arbitrary posits, just as much as the concept of causality, which was the one primarily under discussion.

I now return to the obituary. At the age of 12–16 I became acquainted with the elements of mathematics, including the foundations of differential and integral calculus. In this, fortunately for me, I came across books in which not too much attention was paid to logical rigor, but everywhere the main idea was clearly brought out. The whole occupation was truly absorbing; it contained flights of thought whose force of impression was not inferior to the “miracle” of elementary geometry—the basic idea of analytic geometry, infinite series, the concept of the differential and the integral. I was also fortunate enough to gain an idea of the most important results and methods of the natural sciences from a very good popular edition, in which the exposition was almost everywhere limited to the qualitative aspect.

question (Bernstein’s popular-science books for the people—a work in 5–6 volumes); I read these books without drawing breath. By the time, at the age of 17, I entered the Zurich Polytechnic as a student of physics and mathematics, I was already somewhat familiar with theoretical physics.

There I had excellent teachers (for example, Hurwitz, Minkowski), so that, properly speaking, I could have obtained a solid mathematical education. I, however, spent most of my time working in the physics laboratory, carried away by direct contact with experiment. I used the rest of my time chiefly to study at home the works of Kirchhoff, Helmholtz, Hertz, and so on. The reason that I neglected mathematics to a certain extent was not only the predominance of scientific interests over mathematical ones, but also the following peculiar feeling. I saw that mathematics was divided into a multitude of special fields, and that each of them could occupy the whole short life allotted to us. And I saw myself in the position of Buridan’s ass, unable to decide which armful of hay to take. The point was evidently that my intuition in the field of mathematics was not sufficiently strong for me to distinguish confidently what was fundamental and important from the rest of erudition, without which one could still get by. Moreover, my interest in the investigation of nature was undoubtedly stronger; as a student it was not yet clear to me that access to the deeper fundamental problems in physics requires the subtlest mathematical methods. This became clear to me only gradually, after many years of independent scientific work. Of course, physics too was divided into special fields, and each of them could absorb a short working life, without satisfying the thirst for deeper knowledge. The enormous quantity of insufficiently connected empirical facts had a depressing effect here as well. But here I soon learned to ferret out what could lead into depth and to discard everything else, everything that burdens the mind and distracts from the essential. There was, however, the hitch that for the examination one had to cram into oneself—whether one wanted to or not—all this wisdom. This compulsion frightened me so much that for a whole year after passing the final examination any reflection on scientific problems was poisoned for me. In this connection I must say that in Switzerland we suffered from such compulsion, which stifles genuine scientific work, considerably less than students suffer in many other places. There were only two examinations in all; otherwise one could do more or less whatever one wanted. It was especially good for someone who, like me, had a friend who conscientiously attended all the lectures and diligently worked through their content. This gave freedom in the choice of occupation right up to several months before the examination—a freedom of which I made extensive use; and the bad conscience connected with it I pri-

CREATIVE AUTOBIOGRAPHY

accepted it as an inevitable, and moreover a considerably lesser, evil. In essence it is almost a miracle that contemporary methods of instruction have not yet completely stifled sacred curiosity, for this tender plant requires, along with encouragement, above all freedom—without it, it inevitably perishes. It is a great mistake to think that a sense of duty and compulsion can help one find joy in looking and searching. It seems to me that even a healthy predatory animal would lose its greed for food if one managed, by means of a whip, to force it to eat continuously, even when it was not hungry, and especially if the food compulsorily offered had not been chosen by it.

Let us now turn to physics, as it appeared at that time. Although in certain areas it flourished, in matters of principle a dogmatic stagnation prevailed. In the beginning (if such a thing ever existed), God created Newton’s laws of motion together with the necessary masses and forces. That is all; everything else must be obtained deductively, as the result of developing the proper mathematical methods. Relying on this foundation and, in particular, applying partial differential equations, the nineteenth century yielded so much that it should astonish any thinking person. Newton was probably the first to demonstrate, in his theory of the propagation of sound, the fruitfulness of the method of differential equations in partial derivatives. Euler had already created the foundations of hydrodynamics. But the more detailed construction of the mechanics of discrete masses as the foundation of all physics was an achievement of the nineteenth century. What made the greatest impression on a student was not so much the construction of the apparatus of mechanics itself and the solution of complex problems, as the achievements of mechanics in domains which at first sight seemed entirely unrelated to it: the mechanical theory of light, which regarded light as wave motion of a quasi-rigid elastic ether, and, above all, the kinetic theory of gases. Here one should mention the independence of the heat capacity of monatomic gases from atomic weight, the derivation of the equation of state of a gas and its connection with heat capacity, and, most important, the numerical relationship among the viscosity, thermal conductivity, and diffusion of gases, which also yielded the absolute dimensions of the atom. These results served at the same time as a confirmation of mechanics as the foundation of physics and as a confirmation of the atomic hypothesis, which by then had already become firmly established in chemistry. In chemistry, however, only the ratios of atomic masses played a role, not their absolute magnitudes; therefore atomic theory could be regarded there rather as a vivid analogy than as knowledge of the actual structure of matter. Independently of this, the deepest interest was also aroused by the fact that the statistical theory of classical mechanics was able to derive the fundamental laws of thermodynamics; in essence this had already been done by Boltzmann.

It is therefore not surprising that the physicists of the past century saw in classical mechanics an unshakable foundation for all physics.

and even for all of natural science; they ceaselessly tried to base themselves on mechanics and on Maxwell’s theory of electromagnetism, which was slowly making its way. Maxwell and Hertz, in their conscious thinking, also regarded mechanics as the reliable foundation of physics, although in historical perspective it must be acknowledged that it was precisely they who undermined confidence in mechanics as the foundation of the foundations of all physical thought. Ernst Mach, in his history of mechanics, shook this dogmatic faith; for me as a student, this book had a profound influence precisely in this respect. I see Mach’s true greatness in his incorruptible skepticism and independence; in my younger years Mach’s gnoseological attitude also made a strong impression on me, though today it seems to me untenable in essential points. Namely, he did not sufficiently emphasize the constructive and speculative character of all thought, especially scientific thought. As a result, he condemned theory precisely in those places where its constructive-speculative character appears undisguised, for example in kinetic theory.

Before undertaking a critique of mechanics as the foundation of physics, it is first necessary to state several general propositions about the points of view, or criteria, from which physical theories may in general be criticized. The first criterion is obvious: a theory must not contradict the data of experience. But however obvious this requirement may seem in itself, its application turns out to be quite subtle. The point is that often, if not always, one can preserve a given general theoretical basis if only one adapts it to the facts by means of more or less artificial auxiliary assumptions. In any case, in this first criterion what is involved is the testing of the theoretical basis against the available empirical material.

In the second criterion, what is involved is not the relation to empirical material, but the premises of the theory itself, what might be called, briefly—though not entirely clearly—the “naturalness” or “logical simplicity” of the premises (the basic concepts and the basic relations between them). This criterion, whose exact formulation presents great difficulties, has always played a large role in choosing between theories and in evaluating them. What is at issue here is not simply some enumeration of logically independent premises (if such a thing is even possible unambiguously), but a kind of weighing and comparing of incommensurable qualities. Furthermore, of two theories with equally “simple” fundamental propositions, preference should be given to the one that more strongly restricts the possible a priori qualities of systems (i.e., contains the most definite assertions). With regard to the “domain of applicability” of theories, I can say nothing here, since we are considering only such theories whose subject is the entire totality of physical phenomena.

The second criterion may be briefly characterized as the criterion of the “inner perfection” of a theory, whereas the first concerns

to its “external justification.” To the “inner perfection” of a theory I also include the following: a theory appears to us more valuable when it is not a logically arbitrary choice among approximately equivalent and analogously constructed theories.

I shall not justify the insufficient definiteness of my statements in the last two paragraphs by a lack of space allotted to me in print; I openly admit that I cannot at once, and perhaps am not at all able to, replace these hints with precise definitions. I believe, however, that a more precise formulation is possible. In any case, we see that among the “augurs” there is, for the most part, complete agreement in judgments about the “inner perfection” of theories and, in particular, about the degree of their “external justification.”

We now turn to a critique of mechanics as the foundation of physics.

From the point of view of the first criterion (verification by experience), the inclusion of wave optics in the mechanical picture of the world should have aroused serious doubts. If light is to be regarded as wave motion in an elastic body (in the ether), then this body must be an all-pervading medium. Because of the transverse nature of light waves, this medium must be, in the main, similar to a solid body; yet it must be incompressible, so that longitudinal waves should not exist. This ether had to lead, alongside ordinary matter, a ghostly existence, since it seemed to offer no resistance whatever to the motion of “ponderable” bodies. In order to explain the refractive indices of transparent bodies, as well as the processes of emission and absorption of light, it would have been necessary to assume intricate interactions between two kinds of matter; not only was this not done, but no one even seriously attempted it.

Furthermore, electromagnetic forces compelled the introduction of electric masses, which, although they possessed no noticeable inertia, nevertheless acted upon one another; in contrast to the force of gravitation, this interaction had a polar character.

The cause which, in the end, induced physicists—after long hesitation—to abandon the belief in the possibility of constructing all physics on the basis of Newtonian mechanics was the electrodynamics of Faraday–Maxwell. This theory, together with Hertz’s experiments confirming it, showed that there exist electromagnetic processes essentially detached from all ponderable matter, namely waves that are oscillations of electromagnetic “fields” in empty space. Whoever wished to preserve mechanics as the foundation of physics had to give a mechanical interpretation of Maxwell’s equations. This was undertaken with the utmost zeal, but quite fruitlessly, while the equations themselves were revealing their fruitfulness to an ever greater extent. People became accustomed to operating with these fields as with independent realities, without going into

in their mechanical nature. Thus, almost imperceptibly, the view of mechanics as the foundation of physics was abandoned; this happened because the adaptation of mechanics to experimental facts proved hopeless. Since then there have existed two systems of elementary concepts: on the one hand, material points interacting at a distance, and on the other, the continuous field. This state of physics, in which its unified foundation is absent, is as it were transitional; despite all its unsatisfactoriness, it has by no means yet been overcome — — —.

Now on the critique of mechanics as the foundation of physics from the standpoint of the second, “internal,” criterion. In the present state of science, when the mechanical foundation has already been abandoned, criticism of this kind can have only methodological interest. It is, however, very useful as an example of the kind of argumentation which in the future, when choosing between theories, must play a role that is the greater the farther their basic concepts and axioms are removed from what is directly observed; under such circumstances the comparison of the conclusions of a theory with experience becomes ever more complex and difficult. Here one should mention first of all a consideration of Mach’s which, incidentally, was already quite clear to Newton as well (the bucket experiment). From the standpoint of a purely geometrical description, all “rigid” systems of reference are logically equal. However, the equations of mechanics (and already Newton’s first law) are valid only in some of these systems of reference, namely in “inertial” systems, which form a special class. In this connection the character of the reference system as a material body proves to be inessential. The necessity of taking precisely an inertial system of reference must therefore be conditioned by something lying outside those objects (masses, distances) of which the theory speaks. As such a determining circumstance Newton introduced “absolute space” as a certain omnipresent active participant in all mechanical processes. By “absolute” Newton evidently means “not subject to the influence of masses and their motions.” The situation is aggravated by the fact that the existence is assumed of an infinite multitude of inertial systems moving uniformly and without rotation relative to one another, and these systems of reference are assumed to be singled out among all the other rigid systems of reference.

In Mach’s opinion, in a truly rational theory inertia, like the other Newtonian forces, must originate from the interaction of masses. For a long time I considered this opinion in principle correct. It implicitly assumes, however, that the theory on which everything is based must belong to the same general type as Newtonian mechanics: its basic concepts must be masses and the interactions between them. Meanwhile, it is not difficult to see that such an attempted solution does not accord with the spirit of field theory.

Nevertheless, Mach’s critique is in itself quite well founded. This is especially clearly seen from the following analogy. Let us imagine people constructing mechanics; suppose, moreover, that they know only a small part of the earth’s surface and have no possibility of seeing the stars. They will be inclined to ascribe special physical properties to the vertical measurement of space (the direction of the acceleration in falling). On this basis they will come to the conclusion that the earth’s surface is predominantly horizontal. Suppose that they are not amenable to the consideration that space, in a geometrical respect, is isotropic and that therefore the fundamental physical laws cannot be constructed in such a way that the existence of a privileged direction would follow from them; these people will probably be inclined to assert (like Newton) that the vertical is absolute, that “experience shows this,” and that this must be reckoned with. The singling out of verticals before all other directions is entirely analogous to the singling out of inertial systems before other rigid coordinate systems.

Let us now give further arguments, which also concern the question of the internal simplicity and naturalness of mechanics. If one accepts without critical doubts the concepts of space (including geometry) and time, then there are as yet no grounds for objecting to the introduction of action-at-a-distance forces as fundamental concepts, although the concept of action at a distance is not in accord with those ideas which people form on the basis of crude everyday experience. On the other hand, there is another consideration by virtue of which the understanding of mechanics as the foundation of physics appears to us primitive.

Basically there are two laws:

1) The law of motion.
2) The expression for the force (or for the potential energy).

The law of motion is exact, but it is empty until an expression for the force is given. The writing of this expression is, however, connected with wide arbitrariness, especially if one discards the requirement—not self-evident in itself—that the forces depend only on the coordinates themselves (and not, for example, on their time derivatives).

Within such a theory it is also arbitrary that the action of gravitational forces (and electrical forces) issuing from a single point is determined by a potential function \(\left(\frac{1}{r}\right)\). An additional remark: it has long been known that this function is a centrally symmetric solution of the simplest (rotation-invariant) differential equation \(\Delta \varphi = 0\); it would be natural to regard this as a sign that this function should be determined from some spatial law, whereby arbitrariness in the choice of the law for the forces would be eliminated. Strictly speaking, this is the first result that could have suggested the idea of departing from the theory of action at a distance. However, development in this

directions—begun by Faraday, Maxwell, and Hertz—came only later, under the pressure of experimental facts.

I would also like to point out the internal asymmetry of the theory, manifested in the circumstance that the inertial mass entering into the law of motion also enters into the expression for gravitational forces, but not into the expression for other forces. Finally, I would like to point out that the division of energy into two essentially different parts—kinetic and potential energy—must be regarded as something unnatural; Hertz found this so inconvenient that in his last work he even attempted to free mechanics from the concept of potential energy (i.e., force) — — —.

Enough of this. Forgive me, Newton; you found the only path possible in your time for a man of the greatest scientific creative power and strength of thought. The concepts created by you even now remain leading ones in our physical thinking, although we now know that, if we are to strive for a deeper understanding of interrelations, we must replace these concepts by others, farther removed from the sphere of direct experience.

“And is this an obituary?” the surprised reader may ask. In essence—yes, I would like to reply. For the main thing in the life of a man of my sort consists in what he thinks and how he thinks, and not in what he does or experiences. Therefore, in an obituary one may in the main confine oneself to reporting those thoughts that played a significant role in my aspirations. A theory makes the greater impression the simpler its premises are, the more diverse the objects it connects, and the broader the area of its application. Hence the deep impression that classical thermodynamics made on me. It is the only physical theory of general content of which I am convinced that, within the limits of applicability of its basic concepts, it will never be refuted (for the special information of principled skeptics).

The most fascinating subject during my years of study was Maxwell’s theory. The transition from forces acting at a distance to fields as fundamental quantities made this theory revolutionary. The fact that optics found its place in the theory of electromagnetism, which established a connection between the speed of light and the absolute electric and magnetic system of units, and also connected the refractive index with the dielectric constant and led to a qualitative relation between the coefficient of reflection and the metallic conductivity of a body—all this was for me like a revelation. In addition to the transition to field theory, i.e., to the expression of elementary laws by means of differential equations, Maxwell needed only one hypothetical step—the introduction of the electric displacement current in the vacuum and in dielectrics, with its magnetic action—

... by them; this innovation was almost dictated by the properties of the differential equations themselves. In this connection I cannot refrain from noting the remarkable inner similarity between the combination Faraday—Maxwell and the combination Galilei—Newton. The first in each pair intuitively grasped the relations, while the second formulated them exactly and applied them quantitatively.

Penetration into the essence of electromagnetic theory was at that time made difficult by the following peculiar circumstance. Electric and magnetic “field strengths” were regarded, on a par with “displacements,” as primary quantities, while empty space was considered a special case of a dielectric. The bearer of the field was thought to be matter (substance), and not space. And this implied that the bearer of the field has the property of possessing velocity, which, of course, had also to be true for “empty space” (ether). Hertz’s electrodynamics of moving bodies is wholly based on this fundamental premise.

The great merit of H. A. Lorentz was that he brought about a revolution here, and in the most convincing way. According to Lorentz, in principle there exists only the field in empty space. Matter, which is assumed to be atomistic, is the sole bearer of charges; between material particles there is empty space—the bearer of the electromagnetic field, which is produced by the positions and velocities of the point charges seated on the particles. Dielectric properties, conductivity, and so forth are determined exclusively by the character of the mechanical bonds between the particles of which bodies are composed. The charged particles create a field, which in turn acts on the charges of the particles. The corresponding forces determine the motion of the particles according to Newton’s laws. If this is compared with Newton’s system, the change consists in the following: forces acting at a distance are replaced by a field, which also describes radiation. Gravitation is for the most part not taken into account because of its relative smallness; however, it can be included by “enriching” the structure of the field and by a corresponding extension of Maxwell’s equations for the field. A physicist of the present generation will regard the point of view won by Lorentz as the only possible one, whereas at that time it was an astonishingly bold step, without which further development would have been impossible.

If one looks critically at this phase of the theory’s development, what first strikes the eye is its duality, consisting in the fact that the material point in the Newtonian sense and the field, as a continuum, are used side by side as elementary concepts. Kinetic energy and field energy appear as fundamentally different things. This seems all the more unsatisfactory since, according to Maxwell’s theory, the magnetic field of a moving electric charge represented inertia. Why, then, not all inertia? Then there would be only field energy, and the particle would be merely a region

especially high density of this field energy. Then one might have hoped that the concept of a material point, together with the equations of motion of the particle, could be derived from the field equations—and the troublesome duality would be eliminated.

H. A. Lorentz understood this perfectly well. But Maxwell’s equations did not make it possible to establish the conditions of equilibrium of electricity constituting a single particle. Only other, nonlinear field equations might perhaps have done this. However, there was as yet no method that would make it possible to find such equations without lapsing into the most arbitrary arbitrariness. In any case, one could hope to find a new, reliable foundation for all of physics by advancing step by step along the path so successfully marked out by Faraday and Maxwell.

Thus, the revolution begun by the introduction of the field could by no means be regarded as complete. It so happened that, on the threshold of two centuries, independently of this upheaval, yet another crisis of fundamental concepts broke out, whose importance suddenly came to people’s awareness thanks to Max Planck’s investigations of thermal radiation (1900). The history of this crisis is all the more remarkable because it was not influenced, at least in its initial stage, by any of the many discoveries of an experimental nature that followed one another.

On the basis of thermodynamic considerations Kirchhoff came to the conclusion that the energy density and the spectral composition of the radiation enclosed in a cavity with opaque walls at temperature \(T\) do not depend on the nature of these walls. This meant that the density \(\rho\) of monochromatic radiation is a universal function of the frequency \(\nu\) and the absolute temperature \(T\). Thus there arose the interesting problem of determining this function \(\rho(\nu, T)\). What could be obtained theoretically concerning this function? According to Maxwell’s theory, radiation must exert on the walls a pressure determined by the total energy density. From this Boltzmann derived, by a purely thermodynamic route, that the total energy density of radiation \(\left(\int \rho\, d\nu\right)\) is proportional to \(T^4\). Thereby he found a theoretical justification for the empirical regularity already found earlier by Stefan; in other words, he connected it with the foundations of Maxwell’s theory. After this, W. Wien, with the aid of an ingenious thermodynamic argument in which Maxwell’s theory was also used, found that the universal function \(\rho\) of the two variables \(\nu\) and \(T\) must have the form

\[ \rho \simeq \nu^3 f\left(\frac{\nu}{T}\right), \]

where \(f(\nu/T)\) denotes a universal function of the single variable \(\nu/T\). It was clear that the theoretical determination of this universal function \(f\) was of fundamental importance—this was precisely the problem confronting Planck. Careful measurements

led to a fairly accurate empirical determination of the function \(f\). At first Planck, relying on these empirical measurements, succeeded in finding for this function a representation that reproduced them rather well, namely

\[ \rho=\frac{8\pi h\nu^3}{c^3}\,\frac{1}{\exp\left(\frac{h\nu}{kT}\right)-1}, \]

where \(h\) and \(k\) are two universal constants; the first of them led to quantum theory. This formula looks somewhat strange because of its denominator. Does it admit of a theoretical justification? Planck did indeed find a justification, whose imperfections were at first hidden; this latter circumstance was a real blessing for the development of physics. If this formula is correct, then it allows one, with the aid of Maxwell’s theory, to calculate the mean energy \(E\) of a quasi-monochromatic oscillator situated in the radiation field:

\[ E=\frac{h\nu}{\exp\left(\frac{h\nu}{kT}\right)-1}. \]

Planck preferred to try to calculate this latter quantity theoretically. In this attempt thermodynamics no longer helped, just as Maxwell’s theory did not help. But one property of this formula was very encouraging. Namely, for high values of the temperature (with \(\nu\) constant) the formula gave the expression

\[ E=kT. \]

This is the same expression as that given by the kinetic theory of gases for the mean energy of a material point capable of performing elastic oscillations in one dimension. Kinetic theory gives

\[ E=\frac{R}{N}T, \]

where \(R\) is the gas constant, and \(N\) is the number of molecules in a gram-molecule; this constant is connected with the absolute size of the atom. If we equate the two expressions, we obtain:

\[ N=\frac{R}{k}. \]

Thus, one of the constants of Planck’s formula gives exactly the true size of the atom. The numerical value agreed satisfactorily with determinations of \(N\), admittedly not very precise, made on the basis of the kinetic theory of gases.

This was a great success, as Planck clearly realized. However, there is also a reverse side here, a rather unpleasant one, which,

Fortunately, Planck did not notice this at once. Namely, the reasoning requires that the relation \(E = kT\) be valid also for low temperatures. But then Planck’s formula and its constant \(h\) would also have vanished. The correct conclusion from the existing theory would therefore have been the following: either the mean kinetic energy of an oscillator is obtained incorrectly from the theory of gases, which would mean a refutation of mechanics; or else the mean energy of an oscillator is obtained incorrectly from Maxwell’s theory, which would mean a refutation of the latter. Under these circumstances the most probable thing is that both theories are true only in the limiting case, and otherwise are not true; this is in fact the case, as we shall see below. If Planck had come to this conclusion, he might not have made his great discovery, because the very basis of his reasoning would have disappeared.

Let us return to Planck’s reasoning. On the basis of the kinetic theory of gases, Boltzmann found that entropy is equal, up to a constant factor, to the logarithm of the “probability” of the state under consideration. By this he clarified the essence of processes “irreversible” in the thermodynamic sense. On the contrary, from the molecular-mechanical point of view all processes are reversible. If a state determined in the sense of molecular theory is called a microscopic state or, more briefly, a microstate, and a state described thermodynamically a macrostate, then to each macroscopic state there will correspond a large number \((Z)\) of microstates. Then \(Z\) is a measure of the probability of the given macrostate. This idea seems extremely important also because its applicability is not limited to a microscopic description based on mechanics. Planck noticed this and applied Boltzmann’s principle to a system consisting of a very large number of resonators with the same frequency \(\nu\). The macroscopic state is specified by the total vibrational energy of all the resonators; the microstate is specified if the energy of each individual resonator is given (for the given instant). In order that the number of microstates belonging to one macrostate be finite, Planck divided the total energy into a large but finite number of identical energy elements \(\varepsilon\) and posed the question: in how many ways can these energy elements be distributed among the resonators? The logarithm of this number then gives the entropy, and with it (by the thermodynamic route) the temperature of the system. Planck obtained his formula by taking, for the energy elements \(\varepsilon\), the value \(\varepsilon = h\nu\). The decisive circumstance here is that the result is obtained only if one takes for \(\varepsilon\) a definite finite value and, consequently, does not pass to the limit \(\varepsilon = 0\). This form of reasoning obscures the fact that it contradicts the mechanical and electrodynamic foundation on which the derivation rests in all other respects. In reality, however, this derivation implicitly assumes that individual resonators can absorb and emit energy only in “quanta” of magnitude \(h\nu\). This means that the energy

mechanical oscillatory system, just as radiant energy, can be transmitted only in such quanta—a rebuke to the laws of mechanics and electrodynamics. Here the contradiction with dynamics was fundamental, whereas the contradiction with electrodynamics might also have been not so profound. Namely, the expression for the density of radiant energy is compatible with Maxwell’s equations, but it is not a necessary consequence of these equations. That this expression correctly gives important mean values is evidenced at least by the fact that the Stefan–Boltzmann and Wien laws based on it agree with experiment.

All this became clear to me soon after the appearance of Planck’s basic work, so that, although I had no substitute for classical mechanics, I could nevertheless see what consequences this law of thermal radiation led to, both for the photoelectric effect and for other phenomena related to it, connected with transformations of radiant energy, and also for the heat capacities of bodies, in particular of solids. But all my attempts to adapt the theoretical foundations of physics to these results ended in complete failure. It was as if the ground had slipped away from under one’s feet and nowhere could firm ground be seen on which to build. It always seemed to me a miracle that this wavering and wholly contradictory foundation proved sufficient to allow Bohr—a man of brilliant intuition and subtle feeling—to find the most important laws of spectral lines and of the electron shells of atoms, including their significance for chemistry. This seems to me a miracle even now. It is the highest musicality in the realm of thought.

My personal interests in those years were directed not so much toward individual consequences of Planck’s results, however important they might be; my main question was the following. What general conclusions does the radiation formula allow one to draw concerning the structure of radiation and, in general, concerning the electromagnetic basis of physics? Before speaking of this in greater detail, I must briefly mention several investigations relating to Brownian motion and to subjects akin to it (fluctuation phenomena), and based chiefly on classical kinetic theory. Not being acquainted with the earlier investigations of Boltzmann and Gibbs, which in essence exhaust the question, I developed statistical mechanics and the molecular-kinetic theory of thermodynamics founded upon it. In doing so, my chief aim was to find facts that would establish as reliably as possible the existence of atoms of definite finite size.

Not knowing that observations of “Brownian motion” had long been known, I discovered that atomistic theory leads to the existence of an observable motion of suspended microscopic particles. The simplest derivation was based on the following considerations. If the molecular-kinetic theory is correct in principle, then a suspension of visible particles must, like a solution,

molecules have an osmotic pressure obeying the gas laws. This osmotic pressure depends on the true dimensions of the molecules, i.e., on the number of molecules in a gram-equivalent. If the density of the suspension is nonuniform, then the resulting spatial nonuniformity of the osmotic pressure causes an equalizing diffusion motion, which can be calculated from the known mobility of the particles. On the other hand, the same diffusion process may be regarded as the result of random displacements of suspended particles under the action of thermal motion, with the magnitude of the displacements in advance unknown. Equating the values of the diffusion flux obtained in both ways, we arrive at a quantitative expression of the statistical law for these displacements, i.e., at the law of Brownian motion. The agreement of these conclusions with experiment, as well as Planck’s determination of the true magnitude of the molecule from the law of radiation (for high temperatures), convinced the numerous skeptics of that time (Ostwald, Mach) of the reality of atoms. The prejudice of these scientists against atomic theory may undoubtedly be attributed to their positivist philosophical standpoint. This is an interesting example of how philosophical prejudices hinder the correct interpretation of facts even for scientists with bold thinking and subtle intuition. The prejudice—which has survived to this day—consists in the conviction that facts by themselves, without free theoretical construction, can and must lead to scientific knowledge. Such self-deception is possible only because it is not easy to realize that even those concepts which, thanks to verification and prolonged use, seem to be directly connected with empirical material are in fact freely chosen.

The success of the theory of Brownian motion again clearly showed that classical mechanics invariably gives reliable results when it is applied to motions for which the higher time derivatives of the velocity may be neglected. On the recognition of this fact one can build a comparatively direct method that makes it possible to learn something from Planck’s formula about the structure of radiation. Namely, the following conclusion can be drawn. A freely moving mirror (perpendicular to its plane), reflecting quasi-monochromatically, must, in a space filled with radiation, perform something like Brownian motion with a mean kinetic energy equal to \(\frac{1}{2}(R/N)T\) (\(R\) is the constant in the equation of state for one gram-molecule, \(N\) is the number of molecules in a gram-molecule, \(T\) is the absolute temperature). If the radiation experienced no local fluctuations, the mirror would gradually come to rest, since, owing to its motion, more radiation is reflected from its front side than from its rear side. But the mirror must be subject to the action of fluctuations of the pressure it experiences, because the wave trains composing the radiation interfere with one another; these fluctuations

may be computed from Maxwell’s theory. Such a computation shows, however, that these pressure fluctuations are insufficient (especially at small radiation densities) to impart to the mirror the mean kinetic energy \(\frac{1}{2}(R/N)T\). To obtain such a value of the energy, one must assume that there exist pressure fluctuations of another kind, not following from Maxwell’s theory. These fluctuations correspond to the assumption that the energy of radiation consists of indivisible quanta of energy \(h\nu\) (with momenta \(h\nu/c\), where \(c\) is the speed of light), possessing point localization, and that these quanta are reflected as wholes, without being broken up. The argument just given showed in the most vivid and direct way that Planck’s quanta must be ascribed a kind of immediate reality of their own; consequently, with respect to energy, radiation must possess a kind of molecular structure, which, of course, contradicts Maxwell’s theory. Applying to radiation other considerations based directly on Boltzmann’s relation between probability and entropy (where probability is equated with statistical frequency in time), one can arrive at the same result. This dual nature of radiation (and of material particles) is a fundamental property of reality, which quantum mechanics interpreted in an ingenious and astonishingly successful way. Almost all modern physicists consider this interpretation essentially final, whereas to me it seems only a temporary way out; several remarks on this will follow below.

Thanks to considerations of this kind, already soon after 1900, i.e. soon after Planck’s fundamental work, it became clear to me that neither mechanics nor thermodynamics can lay claim to complete exactness (except in limiting cases). Gradually I began to despair of the possibility of arriving at true laws by means of constructive generalizations of known facts. The longer and the more desperately I tried, the more I came to the conclusion that only the discovery of a general formal principle could lead us to reliable results. Thermodynamics appeared to me as a model. There the general principle was given in the proposition: the laws of nature are such that the construction of a perpetuum mobile (of the first and second kind) is impossible. But how is one to find a general principle similar to this one? I obtained such a principle after ten years of reflection from a paradox upon which I had stumbled already at the age of 16. The paradox is the following. If I were to begin moving after a ray of light with speed \(c\) (the speed of light in a vacuum), then I ought to perceive such a ray of light as a spatially oscillating electromagnetic field at rest. But nothing of the kind exists; this is evident both on the basis of experience and from Maxwell’s equations. It seemed intuitively clear to me from the very beginning that, from the point of view of such an observer, everything must take place according to the same

...the same laws as for an observer at rest relative to the earth. Indeed, how can the first observer know or establish that he is in a state of rapid uniform motion?

It can be seen that this paradox already contains the germ of the special theory of relativity. Now, of course, everyone knows that all attempts to explain this paradox satisfactorily were doomed to failure so long as the axiom of the absolute character of time and simultaneity remained rooted, even if unconsciously, in our thinking. To establish the presence of this axiom and to recognize its arbitrariness is, in essence, already to solve the problem. The critical thinking needed in order to feel out this central point was greatly aided, in particular, by reading the philosophical works of David Hume and Ernst Mach.

It was necessary to form a clear conception of what spatial coordinates and the time of a given event mean in physics. The physical interpretation of spatial coordinates presupposed the existence of a rigid reference body (reference system), which, moreover, must be in a more or less definite state of motion (an inertial system). For a given inertial system the coordinates meant the results of certain measurements made with rigid (stationary) rods. (It should always be borne in mind that the assumption that rigid rods exist in principle is naturally suggested by everyday experience, but is in essence arbitrary.) Under such an interpretation of spatial coordinates, the question of the validity of Euclidean geometry becomes a physical problem.

In order to interpret the time of a given event in an analogous way, a means is needed for measuring intervals of time (such a means is a periodic process proceeding in a determinate manner and carried out by a system of sufficiently small spatial dimensions). Clocks fixed at rest relative to an inertial system determine local time. The totality of the local times of all spatial points constitutes the “time” belonging to the chosen inertial system, if, in addition, a method is given for “synchronizing” all these clocks with one another. Obviously, it is by no means necessary a priori that the “times” of different inertial systems determined in this way coincide with one another. The discrepancy would have been noticed long ago if light had not seemed (thanks to the large magnitude of \(c\)) to be a means for establishing absolute simultaneity—at least in the practice of everyday experience.

The assumptions of the (in-principle) existence of (ideal or perfect) measuring rods and clocks are not independent of one another. Indeed, if one considers that the assumption of the constancy of the velocity of light in a vacuum does not lead to contradictions, then light...

a signal, reflected back and forth from mirrors at the ends of a rigid rod, constitutes an ideal clock.

The above-mentioned paradox may be formulated as follows. According to the rules used in classical physics for transforming the spatial coordinates and times of events in passing from one inertial system to another, the following two propositions:

1) the constancy of the speed of light,

2) the independence of the laws (and hence, in particular, of the law of constancy of the speed of light) from the choice of inertial system (the special principle of relativity),

are incompatible with one another (although each, taken separately, is confirmed by experiment).

At the foundation of the special theory of relativity lies the recognition that propositions 1) and 2) are mutually compatible if, in recalculating the coordinates and times of events, one applies transformation rules of a new kind (“the Lorentz transformation”). With the given physical interpretation of coordinates and time, this assertion means not merely a conventional step, but includes definite hypotheses concerning the actual behavior of moving scales and clocks—hypotheses which can be confirmed or refuted by experiment.

The general principle of the special theory of relativity is contained in the postulate: the laws of physics are invariant with respect to Lorentz transformations (which give the transition from one inertial system to any other inertial system). This is a restrictive principle for the laws of nature, which may be compared with the restrictive principle, lying at the foundation of thermodynamics, of the nonexistence of a perpetuum mobile.

Let us first say a few words about the relation of the theory to “four-dimensional space.” A very widespread error is the opinion that the special theory of relativity somehow discovered, or newly introduced, the four-dimensionality of the physical manifold (continuum). Of course this is not so. The four-dimensional manifold of space and time also lies at the foundation of classical mechanics. Only in the four-dimensional continuum of classical physics do the “sections” corresponding to a constant value of time possess absolute (i.e., not dependent on the choice of reference system) reality. Thus the four-dimensional continuum naturally decomposes into three-dimensional and one-dimensional (time), so that four-dimensional consideration is not imposed as necessary. The special theory of relativity, on the contrary, creates a formal dependence between the manner in which spatial coordinates, on the one hand, and the time coordinate, on the other, must enter the laws of nature.

Minkowski’s important contribution to the theory consists in the following. Before Minkowski’s study, to verify the invariance of a physical law it was necessary to carry out the Lorentz transformation on it to the end. Minkowski, however, succeeded in introducing such a formal apparatus that the mathematical form of the law itself already ensures its invariance;

with respect to Lorentz transformations. By creating four-dimensional tensor calculus, Minkowski gave to four-dimensional space what ordinary vector calculus gives to three spatial dimensions. He also showed that the Lorentz transformation is nothing other than a rotation of the coordinate system in four-dimensional space (apart from the difference in sign, due to the special character of time).

Let us now make a critical remark about the theory in the form in which it has been characterized above. It may be noted that the theory introduces (besides four-dimensional space) two kinds of physical objects, namely: 1) rods and clocks, 2) everything else—for example, the electromagnetic field, the material point, etc. This is, in a certain sense, illogical; strictly speaking, the theory of rods and clocks ought to be derived from the solutions of the basic equations (taking into account that these objects have an atomic structure and are in motion), and not regarded as independent of them. The usual procedure has its justification, however, since from the very beginning the inadequacy of the adopted postulates for grounding a theory of rods and clocks is clear. These postulates are not so strong that sufficiently complete equations for physical processes could be derived from them. If one does not in general renounce the physical interpretation of coordinates (which in itself would be possible), then it is better to admit such an inconsistency, but with the obligation to get rid of it at a later stage of the theory’s development. This sin, however, must not be legitimized to such a degree as to permit, for example, the use of the conception of distance as a physical entity of a special kind, essentially different from other physical quantities (reducing physics to geometry, etc.).

Let us now clarify what are the finally established truths for which physics is indebted to the special theory of relativity.

1) There is no simultaneity of distant events; hence there is no direct action at a distance in the sense of Newtonian mechanics. True, according to this theory one could introduce actions at a distance propagating with the speed of light, but that would be wholly artificial; the point is that in such a theory there can be no reasonable expression for the energy principle. It therefore seems inevitable to describe physical reality by continuous functions of a point in space. In consequence of this, the material point can no longer be regarded as the fundamental concept of the theory.

2) The law of conservation of momentum and the law of conservation of energy merge into one single law. The inertial mass of a closed system is identical with its energy, so that mass ceases to be an independent concept.

Remark. The speed of light \(c\) is one of the quantities entering into the physical equations as a “universal po-

CREATIVE AUTOBIOGRAPHY

…“constant.” However, if, instead of the second, one takes as the unit of time the time in which light travels \(1\ \mathrm{cm}\), then \(c\) will no longer enter into the equations. In this sense one may say that the constant \(c\) is only an apparent universal constant.

It is well known and generally accepted that, moreover, other universal constants can be eliminated from physics if, instead of the gram and the centimeter, suitable “natural” units are introduced (for example, the mass and radius of the electron).

If one imagines this accomplished, then only “dimensionless” constants will enter into the fundamental equations of physics. With regard to these latter I should like to express one proposition, which for the present cannot be justified by anything other than faith in the simplicity and intelligibility of nature. This proposition is the following: such arbitrary constants do not exist. In other words, nature is so constituted that its laws are, to a large extent, already determined by purely logical requirements, to such a degree that the expressions of these laws contain only constants that admit a theoretical determination (i.e., such constants that their numerical values cannot be changed without destroying the theory).

The special theory of relativity owes its origin to Maxwell’s equations for the electromagnetic field. Conversely, only the special theory of relativity gives Maxwell’s equations a satisfactory formal interpretation. Maxwell’s equations represent the simplest field equations invariant with respect to Lorentz transformations that can be written for an antisymmetric tensor associated with a vector field. All this would be well if we did not know from quantum phenomena that Maxwell’s theory does not convey the energetic properties of radiation. But for the solution of the question of how exactly Maxwell’s theory should be modified (and the modification must be a natural one), the special theory of relativity does not provide sufficient indications. And to Mach’s question, “Why are inertial systems physically distinguished relative to other frames of reference?” this theory likewise gives no answer.

The fact that the special theory of relativity represents only the first step in a necessary development became clear to me only when I attempted to include gravitation as well within the framework of this theory. In classical mechanics, interpreted in the spirit of field theory, the gravitational potential is represented as a scalar field (the simplest theoretical possibility for a field with one single component). Such a scalar theory of gravitation can easily be made invariant with respect to the Lorentz transformation group. Thus the following program presents itself as natural: the complete physical field consists of a scalar field (gravitation) and a vector field (the electromagnetic field); further discoveries might have…

force us to introduce still more complicated fields, but for the time being there would be no need to worry about this.

The possibility of carrying out this program, however, seemed doubtful from the very beginning. The point is that the theory had to combine the following things:

1) From general considerations of the special theory of relativity it was clear that the inertial mass of a physical system must increase with an increase in total energy (in particular, with an increase in kinetic energy);

2) From very precise experiments (especially from Eötvös’ experiments with torsion balances) it was known empirically with very great accuracy that the gravitational mass of a body is exactly equal to its inertial mass.

From 1) and 2) it followed that the weight of a system depends in a quite definite and known way on its total energy. If the theory did not give this, or gave it only with difficulty, then it had to be rejected. This condition can be expressed most simply as follows: in the fall of a system in a given gravitational field, the acceleration does not depend on the nature of the falling system (and therefore, in particular, not on the energy contained in it).

It turned out, however, that within the framework of the proposed program this elementary state of affairs could not in general be taken into account properly; at any rate not without difficulty. This convinced me that within the framework of the special theory of relativity there is no place for a satisfactory theory of gravitation.

And then it occurred to me: the fact of the equality of inertial and gravitational mass—or, in other words, the fact that the acceleration of free fall does not depend on the nature of the falling substance—also admits of another expression. It can be expressed as follows: in a gravitational field (of small spatial extent) everything happens as it does in space without gravitation, if instead of an “inertial” system of reference one introduces a system accelerated relative to it.

Thus, if one considers the behavior of bodies in an accelerated system of reference to be due, as it were, to a “true” gravitational field (and not merely an apparent one), then this system of reference may be regarded as “inertial” with the same right as the original system.

If one regards as possible any gravitational fields whatever, extending arbitrarily far and not restricted by boundary conditions, then the concept of an inertial system becomes meaningless. The concept of “acceleration with respect to space” then loses all meaning, and with it the principle of inertia as well; moreover, Mach’s paradox also disappears.

Thus, the equality of inertial and gravitational mass leads quite naturally to the thought that the basic requirement of the special theory of relativity (the invariance of laws with respect to the Lorentz transformation) is too narrow, i.e., that it is necessary

postulate the invariance of laws also with respect to nonlinear transformations of coordinates in the four-dimensional continuum.

This took place in 1908. Why were another 7 years needed in order to construct the general theory of relativity? The main reason is the following: it is not so easy to free oneself from the idea that coordinates have a direct metrical meaning. The turning point came about in approximately the following way.

We start from empty space without a field, in the form in which it is considered—in an inertial frame of reference—in the special theory of relativity. This is the physically simplest possible case. Let us now imagine a noninertial system introduced in such a way that it moves relative to the inertial system in one direction (in the three-dimensional sense) with constant acceleration (suitably defined). With respect to this system there arises a static parallel gravitational field. In this case the system of reference may be taken to be rigid, with a three-dimensional Euclidean metric. But in a uniformly accelerated system, in which there is a static field, clocks do not run as identically constructed clocks in a stationary system do. From this particular example it is already clear that the immediate metrical meaning of the coordinates is lost, if one admits nonlinear coordinate transformations at all. But this must be done if one seeks to make the equality of heavy and inertial mass fundamental to the theory, and if one seeks to overcome Mach’s paradox concerning inertial systems.

But once one has to abandon assigning to coordinates a direct metrical meaning (the difference of coordinates being equal to a measured length or to an interval of time), one can no longer avoid recognizing the equivalence of all coordinate systems obtained by continuous transformations.

Accordingly, the general theory of relativity proceeds from the following fundamental proposition. The laws of nature must be expressed by equations that are covariant with respect to the group of continuous transformations of coordinates. This group thus takes the place here of the group of Lorentz transformations of the special theory of relativity; this latter group is a subgroup of the first group.

By itself this requirement cannot, of course, yet serve as a sufficiently definite starting point for deriving the fundamental equations of physics. First of all, it may even be disputed whether this requirement contains an actual restriction on physical laws; indeed, if a given law is first postulated only for certain coordinate systems, it can always be reformulated so that the new formulation already has a generally covariant form. Moreover, it is clear from the very beginning that there exists an innumerable multitude of field equations admitting

such a generally covariant formulation. The outstanding heuristic significance of the general principle of relativity consists in the following: it leads us to seek those systems of equations which, while being generally covariant, are at the same time the simplest; among these systems we must look for the field equations expressing the properties of physical space. Fields obtained from one another by coordinate transformations reflect one and the same reality.

For the investigator in this field the chief question is the following: of what mathematical character will be the quantities (functions of the coordinates) through which the physical properties of space (the “structure”) are expressed? And only after that: what equations do these quantities satisfy?

Today we still cannot answer these questions with certainty. The path which I followed in the first formulation of the general theory of relativity may be characterized as follows. Even if we do not know what the variables (the field structure) are by means of which physical space should be described, one particular case is reliably known to us: the “field-free” space of the special theory of relativity. Such a space is characterized by the fact that, in a suitably chosen coordinate system, the expression pertaining to two neighboring points,

\[ ds^{2}=dx_{1}^{2}+dx_{2}^{2}+dx_{3}^{2}-dx_{4}^{2} \tag{1} \]

represents a measurable quantity (the square of the distance), and therefore has real physical meaning. Referred to an arbitrary system, this quantity is expressed as

\[ ds^{2}=g_{ik}dx_i dx_k, \tag{2} \]

where the indices run through the values from 1 to 4. The quantities \(g_{ik}\) form a symmetric tensor. If, after carrying out a transformation on expression (field) (1), one obtains \(g_{ik}\) with nonvanishing first derivatives with respect to the coordinates, then, relative to this coordinate system, there exists as it were a gravitational field (in the sense of the reasoning set forth above), namely a gravitational field of a quite special kind. Thanks to Riemann’s investigations of \(n\)-dimensional metric spaces, this special field can be invariantly characterized as follows:

1) The Riemann curvature tensor \(R_{iklm}\), formed from the coefficients of the metric (2), is equal to zero.

2) With respect to an inertial system (in which expression (1) is valid), the trajectory of a material point is a straight line, and thereby an extremal (geodesic line). This latter assertion represents such a characterization of the law of motion as rests upon expression (2).

The general law of physical space must be a generalization of the law just written down. Here I assumed that there are two levels of generalization:

a) the pure gravitational field,
b) the general field (in which quantities also occur that in some way correspond to the electromagnetic field).

Case a) was characterized by the fact that, although the field could still be represented by a Riemannian metric (2) with the corresponding symmetric tensor, there was no representation of the form (1) (except, as it were, in the infinitely small). This means that in case a) the Riemann tensor does not vanish. It is clear, however, that in this case field equations must hold which express a law that is a generalization (weakening) of the former law. If it is required that these equations also be of second order and linear in the second derivatives, then these conditions are satisfied only by the equations

\[ 0=R_{kl}=g^{lm}R_{lklm}, \]

obtained from the preceding ones by a single contraction. Only these equations could be regarded as the field equations in case a). Further, it is natural to assume that in case a) as well the geodesic line continues to give the law of motion of a material point.

At that time an attempt to find a representation for the complete field b) and to obtain equations for it seemed to me hopeless, and I did not venture upon it. I preferred to establish preliminary formal frameworks for the representation of the whole of physical reality. This was necessary in order to be able to investigate, at least provisionally, the suitability of the basic idea of general relativity. It proceeded as follows.

In Newton’s theory, one may write, as the law for the gravitational field, the equation

\[ \Delta\varphi=0 \]

(where \(\varphi\) is the gravitational potential), which must be satisfied in places where the density \(\rho\) of matter is equal to zero. In the general case one should put

\[ \Delta\varphi=4\pi k\rho \qquad (\rho \text{ is the mass density}) \]

(Poisson’s equation). In the relativistic theory of the gravitational field, \(R_{ik}\) takes the place of \(\Delta\varphi\). On the right-hand side we must then put, instead of \(\rho\), also a tensor. Since we know from the special theory of relativity that (inertial) mass is equal to energy, the tensor of energy densities should be placed on the right-hand side—more precisely, of the total energy density, since it does not belong to the pure gravitational field. We thus arrive at the field equations

\[ R_{ik}-\frac{1}{2}g_{ik}R=-\varkappa T_{ik}. \]

The second term on the left-hand side was added for formal reasons, namely, the left-hand side is written so that its divergence, in the sense of the absolute differential calculus, is identically equal to zero. The right-hand side includes everything that cannot yet be combined in a single field theory. Of course, I did not for a minute doubt that such a formulation is only a temporary way out of the situation, undertaken with the aim of giving the general principle of relativity some closed expression. This formulation was, after all, in essence no more than a theory of the gravitational field somewhat artificially detached from the unified field (Gesamtfeld) of still unknown structure.

In the theory sketched, apart from the requirement of invariance of the equations with respect to the group of continuous transformations of the coordinates, only the limiting case of the pure gravitational field and the connection of this field with the metric structure of space can perhaps lay claim to unconditional (final) significance. Therefore we shall now speak only about the equations of the pure gravitational field.

A distinctive feature of these equations is, on the one hand, their complex structure, especially their nonlinear character with respect to the field variables and their derivatives, and, on the other hand, their uniqueness, i.e. the logical necessity with which the group of transformations determines the form of these complex equations. If we were to stop at the special theory of relativity, i.e. at invariance with respect to the Lorentz group, then the field equations

\[ R_{ik}=0 \]

would remain invariant also within the framework of this narrower group. But from the point of view of the narrower group, above all, there would be no reason whatever to suppose that gravitation must be described by so complex a system of quantities (a structure) as the symmetric tensor \(g_{ik}\). Even if sufficient reasons for this could be found, it would turn out that there exists an innumerable number of field equations constructed from the quantities \(g_{ik}\), all of which are covariant with respect to Lorentz transformations (but not with respect to the general group). Even if by chance, from all conceivable laws invariant in the Lorentz group, one succeeded in guessing precisely the one to which the broader group belongs, nevertheless we would not have attained that degree of knowledge which the general principle of relativity gives us. For from the point of view of the Lorentz group, two solutions connected by a nonlinear coordinate transformation would have to be regarded as physically different, which is incorrect, since from the point of view of the general group they give only two different representations of one and the same field.

One more general remark about the structure of the field and the group. It is clear that, generally speaking, a theory appears to us the more perfect, the simpler the “structure” of the field underlying it and the broader...

that group with respect to which the field equations are invariant. But these two requirements, obviously, come into conflict with one another. According to the special theory of relativity (the Lorentz group), one can, for example, write a covariant equation already for the simplest conceivable structure (a scalar field), whereas in the general theory of relativity (a broader group of continuous coordinate transformations) invariant field equations exist only for a more complicated structure, namely for a symmetric tensor. In support of the fact that in physics one must require invariance with respect to the broader group, we adduced physical arguments; from a purely mathematical point of view I see no necessity to sacrifice the simpler structure of the field to the breadth of the group*).

The group of general relativity leads for the first time to the fact that the simplest invariant law will no longer be linear and homogeneous in the field variables and their derivatives. This is a circumstance of fundamental importance, and for the following reason. If the field equations are linear (and homogeneous), then the sum of two solutions will again be a solution; this is the case, for example, for Maxwell’s field equations in empty space. In such a (linear) theory of field equations, the field equations are insufficient for deriving the law of interaction between objects which are described (each separately) by solutions of the system of field equations. Therefore, in the earlier theories, alongside the field equations, special equations were necessary that determined the motion of material objects under the action of the field. True, initially in the relativistic theory of gravitation there was postulated, alongside the laws for the field and independently of it, also a law of motion (the geodesic line). But subsequently it became clear that it is not necessary, and indeed is not possible, to introduce the law of motion independently, and that it is implicitly contained in the law for the gravitational field.

The essence of this rather complex state of affairs can be represented more vividly in the following way. A single stationary material point is represented by a gravitational field which, of course, is regular everywhere except at the place where the material point itself is located; at this place the field has a singularity. If, by integrating the field equations, one computes the field corresponding to two stationary material points, then it will have, besides singularities at the material points, also a singular line connecting the material points with one another. But one can prescribe the motion of the material points in such a way that the gravitational field determined by them outside the material

*) To remain with the narrower group and at the same time to take a more complicated field structure (the same as in the general theory of relativity) means naïve inconsistency. Sin remains sin, even if it is committed by men who are otherwise highly respectable.

... points nowhere had singularities. These will be precisely those motions which, to a first approximation, are described by Newton’s laws. Thus, one may say: the masses move in such a way that the field equations admit solutions having no singularities in the space outside the masses. This property of the equations of gravitation is directly connected with their nonlinearity, and this, in turn, is due to the broader group of transformations.

Here, however, the following objection could be raised. If singularities are admitted at the locations of material points, then what justification is there for forbidding singularities in the remaining space? This objection would be justified if the equations of gravitation could be regarded as the equations of a single complete field. Under the existing state of affairs, however, we must say that the field of a material particle can be regarded as a pure gravitational field with ever less right the nearer we come to the particle itself. If we had equations for a single complete field, then it would have to be required that the particles themselves, too, could be represented as solutions of the complete field equations, having no singularities anywhere. Only then would the general theory of relativity become a closed theory.

Before passing to the question of completing the general theory of relativity, I must state my position with respect to that physical theory which, of all the physical theories of our time, has achieved the greatest successes. I have in mind statistical quantum mechanics, which acquired a coherent logical form about twenty-five years ago (Schrödinger, Heisenberg, Dirac, Born). It is the only modern theory that gives a coherent explanation of what we know about the quantum character of micromechanical processes. This theory, on the one hand, and the theory of relativity, on the other, are both, in a certain sense, considered correct, although the fusion of these theories has not yet been achieved, despite all efforts. Connected with this must be the fact that among contemporary theoretical physicists there are entirely different opinions about what the theoretical foundation of future physics will look like. Will it be a field theory? Will it be a theory that is, in the main, statistical? I shall say here briefly what I think about this.

Physics is the endeavor to comprehend being as something that is conceived as independent of perception. In this sense one speaks of the “physically real.” In pre-quantum physics there was no doubt as to how this should be understood. In Newton’s theory reality was represented by material points in space and in time; in Maxwell’s theory—by the field in space and in time. In quantum mechanics this is less clear. If one asks whether the function \(\psi\) of quantum theory represents some real state of affairs, a real

in the same sense as a system of material points or an electromagnetic field, people are slow to give the simple answer “yes” or “no.” Why? The function \(\psi\) (at a definite moment of time) expresses the following: what is the probability that a certain physical quantity \(q\) (or \(p\)) will be found in a certain specified interval if I measure it at the moment \(t\)? Here probability must be regarded as a quantity accessible to experimental determination, i.e. as an unconditionally “real” quantity. I shall be able to determine it if I create the very same function \(\psi\) very many times and each time measure \(q\). But how does the matter stand with an individual measurement of \(q\)? Did the corresponding individual system possess the given value \(q\) already before the measurement? Within the framework of the theory there is no definite answer to this question, because measurement is, after all, a process involving a finite external intervention in the system; therefore one may imagine that the system receives a definite (namely, the measured) numerical value \(q\) (or \(p\)) only as a result of the measurement itself. For the further discussion I shall imagine two physicists, \(A\) and \(B\), who adhere to different understandings of the real state described by the function \(\psi\).

\(A.\) An individual system possesses (before measurement) a definite value \(q\) (or \(p\)) for all variables of the system; this is the value that is established when these variables are measured. Proceeding from this understanding, he will declare: the function \(\psi\) is not an exhaustive representation of the real state of the system; it expresses only what we know about the system from previous measurements.

\(B.\) An individual system does not possess (before measurement) a definite value \(q\) (or \(p\)). The measured value arises only thanks to the act of measurement, with the probability corresponding to that value obtained from the function \(\psi\). Proceeding from this understanding, he will declare (or at least has the right to declare): the function \(\psi\) is an exhaustive representation of the real state of the system.

And now we shall bring the following case to the attention of both these physicists. Suppose there is a system consisting (at the moment \(t\) under consideration) of two subsystems \(S_1\) and \(S_2\), which at this moment are spatially separated and do not interact in any noticeable way in the sense of classical physics. Let the whole system be completely described, in the sense of quantum mechanics, by a known wave function, namely the function \(\psi_{12}\). All quantum theorists agree among themselves on the following. If I perform a complete measurement on \(S_1\), then from the results of the measurement and from \(\psi_{12}\) I obtain a quite definite wave function \(\psi_2\) of the system \(S_2\). In this case the character of \(\psi_2\) depends on what kind of measurement has been performed on \(S_1\). And so it seems to me that one can speak of the real state of affairs in the subsystem \(S_2\). Of this real state of affairs we know in advance still less than of the system described by the wave function. But one supposi-

ALBERT EINSTEIN

the proposition seems to me indisputable. The real state of affairs (state) of the system \(S_2\) does not depend on what is done with the system \(S_1\), spatially separated from it. But, depending on what kind of measurement I perform on \(S_1\), I obtain for the second subsystem different \(\psi_2\) \((\psi_2, \psi_2^1,\ldots)\). The real state of \(S_2\), however, must be independent of what happens in \(S_1\). Thus, for one and the same real state \(S_2\), different functions \(\psi_2\) may be found (depending on the choice of measurement on \(S_1\)). (This conclusion could have been avoided in only one of two ways. Either one must suppose that the measurement on \(S_1\) changes (telepathically) the real state \(S_2\), or else one must deny that things spatially separated from one another can in general have independent real states. Both appear to me completely unacceptable.)

And so, if physicists \(A\) and \(B\) consider this reasoning correct, then \(B\) will have to give up the admission that the function \(\psi\) is a complete description of the real state of affairs. For in that case it would be impossible for two different wave functions to correspond to one and the same state of affairs (in \(S_2\)).

Then the statistical character of the contemporary theory would be a necessary consequence of the incompleteness of the description of systems in quantum mechanics, and there would no longer be any basis for believing that in the future physics must be founded on statistics.

My opinion amounts to this: if one takes as a basis certain concepts borrowed chiefly from classical mechanics, then contemporary quantum theory may be regarded as the best formulation of real relations. However, I do not think that this theory is a suitable starting point for future development. This is the point at which my expectations diverge from those of the majority of contemporary physicists. They are convinced that the essential features of quantum phenomena (the, as it were, discontinuous and not temporally determined changes of the state of a system, the corpuscular and at the same time wave properties of elementary formations carrying energy) cannot be taken into account by a theory describing the real state of things by continuous functions of the coordinates satisfying certain differential equations. They also think that in this way it will be impossible to interpret the atomic structure of matter and radiation. They expect that the systems of differential equations that might be involved in such a theory generally have no regular (singularity-free) solutions in all four-dimensional space. But above all they believe that, apparently, the discontinuous character of elementary processes can be represented only by a theory that is essentially statistical; in such a theory the discontinuous changes of systems must be taken into account

by means of a continuous change in the probabilities of possible states.

All these remarks seem to me rather weighty. But the main question, as it seems to me, is the following.

What direction promises success in the present state of theories? In choosing a direction I am inclined to be guided by my experience in constructing the theory of gravitation. The equations of this theory give, in my opinion, more hope of obtaining something exact than all the other equations of physics. Let us take for comparison, for example, Maxwell’s equations for empty space. They are a formulation corresponding to observations of infinitely weak electromagnetic fields. This empirical origin already determines their linear form; but we have already indicated that true laws cannot be linear. Linear laws satisfy, with respect to their solutions, the principle of superposition and, consequently, say nothing about the interactions of elementary formations. True laws cannot be linear and cannot be obtained from linear laws. The theory of gravitation taught me something else as well: a collection of empirical facts, however extensive it may be, cannot lead to the establishment of such complex equations. A theory can be tested by experience, but there is no path from experience to the construction of a theory. Equations of such a degree of complexity as the field equations of gravitation can be found only by finding a logically simple mathematical condition that determines completely, or almost completely, the form of these equations. But once such sufficiently rigid formal conditions have been established, very little factual data are needed for the construction of a theory. In the case of the gravitational equations, such formal conditions are: the existence of four dimensions and the assumption that the structure of space is determined by a symmetric tensor. These conditions, together with the requirement of invariance with respect to the group of continuous transformations, determine the form of the equations practically quite unambiguously.

Our task is to find the equations for the complete field. The sought structure of the field must be a generalization of the symmetric tensor. The group must not be narrower than the group of continuous coordinate transformations. If one now introduces a more complicated structure, then this group will no longer determine the equations as rigidly as in the case of a structure characterized by a symmetric tensor. Therefore, it would be best of all if it were possible to enlarge the group again, by analogy with the step that led from special relativity to general relativity. I tried, in particular, to bring in here the group of complex coordinate transformations. All such attempts were unsuccessful. I also rejected an explicit or hidden increase in the number of dimensions.

of space. This direction was indicated by Kaluza, and it still has its adherents (in its projective variant). We restrict ourselves to four-dimensional space and to the group of continuous real transformations of coordinates. After many years of fruitless searches, I regard as logically the most satisfactory the solution whose outline is given below.

Instead of the symmetric \(g_{ik}\) \((g_{ik}=g_{ki})\), a nonsymmetric tensor \(g_{ik}\) is introduced. This quantity is composed of a symmetric part \(s_{ik}\) and an antisymmetric part \(a_{ik}\), which may be real or purely imaginary. We have:

\[ g_{ik}=s_{ik}+a_{ik}. \]

From the point of view of group properties, such a unification of \(s_{ik}\) and \(a_{ik}\) is artificial, since each of these quantities separately has the character of a tensor. However, it turns out that these \(g_{ik}\) (considered as a whole) play in the construction of the new theory the same role as the symmetric \(g_{ik}\) in the theory of the gravitational field.

This generalization of the structure of space appears natural also from the point of view of our physical knowledge, because we know that the electromagnetic field is associated with a skew-symmetric tensor.

Furthermore, for the theory of gravitation it is essential that from the symmetric \(g_{ik}\) one can form the scalar density \(\sqrt{|g_{ik}|}\), as well as the contravariant tensor \(g^{ik}\) according to the definition

\[ g_{ik}g^{il}=\delta_k^l \quad (\delta_k^l \text{ is the Kronecker tensor}). \]

The quantities thus formed, as well as the tensor densities, admit a completely analogous definition also for nonsymmetric \(g_{ik}\).

Furthermore, in the theory of gravitation it is essential that for a given symmetric field \(g_{ik}\) one can define a field \(\Gamma^l_{ik}\), symmetric in the lower indices, whose geometrical meaning consists in the fact that it determines the parallel displacement of a vector. Analogously, for nonsymmetric \(g_{ik}\) one can define nonsymmetric \(\Gamma^l_{ik}\) by the formula

\[ g_{ik,l}-g_{sk}\Gamma^s_{il}-g_{is}\Gamma^s_{lk}=0. \tag{A} \]

This relation coincides with the corresponding relation for symmetric \(g\) with the only difference that here, of course, attention must be paid to the position of the lower indices in the quantities \(g\) and \(\Gamma\).

As in the real theory, from \(\Gamma\) one can form the curvature \(R_{iklm}\) and from it, by contraction, the curvature \(R_{kl}\). Finally,

Creative Autobiography

Using a certain variational principle and the relations (A), one can find mutually compatible field equations:

\[ \breve{\mathfrak{g}}^{ik}{}_{,s}=0 \quad \left(\text{where } \breve{\mathfrak{g}}^{ik} = \frac{1}{2}\left(g^{ik}-g^{ki}\right)\sqrt{\|g_{ik}\|}\right), \tag{B_1} \]

\[ \breve{\Gamma}^{s}_{ls}=0 \quad \left(\text{where } \breve{\Gamma}^{s}_{ls} = \frac{1}{2}\left(\Gamma^{s}_{ls}-\Gamma^{s}_{sl}\right)\right), \tag{B_2} \]

\[ \overline{R}_{kl}=0, \tag{C_1} \]

\[ \breve{R}_{kl,m}+\breve{R}_{lm,k}+\breve{R}_{mk,l}=0. \tag{C_2} \]

Here each of the equations \((B_1)\), \((B_2)\) is a consequence of the other, provided (A) is satisfied. The symbol \(\overline{R}_{kl}\) denotes the symmetric part, and the symbol \(\breve{R}_{kl}\) the antisymmetric part, of the quantity \(R_{ik}\).

In the case where the antisymmetric part \(g_{ik}\) is equal to zero, these formulas reduce to (A) and \((C_1)\). This will be the case of a pure gravitational field.

It seems to me that these formulas represent the most natural generalization of the equations of gravitation*). Testing their physical suitability is an extremely difficult task, because approximations are of no help here. The question is the following: What solutions exist for these equations that have no singularities in all space?

This account will have achieved its purpose if it has shown the reader how the efforts of an entire life are connected with one another and why they led to expectations of a certain kind.

*) If it is at all possible to proceed along the path of an exhaustive representation of physical reality on the basis of the concept of the continuum, then, in my opinion, there is a fairly high probability that the theory proposed here will be confirmed.

Submission history

CREATIVE AUTOBIOGRAPHY\*