On the Phase Contrast Method in Microscopy\*
S. M. Rytov
Submitted 1950 | SovietRxiv: ru-195001.25654 | Translated from Russian

Abstract

A lecture delivered at the colloquium of the P. N. Lebedev Physical Institute of the Academy of Sciences of the USSR.

Full Text

On the Phase Contrast Method in Microscopy*

S. M. Rytov

Despite the fact that the phase contrast method possesses interesting features and, owing to its advantages, has already found fairly wide application in microscopy, and from the purely optical point of view has been treated in considerable detail in our literature1, physicists know it comparatively little. This is not difficult to explain, considering that what is involved is a narrow question in the field of microscopic technique. However, the phase contrast method can also be approached somewhat differently, namely from the standpoint of the theory of oscillations. With such an approach this optical method appears as a special case of the manifestation of certain general oscillatory regularities, which in their totality embrace not only a number of questions concerning the formation of an optical image, but also, for example, questions of modulation in radio. In this more general connection the phase contrast method may be of more than special interest, and it is in precisely this way that it is considered below. After this remark it should no longer seem strange that at first the discussion will not concern the phase contrast method at all, nor even optics, but modulation in radio. The ideas and conclusions that can be gleaned here will not remain without application in optical questions and, moreover, may serve to clarify them better.

1. Modulation in Radio

As is known, for the transmission of speech, music, television—in general, of signals—it is necessary to disturb the sinusoidality of oscillations in the radio transmitter, i.e. to depart from the harmonic oscillation \(Ae^{i\psi}\), in which by definition \(A=\mathrm{const.}\) and \(\psi\) is a linear

function over the entire time interval \(t\) from \(-\infty\) to \(+\infty\). Modulation deviations from sinusoidality, which express the content of the radio transmission, are characterized by their slowness. This means that for \(A=A(t)\) (amplitude modulation) or for \(\dot\psi=\dot\psi(t)\), i.e., a nonlinear dependence of \(\psi\) on \(t\) (frequency or phase modulation), the changes in \(A\) and \(\dot\psi\) over a “period” \(\frac{2\pi}{\dot\psi}\) of the modulated oscillation itself are sufficiently small. Such a formulation is somewhat vague, but it gives a known idea of the nature of the process, and for what follows one can do without making it more precise.

Fig. 1

Fig. 1.

In the usual vector diagram (Fig. 1), an amplitude-modulated oscillation can be represented by a vector whose phase angle \(\varphi\) is unchanged, while its modulus changes, being the sum of a constant vector of length \(A_0\) (the modulated oscillation, or “carrier”) and a collinear with \(A_0\) variable vector \(a(t)\) (the modulating oscillation).

Fig. 2

Fig. 2.

A frequency- or phase-modulated oscillation is represented by a vector of constant length \(A_0\) (Fig. 2), which rotates as the phase \(\varphi(t)\) changes. If this rotation reduces to sufficiently small oscillations of the vector about some phase value \(\varphi_0\), then it can be interpreted as the addition to the constant vector \(A_0\) of a variable vector \(a(t)\) perpendicular to it.

One may distinguish two groups of questions connected with the action of such modulated oscillations on receiving apparatus. On the one hand, these are questions concerning the selection and reproduction of the form of the received oscillation; on the other, questions of demodulation, i.e., the extraction from the received high-frequency modulated oscillation of the modulating oscillation with low frequencies. A process of this kind, the inverse of modulation, like, for example—

than the modulation process itself, can be accomplished only by means of systems capable of transforming frequencies. These include systems with variable parameters (in particular, so-called synchronous detection, invented by engineer E. G. Molot) and nonlinear systems—devices with a nonlinear current-voltage characteristic. For the most part it is precisely such nonlinear devices—tube or crystal detectors—that are used. Conversely, for selection, linear harmonic devices are most often employed: simple oscillatory circuits or networks composed of them (band-pass filters). It is precisely for such harmonic systems that the expansion of the oscillations under investigation into harmonic oscillations, i.e. expansion in a Fourier series or Fourier integral, is of interest and has physical content.

A modulated oscillation can be represented in the form

\[ s=f(t)e^{i\omega_0 t}, \tag{1} \]

where

\[ f(t)=A(t)e^{i\varphi(t)} \tag{2} \]

is the modulating function. Its modulus \(A(t)\) expresses amplitude modulation, and its argument \(\varphi(t)\) phase modulation. Obviously, in order to expand \(s\) in a Fourier series, it is sufficient to carry out the corresponding expansion for \(f(t)\). If, for example, the modulation is periodic with period

\[ T=\frac{2\pi}{\Omega}, \]

then

\[ f(t)=\sum_{n=-\infty}^{+\infty} c_n e^{in\Omega t} \tag{3} \]

and, consequently,

\[ s=\sum_{n=-\infty}^{+\infty} c_n e^{i(\omega_0+n\Omega)t}. \tag{4} \]

Fig. 3

Fig. 3

Thus, in addition to the carrier frequency \(\omega_0\), the spectrum contains side frequencies \(\omega_0 \pm \Omega\), \(\omega_0 \pm 2\Omega,\ldots\), located on both sides of \(\omega_0\) at intervals that are multiples of \(\Omega\) (Fig. 3). The amplitudes and phases of all the harmonics of the spectrum are completely determined by the form of the modulating function \(f(t)\) through its (in the general case complex) Fourier coefficients \(c_n\).

Strictly speaking, in the problem under consideration, which concerns oscillations in time, one should use the actual

form of Fourier expansion, regarding the frequency \(\omega\) always as positive, since oscillations with frequencies \(+\omega\) and \(-\omega\) are physically indistinguishable. But for modulated oscillations, characterized by the slowness of deviations from harmonicity [in (4) this means that \(\Omega \ll \omega_0\)], any appreciable part of the spectrum occupies a comparatively narrow band of frequencies near \(\omega_0\), and the negative frequencies \(\omega_0-N\Omega<0\) begin with such high harmonic numbers \(N\) that they are of practically no interest (\(c_N\simeq 0\)). It is precisely for this reason that complex notation is possible and expedient, and not merely because such notation is more convenient. When optical questions and spatial oscillations are discussed, it is precisely the complex form that will be physically adequate. The presence of the same form in the treatment of radio modulation will facilitate comparison of the two questions.

The problem of selection consists in “passing” the desired station and rejecting the others. Thus, the resonance curve of an oscillatory circuit, or the pass band of a filter, must be so narrow that, when tuned to \(\omega_0\), it does not respond to other stations. But at the same time the curve must not be too narrow, since otherwise the filter will cut off and spoil the spectrum of the received oscillation. When the pass band is narrowed, the cutting-off begins with the high-frequency components of the modulating function; ever lower frequencies of it remain, and the form of the forced (“passed”) oscillation is smoothed more and more, coarsened, and, finally, if one reaches a selectivity that leaves only the carrier, the last traces of modulation in the response of the system disappear—there is no transmission.

One can arrive at the same result by considering not the frequency composition of the modulated oscillation and the width of the pass band, i.e. arguing not in the spectral language but in the language of times, namely, by comparing the times of variation of \(f(t)\) with the settling time of the circuit or filter. Indeed, a resonant system manages to follow only those modulation changes of the force acting on it which occur slowly in comparison with its settling time. But the more selective the harmonic system, the smaller the width \(h\) of its pass band, the greater its settling time \(\tau\) and, consequently, the more strongly it smooths the modulation. Qualitatively we arrive at the same result as in the spectral approach, and it can be shown that the agreement is not only qualitative but also quantitative. Here \(h\) and \(\tau\) are connected by a relation of the type

\[ h\cdot\tau \gg \mathrm{const}. \]

A relation of the same kind holds between the duration of a signal or pulse \(\Delta t\) and the width (spread) of its spectrum \(\Delta\omega\):

\[ \Delta\omega\cdot\Delta t \gg \mathrm{const}. \]

It is not hard to notice that these two equivalent ways of describing oscillatory phenomena— in the language of frequencies and in the language of times—strongly recall the so-called complementary representations ($t$- and $p$-representations) in wave mechanics, while the relation between $h$ and $\tau$ or between $\Delta\omega$ and $\Delta t$, at least externally, repeats the well-known uncertainty relation. There is no possibility here of dwelling on this similarity in greater detail, since a more thorough analysis would lead us too far aside; nevertheless, it must be emphasized that the indicated parallels are not accidental. They are rooted in the very spectral approach that permeates both linear radiophysics and wave mechanics. Of course, it would be incorrect to infer from the existence of this similarity between certain laws that the physical phenomena themselves are identical.

Returning to questions of radio modulation, suppose that, with the aid of a linear harmonic spectrum, the transmission of the station of interest to us has been singled out. Let us denote the “passed” oscillation, whose spectrum, generally speaking, is somewhat truncated and distorted in comparison with (4), by

\[ \tilde{s}=\tilde{f}(t)e^{i\omega_0 t}. \tag{5} \]

From this it is now necessary to extract the modulating signals themselves—telegraph signs, speech, music, etc.; that is, it is necessary to return to low frequencies, to carry out demodulation. As has already been noted, radio engineering has at its disposal a number of methods for this purpose, but, with what follows in mind, we shall consider only one—quadratic detection.

An ideal quadratic detector has the volt-ampere characteristic $i=\mathrm{const}\cdot V^2$. Real nonlinear conductors, for sufficiently small $V$, behave for the most part precisely as a quadratic detector, since in the expansion of their characteristic in a power series in $V$ the term with $V^2$ (the first nonlinear term) can usually be made dominant over the sum of the higher nonlinear terms. The simplest demodulator consists of a detector and an instrument that responds to low-frequency oscillations, for example a telephone (Fig. 4). Let the oscillation (5) be applied to this device. Then, if for simplicity we assume that the resistance of the detector is much greater than the resistances of the other elements of the circuit, from the quadratic term of the characteristic there will be obtained a current

\[ i=\mathrm{const}\cdot(\operatorname{Re}\tilde{s})^2. \]

Fig. 4.

This current obviously contains all possible combination frequencies between the components of the spectrum $\widetilde{s}$, i.e., frequencies of the form $(\omega_0+n\Omega)\pm(\omega_0+m\Omega)$. They split into high frequencies $2\omega_0+(n+m)\Omega$ and low frequencies $(n-m)\Omega$, which alone can be followed by the registering device (telephone, relay). Usually this device is shunted by a capacitance $C$ (Fig. 4), which readily passes the high-frequency part of the current.

As a result, from the total current $i$ its low-frequency part $I$ is separated out, which can be written in the form

\[ I=\mathrm{const.}\,\frac{\widetilde{s}\widetilde{s}^{*}}{2}. \]

Taking (5) into account, we obtain:

\[ I\sim \widetilde{f}\widetilde{f}^{*} \tag{6} \]

—the demodulator with a square-law detector gives the square of the modulus of the modulating function distorted by the selector. This result is of interest for what follows. With it the path from the input of the receiving device to its output ends, and, using (6), we can now determine what will be heard by us for different types of modulation and methods of selection.

Let us first suppose that the selector has not at all spoiled the spectrum of the oscillation, i.e. $\widetilde{s}=s$ and, consequently,

\[ \widetilde{f}=f=A(t)e^{i\varphi(t)}. \]

Then the response at the output of the receiver will be:

\[ I\sim A^2(t), \tag{7} \]

i.e. it reproduces the square of the modulated amplitude and in no way reflects phase modulation. Thus, with purely phase modulation $(f=A_0 e^{i\varphi(t)})$ one obtains $I=\mathrm{const.}$, i.e. the transmission is inaudible. What does this mean from the spectral point of view? Does it not follow from this that phase modulation gives too weak a spectrum of side frequencies? Such a conclusion would be absolutely incorrect. Precisely one of the features of phase modulation is that, even at small depth corresponding to frequency modulation, it can give an oscillation with an extremely extensive, multiline spectrum. The point is not at all in the small amplitudes of the side frequencies of the spectrum, but in the fact that their phases are unfavorable for square-law demodulation. This is easy to explain by the example of such shallow sinusoidal frequency modulation, in which in the spectrum practically only two side-

frequency. If \(\omega=\omega_0(1+\chi \cos \Omega t)\) and \(\dfrac{\chi\omega_0}{\Omega}<\dfrac{1}{2}\), then the spectrum has the same form as for an oscillation modulated in amplitude according to the law \(A=1+k\cos\Omega t\), with modulation depth \(k=\dfrac{\chi\omega_0}{\Omega}\). The difference lies in the phases of the spectral components: if, in amplitude modulation, their phases are taken to be zero (Fig. 5, a), then in the case of “equivalent” frequency modulation the phase of the carrier will differ by \(\pm 90^\circ\) (Fig. 5, b). Hence a completely different result is obtained under quadratic demodulation. The low

Fig. 5

Fig. 5.

frequency \(\Omega\) is formed in the detector as the difference frequency of two combinations between the carrier and the side frequencies:

\[ (\omega_0+\Omega)-\omega_0=\Omega, \]

\[ \omega_0-(\omega_0-\Omega)=\Omega. \]

In the case of amplitude modulation, both low-frequency oscillations are in phase and, when added, are doubled. In the case of phase modulation, however, they are in antiphase and cancel each other. In the general case of phase modulation with a multilinear spectrum, such mutual cancellation of combination oscillations extends to all the low frequencies present in them, except the zero frequency, i.e., except direct current.

From what has been said, the path that can lead to the detection of phase modulation with a quadratic demodulator is clear. It is evidently necessary somehow to disturb the amplitude-phase relationships in the spectrum of an oscillation modulated in phase, i.e., to specially distort its modulating function \(f\) so as to transform phase modulation—if not completely, then at least partially—into amplitude modulation. This is what is done in radio engineering with the aid of devices called discriminators. Let us consider several examples of the action of a discriminator, which in no way constitute a radio-engineering solution of the problem, but clarify the basic aspect of the matter. Moreover, these examples will be directly useful later on.

Suppose that the “discriminator” performs reception without the carrier, i.e., from the spectrum (4) of the received phase-modulated oscillation \(s=e^{i\omega_0 t+i\varphi(t)}\) it cuts out the oscillation \(c_0 e^{i\omega_0 t}\), without affecting the remaining components. Then

\[ \widetilde f=f-c_0=e^{i\varphi}-c_0, \]

and, according to (6), at the receiver output there will be a current

\[ I\sim \widetilde f\,\widetilde f^{*}=1+c_0^2-2c_0\cos\varphi(t), \tag{8} \]

in which the phase modulation has been manifested in some way. True, this is far from a brilliant reproduction of the transmission: \(I\) depends on \(\varphi\) through the even function \(\cos\), and, consequently, if \(\varphi(t)\) itself is odd, then in the current \(I\) a doubling of the fundamental frequency will result. If \(\varphi\) is small, then \(I\approx(1-c_0)^2+c_0\varphi^2\), i.e., the response decreases quadratically with decreasing \(\varphi\).

Another method of “discrimination” may be to cut off half of the spectrum—either together with the carrier, or while preserving it. It is clear that the phase modulation will necessarily appear in this case, since the demodulator does not produce, in this case, opposite-phase combinational oscillations. The distortion of the waveform will be quite strong here as well. A certain approximation to such single-sideband reception may be simply an asymmetric tuning of the resonant circuit, with a correspondingly selected width of the resonance curve (Fig. 6).

Fig. 6.

Fig. 6.

Fig. 7.

Fig. 7.

Finally, instead of completely eliminating the carrier, one may limit oneself to changing its phase by \(\pm 90^\circ\). The vector diagram immediately shows that, in the case of small phase deviations, this will lead simply to a complete transformation of phase modulation into amplitude modulation (Fig. 7), since after rotation by \(\pm 90^\circ\) the modulated vector \(A_0\) proves to be collinear with the modulating ...

to the vector \(a\). The distorted modulating function can be represented in the form

\[ \tilde f=f-c_0\pm i c_0=e^{i\varphi}-c_0\mp i c_0. \]

The corresponding response at the receiver output will be:

\[ I\sim \tilde f \tilde f^{*}=1+c_0^2-2c_0\cos\varphi(t)\pm 2c_0\sin\varphi(t). \tag{9} \]

Since the phase modulation is now revealed not only by an even but also by an odd function of \(\varphi\), frequency doubling is excluded. Moreover, for sufficiently small \(\varphi\) we obtain:

\[ I\sim (1-c_0)^2\mp 2c_0\varphi(t), \tag{9'} \]

—that is, the current at the output depends linearly on the modulated phase, i.e. the modulation is reproduced without distortion (this is the case of small \(a\), Fig. 7).

2. FORMATION OF THE OPTICAL IMAGE

Let us now turn to optics, to phenomena which, it would seem, have nothing in common with what has been discussed up to now. Let us consider the formation of an optical image of some two-dimensional

Schematic diagram with normally incident plane waves approaching an object plane \(O\), followed by planes labeled \(L\), \(F\), and \(F'\).

Fig. 8.

object and, in order to exclude everything inessential for the questions of interest to us, let us suppose that the objective is an ideal thin lens \(L\) (Fig. 8), and that the object \(O\) is illuminated by a normally incident plane wave. It is required to find the image, i.e. the distribution of illumination in the plane \(F'\), conjugate with respect to the lens to the plane of the object \(O\). Of course, the question is highly schematized, but nevertheless, even in such a formulation, it can reflect certain points essential for a number of real optical devices, including the microscope. One of these points is (in contrast to what takes place

in the telescope) the noncoincidence of the image plane \(F'\) with the principal focal plane of the lens \(F\).

The theory of the optical image can, as is known, be constructed in various ways. Following Rayleigh\(^2\), one may divide the object into “point” elements, take into account the limitation of the light wave from such an element by the aperture of the optical system, and find the corresponding field in the plane \(F'\). The image is then obtained as the intensity of the total field in \(F'\) from all the “point” elements of the object. In its result this path, as Rayleigh showed, is equivalent to another, developed earlier by Abbe\(^3\). In Abbe’s theory the field beyond the object is decomposed into plane waves and the behavior of each such wave is then traced. With this method of consideration the principal focal plane \(F\) of the lens acquires a special role, since in this plane the waves are focused, giving the Fraunhofer diffraction pattern of the object. Abbe calls this pattern the “primary image”; the image proper, i.e. the result of the interference in the plane \(F'\) of coherent waves “emitted” by the elements of the primary image, Abbe calls the “secondary image.” For questions connected with the phase-contrast method, Abbe’s method is the most appropriate.

Fig. 9.

Fig. 9.

When applying this method, it is first of all necessary to know what the aggregate of plane waves behind the object is. Obviously, this depends on the field in the plane of the object itself (or in some effective “exit plane,” sufficiently close to the surface of the object, which need not be plane). Let the object be a sufficiently thin plate in which the thickness, refractive index, and absorption depend only on one coordinate \(x\) (Fig. 9). A primary (illuminating) wave, having everywhere constant amplitude and a plane front, will give, after passing through the object, a wave of another character. Variable absorption makes the front of the transmitted wave “spotted” (or, more precisely, “striped,” since the absorption depends only on \(x\)), i.e. the amplitude of the transmitted wave in the plane \(y=0\) will be some function \(A(x)\). Variable thickness and refractive index create a phase shift that is different at different points, i.e. the phase of the wave in the plane \(y=0\) will be some function \(\varphi(x)\). In other words, the front of the transmitted wave will be “crumpled” or “corrugated.” Both kinds of changes can be represented by means of a single complex function

\[ f(x)=A(x)e^{i\varphi(x)}, \tag{10} \]

which expresses the field on the “exit plane” \(y=0\) and is sometimes called the transmittance of the object. Of course, in reality the situation may be more complicated: refraction and variable thickness may also affect the amplitude, etc.; but for a sufficiently thin plate and sufficiently small gradients of the refractive index it may be assumed that, on the plane \(y=0\) immediately adjacent to the object, the functions \(A(x)\) and \(\varphi(x)\) reproduce respectively the distributions of the absorption and refractive structure of the object. One may, of course, proceed more formally and say that \(f(x)\) is a given field on \(y=0\), without going further into the question of the relation of \(f(x)\) to the structure of the object.

If, for simplicity, we assume that the object has a periodic structure with period \(d\), then \(f(x)\) may be represented in the form of a Fourier series

\[ f(x)=\sum_{n=-\infty}^{+\infty} c_n e^{inKx}, \quad \text{where} \quad K=\frac{2\pi}{d}. \tag{11} \]

Obviously, the field behind the object (\(y>0\)) will also be periodic in \(x\) with period \(d\) and, as is not difficult to verify, will be expressed by the series

\[ \Phi(x,y)=\sum_{n=-\infty}^{+\infty} c_n e^{i\left(nKx+\sqrt{k^2-n^2K^2}\,y\right)}. \tag{12} \]

Indeed, \(\Phi(x,y)\) satisfies the two-dimensional wave equation and, at \(y=0\), becomes \(f(x)\).

But expression (12) for \(\Phi(x,y)\) is nothing other than a set of plane waves; the \(n\)-th wave travels at an angle \(\theta_n\) to the optical axis \(y\), where

\[ \sin \theta_n=\frac{nK}{k}=\frac{n\lambda}{d}, \]

and the amplitude and phase of this wave are expressed by the coefficient \(c_n\), which is wholly determined by the transmittance of the object \(f(x)\). Thus, each complex harmonic of \(f(x)\) gives its own plane wave. For \(|n|<N=\dfrac{d}{\lambda}\), these are ordinary traveling plane waves with real angles \(\theta_n\). In the principal focal plane \(F\) of the lens \(L\), they are focused into separate lines—the spectra of the diffraction pattern. One may say that in the “primary image” on the plane \(F\) the Fourier decomposition of the transmittance \(f(x)\) is physically realized, but only as limited by harmonics whose indices satisfy \(|n|<N\). Waves with \(|n|>N\) have a different character: they propagate parallel to the \(x\)-axis, and their amplitudes expo-

nentially decrease with \(y\). Such waves do not pass through the lens and take part neither in the formation of the diffraction pattern nor of the image.

From the points of the “primary image” the light beams with \(|n|<N\) again diverge (Fig. 10) and, precisely on the plane \(F'\) conjugate with the object plane, overlap edge to edge, being superposed with the same amplitudes and phases (the same \(c_n\)) with which they left the object plane \(y=0\). Thus, on \(F'\), up to scale, the field \(f(x)\) truncated on the side of the high harmonics is reproduced. If \(d<\lambda\), then even the first beams will be cut off; only the zeroth will pass, and on \(F'\)—even with such an ideal instrument as an infinitely thin lens—uniform illumination will be obtained: light with wavelength \(\lambda\) cannot give an image of so fine a structure. This fundamental limit of the resolving power of an instrument, imposed by the wave nature of light, was pointed out by Abbe and Helmholtz.

Fig. 10.

Fig. 10.

In any real optical instrument the aperture is finite, which can lead only to an still stronger limitation of the series (11), i.e. not only evanescent waves will fail to pass, but also part of the traveling waves for which \(|\theta_n|\) exceeds some angle \(\alpha<\dfrac{\pi}{2}\). Let us denote the function represented, correspondingly, by the truncated series \(f(x)\), by \(\tilde f(x)\). Then the light oscillation on the plane \(F'\) (up to the scale in \(x\)) will be \(\tilde f(x)e^{i\omega_0 t}\), and the distribution of its intensity will be proportional to

\[ \overline{\left[\operatorname{Re}\left(\tilde f(x)e^{i\omega_0 t}\right)\right]^2}^{\,t} = \frac{\tilde f\,\tilde f^{*}}{2}. \]

Thus, in the image we obtain the illumination

\[ I\sim \tilde f\,\tilde f^{*}. \tag{13} \]

Comparing with one another the basic propositions of the theory of the reception and demodulation of modulated oscillations in radio and the theory of the formation of an optical image, it is not difficult to notice a far-reaching parallelism between the two problems.

The transmittance of the object \(f(x)\) plays the same role as the modulating function \(f(t)\); changes of the amplitude \(A(x)\) caused by absorption are analogous to the amplitude modulation \(A(t)\), while changes of the phase \(\varphi(x)\), caused by variations in the thickness and refractive index of the object, are analogous to the phase modulation \(\varphi(t)\). Thus, pure amplitude modulation in radio corresponds to absorption structures in optics, and pure phase modulation to refraction or phase structures. The “primary image”—diffraction spectra with complex amplitudes \(c_n\)—plays the role of the carrier and side frequencies in the spectrum of the modulated oscillation. The finite aperture of an optical instrument, or any other changes introduced into the “primary image,” are analogous to the pass band of the selector and to the action of “discriminators” in radio. Finally, quadratic demodulation at the receiver output is the transition to the intensity of light “at the output” of the optical instrument. One may say that the quadratic detector and the telephone in a radio receiver do the same thing as the eye or a photocell in an optical instrument*).

Of course, such a far-reaching similarity holds only under certain limitations, which must not be lost sight of. There are not only similarities but also differences, and not merely external and obvious ones, but deeper ones—differences in the laws. They are ultimately due to the difference in the physical nature of temporal and spatial problems.

First, as has already been noted, in the case of oscillations in time, positive and negative frequencies are physically indistinguishable. In the spatial “oscillation” \(f(x)\), however, positive and negative frequencies (i.e., wave numbers) are essentially different, since they correspond to different directions of propagation of the waves that have passed through. Therefore, in the temporal problem, the complex Fourier series for \(f(t)\) was a convenient (and not always legitimate) method of notation. Here, however, in the spatial problem, such notation is adequate to the physical aspect of the matter.

Second, evanescent waves from structures smaller than the wavelength have no analogue in the temporal problem. In radio modulation there is no fundamental limit whatsoever for the modulation frequency \(\Omega\), i.e., a limit imposed by the very nature of the oscillations. We take \(\Omega \ll \omega_0\) on the basis of certain conditions of a practical kind, connected with the transmission of signals and with the corresponding apparatus. Oscillations with \(\Omega \gg \omega_0\) are also entirely possible. In optics, however, we can reach structures with a period \(d\) of the order of \(\lambda\), i.e., up to “frequencies” \(K\) of the order of \(k\). For \(K > k\), fundamentally

*) A photographic plate responds not simply to intensity, but accumulates the action over time.

the very propagation of light behind the structure changes, independently of what optical instruments are placed there.

But if, in the optical problem as well, we restrict ourselves to the case of sufficiently large and smooth structures \((K \ll k)\), then one may forget about the existence of a fundamental limit of resolving power, since then \(|nK|\) becomes comparable with \(k\) only for such large orders of the spectra \(|n| = N\) for which, practically, \(c_N \approx 0\). We may then with full justification say that the object spatially modulates the light wave passing through it, and consistently trace the identity between the laws of this spatial modulation in optics and the laws of temporal modulation in radio.

Bearing this in mind, let us now turn to the question of what we shall see for various structures of the object and various methods of observing (forming) the optical image.

3. IMAGES OF AMPLITUDE AND PHASE STRUCTURES

Let us first suppose that the aperture is so large that it practically does not in any way spoil the “modulating function,” i.e.

\[ \widetilde{f}(x)=f(x)=A(x)e^{i\varphi(x)}. \]

Let us emphasize once more that the function \(f(x)\) is now assumed to be so coarse and smooth that one may take into account only the aperture distortions of the field (12) in the “primary image.” According to (13), the intensity distribution in the image will be:

\[ I \sim A^2(x). \]

We shall therefore see the amplitude (absorption) structure of the object, while the phase structure will not appear at all. If the object is purely refractive, transparent, i.e. \(f(x)=e^{i\varphi(x)}\), then \(I=\mathrm{const}\)—the field of view will be illuminated uniformly. Thus, “just like that,” without additional measures, the refractive structure is invisible, just as phase modulation is inaudible in radio. And here, in optics, this by no means signifies that refractive structures produce too weak a diffraction, as a result of which the image supposedly “drowns” in the bright light of the zero beam. Such a view of the matter formerly had some currency among microscopists, but, as has been said, it is incorrect. Purely phase structures, like phase modulation in radio, can produce an exceedingly large number of intense “satellites” (diffraction spectra), but the amplitude-phase relations in them are unfavorable for “quadratic demodulation.” The diffracted beams, interfering with one another, give on the image plane \(F'\) the same uniform distribution as on the object itself.

A beautiful illustration of what has been said may be furnished by the diffraction of light by ultrasonic waves. Even at a low intensity of ultrasound it is easy here to obtain an intense diffraction pattern with a large number of spectra; however, if the ultrasonic column being illuminated is sufficiently thin and the wavelength of the ultrasound is not too small, then the structure is purely phase, refractive, and therefore invisible4.

But, in speaking of phase structures, one should by no means imagine only transparent objects which modulate the light passing through them in phase owing to inhomogeneities of thickness and refractive index. The concept of a phase structure is broader than that of a refractive one. In the well-known Foucault shadow method for testing

Fig. 11.

Fig. 11.

Fig. 12.

Fig. 12.

the sphericity of concave mirrors5, one is likewise dealing with a phase structure, although it is obtained as a result of reflection and not refraction. The light traveling from the mirror into the observer’s pupil (Fig. 11) is reflected at all points of the mirror surface with the same amplitude, but small deviations of this surface from sphericity create, in the reflected wave, deviations from the correct phase, i.e. they “crumple” or “warp” the front of the reflected wave. The “knife” \(K\) is precisely what serves to reveal this phase modulation; without it the structure (the irregularities of the mirror) remains invisible. In what exactly the action of the “knife” consists it is useful to consider somewhat further below.

Another kind of reflecting phase structure was recently realized by the author together with I. L. Fabelinskii6. On the reflecting face of a total-reflection prism, by evaporation in vacuum through a special mask, a grating was deposited in the form of strips of silver separated by empty gaps. If light is incident on such a grating from the glass side (Fig. 12), then the “jump” of phase in the reflected wave is different on the silver and on the regions of total internal reflection, while the amplitude is practi-

...tically identical. A typical phase structure is obtained, which, under uniform intensity, gives a “corrugated” front of the light wave, shown in the right-hand part of Fig. 12. In a photograph of the prism taken in reflected light in this way, the phase structure (grating) is almost invisible (Fig. 13), whereas in the principal focal plane of the objective a diffraction pattern is obtained (Fig. 14, a), moreover one that is more intense than in the reflection of light from the grating on the air side (Fig. 14, b), when it acts as an amplitude structure and reflects a smaller quantity of light.

What, then, must be done to make the phase structure visible? Obviously, as in the case of phase modulation in radio, it is necessary to introduce such changes into the “primary image” that the square of the modulus of the distorted transmittance function \(\widetilde{f}(x)\) (i.e., the intensity of light in the image plane \(F'\)) should depend on the modulated phase \(\varphi(x)\). This is precisely what is done in optics, using various techniques for this purpose.

Fig. 13
Fig. 13.

Fig. 14
Fig. 14

We have seen that the carrier-suppression method makes it possible to hear phase modulation. There exists an analogous optical method—the so-called dark-field method, which consists in eliminating the zero-order diffraction beam from the “primary image.” This can be achieved simply by blocking the zero-order beam with a special stop placed in the principal focal plane \(F\) of the objective. Another way consists in using a special system of object illumination (cardioid condenser), but since it falls outside our scheme (the illuminating wave is not plane), we shall confine ourselves to the first method.

In the case of phase structures, a drawback of the dark-field method is the possibility that the period of the visible image may be reduced by a factor of two in comparison with the true period of the structure [cf. the remark to formula (8)]. In the case of amplitude objects this method gives an identical image of complementary structures and may lead to only bright contours being visible—the boundaries between transmitting and absorbing regions of the object.

Indeed, if the transmittance of the object is \(0 \leq A(x) \leq 1\) and, consequently, \(c_0=\overline{A(x)}\) (the mean value), then in dark field we have:

\[ \widetilde{f}(x)=A(x)-\overline{A(x)}, \]

whence it follows that the distribution of intensity in the image will be:

\[ I\sim (A-\overline{A})^2. \]

A complementary structure is one whose transmittance is \(1-A(x)\), and hence \(c_0=1-\overline{A(x)}\). In this case

\[ \widetilde{f}(x)=1-A(x)-[1-\overline{A(x)}]=\overline{A(x)}-A(x), \]

from which it is clear that the intensity distribution obtained will be the same. If the object consists of transparent \((A=1)\) and opaque \((A=0)\) regions, then the former will have, in the image, intensity \(I\sim(1-\overline{A})^2\), and the latter \(I\sim\overline{A}^{\,2}\). For \(\overline{A}=\frac{1}{2}\) the intensity of both is the same, and only the contours of the boundaries between them will be visible.

However, when the question is only to see the object and not to discern its structure (for example, in ultramicroscopy), the merits of the dark-field method are undoubted and generally known.

Recently the dark-field method for observing nonmicroscopic objects (striae in glass, heat flows, etc.) was considerably improved by S. M. Raiskii. In the method proposed by him, the source is a grating or mesh \(S\) (Fig. 15), through which, owing to its considerable area, a large quantity of light can be transmitted. The condenser \(L_1\), behind which the observed object \(O\) is located, gives in the plane \(S'\) the “primary image,” i.e. the image of the mesh \(S\). In the usual technique the source of light is a slit or an aperture of some other shape, and in \(S'\) a diaphragm of the corresponding shape is placed, closing the image of the source “edge to edge.” Hence there naturally follow high requirements on the quality of the image of the source, i.e. on the optical system \(L_1\), and—if

the object is a liquid in a vessel—so also do the optical properties of this vessel. In Rayleigh’s method all these conditions disappear, since the diaphragm is replaced by a negative image of the grid \(S\), obtained by exposing a photographic plate placed precisely in the plane \(S'\). All the defects of the image of the grid \(S\), due to the poor objective \(L_1\) and the poor vessel at \(O\), are also present on the negative. With the negative correctly positioned in the plane \(S'\), the intensity behind it is equalized, i.e. the “primary image” is compensated. Consequently, in the plane \(F'\), conjugate with respect to the objective \(L_2\) to the object \(O\), there will be no image of the object; the vessel with all its defects will become invisible. An image will be produced only by the additional inhomogeneities of the object that interest us (for example, flows or ultrasonic waves in the liquid filling the vessel), which will cause

Fig. 15.

Fig. 15.

additional distortions of the “primary image” in \(S'\), which are absent on the negative placed there. Thus, only the objective \(L_2\) must have good optical qualities. Among other advantages of this method one should point out the comparative ease of adjustment and the high aperture ratio of the device. Of course, with respect to the similarity between the image and the structure of the object itself, when this structure is phase-like, the described modification of the dark-field method has no advantage.

Another possible method of detecting phase modulation in radio, as we have seen, consisted in cutting off half the spectrum, including or excluding the carrier. An analogous method has long been known in optics—this is the above-mentioned Schneidenverfahren (cutting method) of Foucault (1859), carried out either, as by Foucault himself and later by Toepler\(^8\), with the aid of a “knife,” or, as is done in microscopy practice, by means of oblique illumination of the object.

Toepler developed his stripe method for observing “macroscopic” phase structures (schlieren, flows, etc.). The scheme of his arrangement is the same as in Fig. 15, but the light source is an aperture with a straight-line edge, and in the plane \(S'\) a “knife” is placed whose “blade” is parallel to the image

of this edge. In Fig. 16 a diagram of Töpler’s arrangement with a source in the form of a horizontal slit is shown. If the “knife” is placed close against the image of the slit, but does not cover it, then in the presence of the object \(O\) one half of the “primary image” will pass, including the zero-order beam. Further movement of the “knife” toward the zero-order beam will lead to a gradual darkening of the field of view, to the cutting off of half of the spectra together with the “carrier.” With oblique illumination of the object in the microscope, i.e., with oblique incidence of a plane wave on the object \(O\) (Fig. 10), the diffraction pattern in the plane \(F\) is correspondingly shifted upward or downward. If one imagines that a diaphragm has been placed in this plane, then, as the angle of inclination of the illuminating wave is increased, the zero-order spectrum will approach the edge of this diaphragm and then go beyond it; that is, here the same change of the “primary image” is effected as in Töpler’s method.

Fig. 16.

Fig. 16.

As in the dark-field method (covering only the zero-order beam), in the described “Foucault knife” method, which is often also called the shadow method, the distribution of intensity in the image of the object differs greatly from the structure of the latter. In this case a peculiar phenomenon is observed, noted already by Töpler: if the zero-order beam is not completely covered, then in the observed distribution of illumination an asymmetry is obtained, creating the impression of relief. The object appears in the image as an uneven surface illuminated from the side. With the zero-order beam completely covered, this effect disappears—a symmetric structure gives a symmetric distribution of intensity in the image. However, even in this case the reproduction of the structure is not similar, and moreover, as in the dark-field method, owing to the covering of the zero-order beam a large loss in the amount of light is obtained.

Among other purely optical methods of detecting phase structures, one should also mention defocusing of the optical system, but it is advisable to dwell on this method later. Finally, a method very straightforward in idea—the so-called forehead method—is possible and is indeed used in biological microscopy.

method of making transparent specimens visible is to stain them. But in biology this is far from the best solution to the problem, since chemical treatment for the most part kills living specimens or else changes their structure. The method of phase contrast, to which we shall now turn, is free of all the enumerated shortcomings.

4. THE METHOD OF PHASE CONTRAST

Among the possible radio methods of detecting phase modulation under square-law detection, the change of the phase of the carrier by \(\pm 90^\circ\) was indicated above. The method of phase contrast, proposed in 1934 by the Dutch physicist Zernike[^9], is an optical analogue of such a procedure. It amounts to the fact that the zero-order diffraction beam is not removed from the “primary image,” but only the phase of this beam is changed by \(\pm 90^\circ\)*). Fig. 7 explains how, in this case, phase modulation is transformed into amplitude modulation, and this remains valid, of course, also for spatial modulation in optics. Equally, formulas (9) and (9′), if \(t\) is replaced in them by \(x\), give the spatial distribution of intensity in the image of the structure.

Initially Zernike proposed the method of phase contrast as a substitute for the “shadow method” for testing astronomical mirrors. In this form this method had already been used since 1938 at the State Optical Institute, and for testing not only mirrors but also objectives[^10]. With approximately equal sensitivity, the advantage of phase contrast in comparison with the “Foucault knife-edge” method is that it gives not a “shadow” picture, but directly shows the humps and depressions on the surface under investigation, facilitating corrections during polishing. However, as D. D. Maksutov convincingly showed[^11], the possibilities of the “shadow method” are extremely varied (not only control of sphericity, but also measurement of radii of curvature, aberrations, astigmatism, field curvature, the quality of flat and aspherical surfaces, investigation of inhomogeneities and striae in glass, etc.). It is not excluded that the method of phase contrast, which is not mentioned in the cited work of D. D. Maksutov, has advantages also in some other of the listed applications of the “shadow method.” The idea of extending phase contrast to all cases of observation of phase “macro” objects was expressed by L. M. Mandelstam soon after the appearance of Zernike’s first reports. Only much later

*) Academician S. I. Vavilov informed the author that analogous ideas had been expressed as early as 1932 by D. S. Rozhdestvensky. Unfortunately, these considerations were not published.

and in a greatly limited form, but with quite positive results, some experiments of this kind were carried out at FIAN\({}^{12}\).

In 1935 Zernike described the phase-contrast method as applied to the microscopic observation of refractive objects\({}^{13}\). A few years later specially equipped microscopes were produced\({}^{14}\), which at the present time are also being manufactured by the domestic optical industry.

How, then, is the change in phase of only the zero diffraction beam by \(\pm 90^\circ\) achieved? In the case of a periodic structure of the object under observation, giving in the principal focal plane of the objective non-overlapping discrete spectra, this problem is solved quite simply. In this plane a plane-parallel glass plate is placed, having a thickening or, conversely, a thinning at the place where the zero spectrum passes (Fig. 17).

Fig. 17.

Fig. 17.

The width of the projection or depression on this phase plate must be such that the zero spectrum fits entirely on it, while the nearest \(\pm 1\)st orders already pass by it. Then the zero beam acquires a retardation or advance in phase by

\[ |\Delta\varphi|=\frac{2\pi(n-1)\delta}{\lambda}, \tag{14} \]

where \(n\) is the refractive index of the glass, and \(\delta\) is the difference in thicknesses (Fig. 17). Requiring

\[ |\Delta\varphi|=\frac{\pi}{2} \]

and taking \(n\approx 1.5\), we obtain:

\[ \delta=\frac{\lambda}{4(n-1)}\approx\frac{\lambda}{2}. \]

For questions of the technology of manufacturing such plates, see, for example,\({}^{10}\) and the literature cited there.

Fig. 18.

Fig. 18.

Zernike\({}^{13}\) calculated, according to Abbe, the intensity distribution in the image of a diffraction grating, the profile of which is shown in the upper part of Fig. 18 (the phase of the transmitted wave is modulated by \(30^\circ\)), for various methods of observation and illumination. These graphs make it possible to judge the theoretically expected advantages

methods. Curves 1 and 2 refer to the case of central illumination in the presence of defocusing: in 1 the image is considered in a plane somewhat advanced forward (toward the objective) from the plane \(F'\); in 2, displaced backward, the displacement being 8 times greater than in case 1. The intensity distribution departs from uniformity, but in neither case can one speak of any similarity to the structure of the object. Let us note that the method of defocusing has no analogue in the case of reception of phase modulation in radio, and this again is connected with the multidimensional character of the wave (spatial) problem, in contrast to a “one-dimensional” oscillation in time. The structure modulates the wave along the coordinate \(x\), while defocusing is effected along the optical axis of the system—the coordinate \(y\), for which, in the case of oscillations in time, there is no equivalent. True, the final result can also be reproduced in radio, but in a very artificial way—by introducing specially selected phase shifts into each component of the spectrum of the modulated oscillation.

Curve 3 still refers to central illumination, but with the zero beam blocked (dark field). Along with a generally large loss of light one may observe the presence of dark borders at the boundaries of the “thick” and “thin” portions. Curves 4 and 5 correspond to oblique illumination (the strip method); 4 refers to the case where the zero beam still passes, which gives a relief effect; 5 is obtained when the zero beam is cut off—the relief effect is not observed, but the amount of light drops sharply. In both cases there is no similarity between the image and the structure. Finally, curves 6 and 7 illustrate the method of phase contrast: 6—the so-called positive contrast (the zero beam receives a phase advance of \(90^\circ\)); 7—negative (a retardation of the zero beam by \(90^\circ\)). In both cases a contrasty and similar image is obtained while the full luminous flux is preserved.

In practice, however, one has to give up preserving the full amount of light, owing to two circumstances. The first is connected with the form of the diffraction spectra. In the principal focal plane of the microscope objective an image of the aperture diaphragm of the condenser is obtained. In the presence of an object, for example a grating, spectra, which are displaced images of this diaphragm, are superposed upon one another, and the more so the larger the period of the structure (Fig. 19). What has been said applies to an even greater degree to nonperiodic structures. If the diaphragm is a circular aperture, as is the case in Fig. 19, then the area of overlap of the diffracted circles with the zero circle is large, i.e. the action of the phase plate will to a considerable extent affect not only the zero beam, but also the lateral beams nearest to it. This substantially reduces the contrast of the image. More favorable conditions are obtained

when using an annular aperture diaphragm (and, correspondingly, an annular phase plate), since, for the same distance \(r\) between the centers, the overlap area \(S\) is substantially smaller (Fig. 20). The improvement in contrast is so significant that one is forced to accept a reduction in the total amount of light.

Fig. 19.

The second circumstance leading to a further loss of light is again connected with a gain in contrast. The ring on the phase plate is made slightly absorbing. Owing to this, the modulated vector is reduced while the magnitude of the modulating vector remains practically unchanged, i.e., the depth of amplitude modulation (contrast) is increased still further.

Fig. 20.

Figs. 21–25 illustrate the application of the phase-contrast method to several specimens. Phase contrast makes it possible to reveal smaller differences in refractive index than the dark-field method or defocusing (see Fig. 24). It can therefore provide higher accuracy in measuring the refractive index of micro-objects by the well-known method of immersing them in a medium (solution) whose refractive index is varied (by changing the concentration of the solution or the temperature) until the disappearance of

Fig. 21. Live trypanosomes in mouse blood.
\(a\)—bright field, \(b\)—positive phase contrast (\(\times 200\)).

Figure 22

a            b

Fig. 22. Cell of the epithelium of the mucous membrane of the cheek. a — bright field, b — positive phase contrast (×500).

Figure 23

a            b

Fig. 23. Erythrocytes and a leukocyte (in the middle) in fresh blood. a — bright field, b — positive phase contrast (×1000).

Figure 24

a            b

Fig. 24. Glass powder in liquid. a — bright field, b — negative phase contrast (×80).

images of immersed particles. It is also of interest to investigate the possibility of applying phase contrast in metallographic microscopes*). Finally, in principle this method can also be extended to electron microscopy^16.

The implementation of a corresponding device would make it possible to obtain contrast electron-microscopic images of phase (electron-transparent and weakly scattering) objects and, possibly, in a number of cases would replace the method of depositing metal on the specimen. However, difficulties arise here connected with the extremely small de Broglie wavelength. At an electron energy of 100 kev it is about 0.04 Å. If the characteristic dimensions (or period) of the structure under consideration are of the order of 20 Å, then the diffraction angle of the 1st spectra will be 0.002 radian. With the principal focal length of the objective equal to 3–5 mm, this gives a distance between the spectra of \(6—10\cdot 10^{-3}\) mm, i.e. the spot (thickening) on the “phase plate” must be no greater than the indicated value in diameter. As for the thickness of the phase spot, it can be estimated in the following way. The refractive index of electron waves is equal to

Fig. 25. Chromosomes in a cell nucleus. a — bright field, b — positive phase contrast.

Fig. 25. Chromosomes in the cell nucleus. \(a\) — bright field, \(b\) — positive phase contrast.

\[ n=\sqrt{1+\frac{V_1}{V}}\simeq 1+\frac{V_1}{2V}, \]

where \(V\) is the energy of the electrons in the beam (for example, 100 kev), and \(V_1\) —

*) On various applications of the phase-contrast method, see ^1,14,15.

the internal potential of the body is a quantity of the order of several volts (for example, 12 eV in carbon). Thus, \(n - 1 \simeq 10^{-4}\), and formula (14), if one puts \(|\Delta \varphi| = \dfrac{\pi}{2}\), gives \(\delta \simeq \dfrac{\lambda}{4}\cdot 10^4\), i.e. a value of the order of \(100\) Å. These figures show that, from the technical point of view, the transfer of the phase-contrast method into the field of electron microscopy is not a simple problem.

In conclusion it is perhaps worth noting the following. Radio engineers developed phase modulation and methods for its reception without suspecting any phase contrasts in optics. In the same way, all opticians who wrote about the phase-contrast method did not say a word about phase modulation in radio. The theory of oscillations overcomes the partitions that exist by virtue of the historically formed and traditionally inherited separation of the different “departments” of physics, and reveals the commonality of laws in phenomena in which, to outward appearance, there would seem to be only differences. This characteristic and, to the highest degree, valuable feature of the theory of oscillations was repeatedly used, broadly developed, and always emphasized in his lectures by one of the founders of the modern theory of oscillations, Academician L. I. Mandelstam. And the aim toward which the author of this article has striven consists not only in giving an idea of the phase-contrast method and its advantages, but also in illustrating, by this concrete example, the indicated feature of the theory of oscillations.

CITED LITERATURE

  1. E. V. Shpolsky, New methods in microscopy, UFN 32, 376 (1947).
  2. Rayleigh, On the theory of optical images, with special reference to the microscope, Sci. Pap., vol. IV, 235 (1896).
  3. E. Abbe, Gesamm. Abh., vol. I (Jena, 1904).
  4. See, for example, S. M. Rytov, Diffraction of light by ultrasonic waves, § 6. Izv. AN SSSR (ser. phys.), No. 2, 223 (1937).
  5. L. Foucault, Mém. sur la construction des télescopes en verre argenté, Ann. de l’observ. imp. de Paris, vol. V, p. 203.
  6. S. M. Rytov and I. L. Fabelinskii, ZhETF 20, 340 (1950).
  7. S. M. Raiskii, ZhETF 20, 378 (1950).
  8. A. Toeppler, Pogg. Ann. 131, 33 (1867).
  9. F. Zernike, Monthly Not. of the Roy. Astr. Soc. 94, 377 (1934); Physica 1, 689 (1934). See also C. R. Burch, Monthly Not. 94, 384 (1934) and 95, 548 (1935).
  10. Yu. V. Kolomiitsev, Optical-Mechanical Industry, Nos. 8 and 9 (1938).
  11. D. D. Maksutov, Optical-Mechanical Industry, No. 5, 3 (1941).
  1. S. M. Rytov and M. E. Zhabotinskii, Journ. of Phys. (USSR) 11, 92 (1947).

  2. F. Zernike, Phys. Zeits. 36, 848 (1935); Zeits. f. techn. Phys. 16, 454 (1935).

  3. A. Koehler and W. Loos, Naturwiss., No. 4, 49 (1941).

  4. J. Picht, Zeits. f. Instr-nde 56, 363 and 481 (1934) and 58, 1 (1938); K. Michel, Naturwiss., No. 4, 61 (1938); W. Loos, W. Klemm and Smekal, Naturwiss., No. 50.51, 769 (1941); B. O. Payne, J. Sci. Instr. 24, 163 (1947); M. Françon, Mikroskopie 1, 117 (1949); F. J. Keuning, Mikroskopie 5, 49 (1950); H. Wolter, Ann. d. Physik 7, 33 (1950).

  5. D. Gabor, The electron microscope (New York, 1948), p. 54.

  1. Report read at the colloquium of the P. N. Lebedev Physical Institute of the Academy of Sciences of the USSR. 

Submission history

On the Phase Contrast Method in Microscopy\*