A NEW THEORY OF HEARING
H. Fletcher
Submitted 1931 | SovietRxiv: ru-193101.24523 | Translated from Russian

Abstract

Presented at a meeting of the Acoustical Society of America.

Full Text

A NEW THEORY OF HEARING

Harvey Fletcher *

In constructing a theory of hearing, the following experimental data, which admit of quantitative evaluation, must be taken into account: 1) the limits of pitch and sound intensity perceived by the auditory apparatus, 2) the minimal changes in pitch and sound intensity noticed by hearing, 3) subjective tones, 4) the masking of one sound by another, 5) the relation of loudness to the level of sensation, and 6) the binaural effect. In addition to these data there is also a series of other observed facts of great importance which, however, do not admit of quantitative evaluation. Chief among them are: 1) the ability of the ear to analyze a complex sound by resolving it into components, 2) changes in the pitch, loudness, and timbre of a complex sound when some of its components are excluded, 3) the relation of “brightness” and “volume” ** to the frequency and intensity of a sound, and 4) the influence of phase shift of the components on the quality of a consonance. ***

The theories of hearing proposed to explain these experimental data fall into two classes. Some of them may be called theories of spatial “patterns” or “sequences” (space pattern theory), while others—temporal “pattern” theories or

* H. Fletcher, A space time pattern theory of hearing, Jour. Acoust. Soc. of Amer. 1, p. 311, 1930. Reported at a meeting of the Acoustical Society of America. Translated by S. N. Rzhevkin.

** Concepts introduced by psychologists. See L. T. Troland, Jour. of gen. Psychol. 2, p. 28, 1929.

*** The reader may become acquainted with the indicated phenomena from S. N. Rzhevkin’s book Hearing and Speech and from the article of the same title in Uspekhi Fizicheskikh Nauk, vol. VII, p. 231, 1927.

“sequence” theories (time pattern theory). In the first theories it is assumed that the temporal patterns or sequences of the wave motion of the air are transformed into a certain spatially arranged process (pattern) in the inner ear, and the representation of the temporal “pattern” of the wave (frequency) is created in the brain because a definite region of the inner ear is excited, which then sends nerve impulses to the brain. In theories of the second class it is assumed that the temporal patterns of the wave are transmitted directly to the brain. In the author’s opinion both processes take place simultaneously and create a complete representation of the sound being perceived. The theory proposed by the author can best be called a theory of hearing based on spatio-temporal “patterns” or “sequences” (space-time pattern theory).

When a sound wave passes through the air, the particles of the air oscillate backward and forward and produce small changes in pressure, occurring in a known sequence over time, determined by the character of the sound. If some small pressure indicator, instantaneously following all changes of pressure, is placed in the path of the wave, then its readings as a function of time will be expressed by curves similar to those shown in Fig. 1. Each curve characterizes the “time pattern” of the sound to which it pertains. This pattern is transmitted to the tympanic membrane of the ear and, through the chain of auditory ossicles and the oval window, reaches the fluid of the labyrinth. During this transmission the amplitude of the pressure is increased approximately 60-fold owing to the lever action of the auditory ossicles, and also because the total force acting on the tympanic membrane is distributed over the much smaller area of the oval window.* If we knew all the mechanical properties of the system, we could precompute all possible changes of the temporal pattern

* This is true only if the tympanic membrane oscillates as a whole, without being broken up by nodal lines.

pressure transmitted into the cochlea. For low sound intensity and for frequencies in the middle range, the time sequence of pressure in this transmission probably changes only slightly; components lying in the low and high ranges are apparently transmitted with attenuation. Thus, in any case, a temporal pressure pattern is created in the cochlear fluid near the oval window, similar in form (though not entirely exactly) to the pattern existing in the air.

Fig. 1

Fig. 1

Nonlinear characteristics of the middle ear

If the intensity of sound increases, other conditions being equal, then the temporal pressure patterns in the cochlear fluid change not only in magnitude but also in form. For a high sound intensity, the movements of the elastic membranes of the middle ear pass beyond the range in which the displacement is proportional to the acting force. Moreover, the elasticity of the membranes is asymmetrical under compression and expansion. It is known that under these conditions there arise frequencies other than those present in the original sound. Thus, for example, at high sound intensity a pure tone of frequency \(f\), when transmitted

in the cochlea it is transformed into a tone containing the frequencies \(f, 2f, 3f, 4f\), etc. The components added to the sound in the process of transmission through the middle ear are called subjective tones; their amplitude can be measured experimentally. Let us present to the ear a pure tone of a known level of intensity,* and at the same time present a second tone differing in frequency from the harmonic under investigation by 2–3 cycles. Then beats will arise between the subjective harmonic under investigation and the second tone. If the intensity of the second tone is chosen so that the beats become most distinct, then its intensity will serve as a measure of the intensity of the subjective harmonic. In Fig. 2 the intensity of subjective harmonics, measured in this way, is shown for exciting tones of different intensities; these curves are constructed on the basis of measurements made by Graham. The abscissae show the number of the harmonic (the fundamental tone is the first harmonic), and the ordinates show the intensity level of the subjective tones. For example, a tone whose fundamental harmonic is at the zero level \((1\,\mu\text{W}/\text{cm}^2)\) will give subjective harmonics lying at the levels \(-1.4\), \(-2.6\), \(-3.7\), and \(-4.8\) bels respectively for the second, third, fourth, and fifth harmonics. An exact determination of the intensity of subjective tones is difficult; however, within the limits of observational errors and individual variations it has been found that the series of curves presented in Fig. 2 characterizes the occurrence of subjective harmonics for tones of any fre-

Fig. 2. Graph with axes: “subjective harmonics”; vertical axis “level relative to fundamental, in bels”; horizontal axis “order of harmonic.”

Fig. 2.

* The level of sound intensity is determined relative to the conventional, so-called zero level of \(1\,\mu\text{W}/\text{cm}^2\) and is measured in bels or decibels. 1 bel corresponds to an intensity level of \(10\,\mu\text{W}/\text{cm}^2\), 2 bels to \(100\,\mu\text{W}/\text{cm}^2\), 3 bels to \(1000\,\mu\text{W}/\text{cm}^2\); \(1\) bel \(= db\); \(db\) = decibel. —Trans.

...frequency. Such a result was to be expected, since the overload of the middle-ear mechanism depends on the physical force of the incident sound, and not on its sensation level.*

Further, one may expect that the overload should depend on the amplitude of vibration of the tympanic membrane and of the auditory ossicles connected with it. In agreement with the experimental data just cited, one must suppose that the (mechanical) impedance of the tympanic membrane must have the character of an elastic reaction, and not of resistance or mass. A recent paper by Tröger** shows that the impedance of the ear has precisely this character, except at high frequencies. It is important to emphasize that the curves in Fig. 2 refer to levels of sound intensity, and not to sensation levels. If the intensity of the subjective tones is expressed relative to the threshold of audibility, then, for example, for the tone in −4 octaves*** (i.e. 62.5 cycles), on the basis of the observations presented above, sensation levels of 4.4; 4.6; 4.4; 3.8 and 3.0 bels are obtained. It is thus clear that the second harmonic has a higher sensation level than the fundamental tone.

Fig. 3 presents the results of observations for four different intensity levels for a tone of 50 vibrations per second. It is clear from the curves that for the two upper curves the sensation level of the first five overtones is higher than for the fundamental tone. It should be stipulated that the data used to construct the curves were of small numerical extent and that measurements with the given ear give results which often may deviate considerably from the average curves presented. Owing to the indicated nonlinearity of the transmission of the middle-ear mechanism, the ratio of the amplitudes of pressure in the cochlea to the amplitudes of pressure in the air in front of the ear will be considerably smaller for strong sounds than for weak ones. As the intensity of the sound increases, the transmitted si-

* The sensation level is reckoned from the threshold of audibility and is measured in bels or decibels.

** Tröger, Phys. Zs. 31, p. 26, 1930, Jan.

*** The pitch of a tone, at Fletcher’s suggestion, is measured in octaves counted upward and downward from the tone of 1000 vibrations per second.
Translator’s note.

system becomes more and more overloaded, and consequently the output correspondingly decreases. For sound levels below −6 bels, the pressure in the cochlea is proportional to the acting pressure. For higher levels, transmission at the fundamental tone falls owing to the occurrence of overtones. If, during the oscillation, the displacement at the oval window were the same on both sides of the neutral position, then only even harmonics would arise. In view of

Fig. 5. Subjective harmonics of a tone at 50 cycles/sec. Vertical axis: sensation level, dB. Horizontal axis: order of the harmonic.

Fig. 5.

the fact that both even and odd harmonics are observed, it is clear that the oscillation is asymmetrical and that a rectifying action takes place. Thus the mechanism of the inner ear possesses both nonlinearity and asymmetry.

Mechanics of the Inner Ear

To simplify the picture, we shall first consider the character of the motions in the absence of subjective harmonics. Let a pure tone be transmitted through the middle ear into the fluid of the cochlea adjoining the oval window. The question is what the motion of the fluid in the cochlea will be, and in particular what the motion of the basilar membrane, in which the nerve endings are located, will be.

From the very beginning I should like to emphasize that a satisfactory solution of this problem has not yet been given, nor

is such, and the solution proposed below. Wegel and Lane*, in their work on the masking of tones, give a good approximation to the solution of the problem on the basis of an electrical analogy of the cochlea. However, they do not make numerical calculations because of the unreliability of the constants entering into the problem. Despite the fact that the simplified solution of the problem developed below may arouse many objections, it may nevertheless be of significance, since it helps one better to understand the mechanism of action of the cochlea.

According to Helmholtz’s resonance theory, the oscillation imparted by the stirrup to the fluid of the cochlea propagates for some distance, depending on the pitch, along the vestibular scala (scala vestibuli), then passes through the basilar membrane (membrana basilaris) and returns back along the tympanic scala (scala timpani) to the round window. The nerve endings are excited at the place where the oscillation passes through the basilar membrane, and send impulses along the auditory nerve, which conducts them to the brain.

Let us consider a small element of the basilar membrane of length \(dx\) and width \(w\), situated at a distance \(x\) from the oval (and round) window, Fig. 4. A sinusoidal pressure \(p\), applied to the oval window, is transmitted through the fluid and causes this element to oscillate. The problem is to determine the law of its oscillations. If the oscillations of the sound pressure are slow, then any surface of equal size along the basilar membrane experiences practically the same force at a given instant of time. Consequently, the forces on the two sides of the element under consideration are equal and opposite, as a result of which the element will be in equilibrium. With increasing frequency an excess pressure is obtained on the upper side of the element, as a result of which the element comes into motion. The force \(F\), caused by this excess pressure, will be

\[ F = k_f w\, dx\, P \cos \omega t, \tag{1} \]

* Wegel and Lane, Phys. Rev. Feb. 1924

where \(P\) is the pressure amplitude at the oval window and \(\dfrac{\omega}{2\pi}\) is the frequency of oscillation. The coefficient \(k_t\) is introduced on the basis of the considerations given above; it varies from unity at high frequencies to zero at very low ones.

The reaction forces opposing the applied force are caused by the inertia of the system, friction in the moving liquid, and the elasticity of the basilar membrane in that region through which the oscillations penetrate. The effective mass is due almost entirely to the liquid oscillating on both sides of the membrane element. It is natural to assume that the effective mass \(m\) is determined by the expression:

\[ m = k_m w\,\delta x \cdot x, \tag{2} \]

i.e., expressed in words, this means that the effective oscillating mass of the liquid is proportional to the distance of the element from the oval window or, in other words, proportional to the amount of liquid moving together with the membrane element. Obviously this would be correct if the motion were represented by the simplified scheme of Fig. 4; here the column of liquid, oscillating as a single whole, has an area equal to the area of the element \(w\delta x\), and extends from the oval window to the element over a distance \(x\) and back over the same distance to the round window. In the present case, obviously, \(k_m = 2\), since the density of the liquid is approximately equal to unity. Although the motion cannot be as simple as has just been assumed, nevertheless the general character of the motion caused by tones of different pitch is probably similar to that described and differs only quantitatively. Therefore we shall consider that equation (2) is correct in the first approximation.

Fig. 4.

Fig. 4.

The stiffness \(S\) is chiefly inherent in the element itself, although, of course, some stiffness is added owing to the reaction of the oval and round windows, and also of the auditory ossicles. The stiffness is proportional to the number of transverse fibers of the membrane passing through the element \(\delta x\), propor-

is proportional to the tension of the membrane \(T\) and inversely proportional to the length of the element \(w\), i.e., in other words, to the width of the membrane at the point \(x\). Thus:

\[ S=k_s\,\frac{\delta x\,T}{w}. \tag{3} \]

It is well known that such a system has a natural frequency determined by the expression:

\[ \omega_0^{\,2}=(2\pi f_0)^2=S/m=\frac{k_s}{k_m}\,\frac{T}{w^2 x}. \tag{4} \]

This expression gives the relation between the frequency \(f_0\) of the tone acting on the ear and the position \(x\) of the greatest excitation on the basilar membrane.

Let us consider the anatomical data in connection with expression (4) and try, on their basis, to determine the position of the greatest excitation for tones of different frequencies. The width \(w\) can be measured very accurately by means of microscopic sections. Here considerable individual deviations are observed; the mean values are given in Fig. 5, curve \(w\). We see that \(w\) changes from \(0.08\) to \(0.5\) mm as \(x\) changes from 1 to 31 mm.

Fig. 5.

Fig. 5.

Anatomical studies show that the tension \(T\) of the transverse fibers decreases as \(x\) increases, but there are no quantitative data in this respect. We assume the possibility of three different types of dependence of \(T\) on \(x\) and give in Fig. 5 the results of the corresponding calculations. In case I it is assumed that the tension \(T\) of the transverse fibers is constant along the entire length of the membrane. In case II it is assumed that the tension is рав-

decreases uniformly in the ratio 10 to 1 from the oval window to the helicotrema. In case III the variation of \(T\) was chosen so as to give a uniform distribution of pitch along the basilar membrane over a range of six octaves. The ratio \(k_s/k_m\) was determined so that, for the resonator at \(x=1\), a frequency of 8000 cycles/sec would be obtained.

From the anatomy of the cochlea it is evident that the position of greatest excitation for high tones must lie near the oval window and, as the tone is lowered, must shift toward the helicotrema*. We see from Fig. 5 that the required region of tone perception can be obtained without any improbable assumptions regarding the limits of variation of the tension of the fibers along the basilar membrane. Even if no variation of the tension \(T\) along the membrane is assumed, it is still possible to obtain a change of resonance within the limits from 275 to 8000 cycles/sec. Most authors who have touched upon this problem have assumed that the square of the resonant frequency is inversely proportional to the length of the transverse fibers \(w\). Under our assumptions, \(\omega_0^2\) proves to be inversely proportional to \(w^2\); this is due to the fact that, according to our assumption, the effective mass, determined by the co-vibrating fluid associated with the given elements of the membrane, decreases in proportion to the decrease of \(w\). In other words, the effective vibrating column has precisely the width \(w\) of the basilar membrane at the point of excitation. A tenfold change in \(T\) is sufficient to explain a change from 86 to 8000 cycles/sec. If we assume that the arcs of Corti begin already at \(x=0.5\) mm and that \(w\) varies as indicated above, then the range rises to 16 thousand cycles/sec. This, of course, is true on the assumption that equation (4) retains its force for small values of \(x\), which can hardly be regarded as entirely correct.

Determination of the minimally perceptible difference in pitch has shown** that the distribution of frequency along the basilar membrane is closest to case III. On this basis

* The helicotrema is the opening connecting the tympanic and vestibular canals, situated at the apex of the cochlea. Translator.

** Wegel and Lane, l. c.

In the further derivations the following simple relation is adopted, connecting the pitch \(P_0\) (or its frequency \(f_0\)) and the position of excitation \(x\) (in mm) on the basilar membrane*:

\[ P_0=\frac{16-x}{5}, \tag{5} \]

or

\[ \lg_{10} f_0=\frac{16-x}{16.6}+3; \tag{6} \]

this relation holds when \(x\) varies from 1 to 31 mm. According to this hypothesis, the positions of the resonance maxima are distributed along the basilar membrane uniformly (on a logarithmic scale) as the pitches of the exciting tone vary from \(-3\) to \(+3\) octaves (from 125 to 8,000 cycles/sec). Exciting tones whose pitches are above or below these limits do not create conditions for resonance in the cochlea. This may be one of the reasons why sensitivity falls above and below these limits.

Having established the relation between \(f_0\) and \(x\), let us consider the form of the vibrations of the basilar membrane. If we were dealing with individual stable resonators, and not with a system of mutually coupled unstable ones, the problem would be much simpler. But, as we have seen, this is not the case. The elasticity of the portion of the membrane corresponding to a given tone remains almost constant, but the effective mass co-oscillating with it, as well as the resistance \(r\), vary.

The continuation of the vibration after the exciting periodic force has ceased to act depends on the ratio of the resistance \(r\) to the mass \(m\), i.e. on the damping factor, which is given (in nepers per second) by the expression

\[ \Delta=0.434\,\frac{r}{m}. \tag{7} \]

It is difficult to make an exact determination of \(\Delta\) on the basis of the anatomy of the cochlea; the considerations given below clarify only those factors which influence damping.

* Here the pitch is expressed in octaves relative to the tone 1000 cycles/sec.; \(\lg_{10} f\); between \(f\) and \(P\) there is the relation: \(P\lg_{10}2+3\). —Trans.

A NEW THEORY OF HEARING

Let us consider a fluid tube of length \(2x\) and cross-section \(w\delta x\), moving slowly back and forth inside a wider tube filled with the same fluid; this is shown in section in Fig. 6. If the cross-section of such a tube is small, then the frictional force is proportional to the interacting surface, namely \(2x\cdot 2(w+\delta x)\), multiplied by the velocity; the proportionality factor will be equal to \(\eta/d\), where \(\eta\) is the coefficient of viscosity, and \(d\) is the effective distance from the surface of the inner tube to the outer walls. The magnitude \(d\) is approximately as shown in Fig. 6. The damping coefficient, under the condition that \(\delta x\) is small in comparison with \(w\), will be approximately equal to

\[ \Delta = 0.868\,\frac{\eta}{\delta x d}. \tag{8} \]

Fig. 6.

Fig. 6.

We see that \(\Delta\) does not depend on the length of the fluid tube, but depends on \(\delta x\). If the thickness \(\delta x\) corresponds to one arc of Corti, then \(\Delta = 30\) bel per sec.; if, however, it is taken equal to the width \(w\), then \(\Delta = 570\) bel per second.

It could be expected that the magnitude \(\Delta\) for the case of the element of the basilar membrane entering into our consideration would lie precisely within these limits. The reasoning given leads to the conclusion that \(\Delta\) should be regarded as independent of \(x\), i.e. the same for different points along the basilar membrane.

Using the expressions found for \(F\), \(m\), and \(S\), we can find the displacement \(y\) of the element \(\delta x\) from the position of equilibrium.

For sinusoidal motion:

\[ y = A\cos(\omega t-\Theta), \tag{9} \]

we find for the amplitude \(A\) the following value:

\[ A=\frac{1}{4\pi^{2}}\,\frac{k_s}{k_m}\, \frac{P}{x\left[\left(f_0^{2}-f^{2}\right)^{2}+\left(0.37 f\Delta\right)^{2}\right]^{1/2}}. \tag{10} \]

and for the phase angle:

\[ \tg \Theta = 0{,}37\,\Delta\,\frac{f}{f_0^2-f^2}. \tag{11} \]

It was indicated that the coefficient \(k_f\) changes from zero at very low frequencies to unity at high frequencies, and that \(k_m\) has a value close to 2. Taking \(k_f=1\); \(k_m=2\); \(\Delta=70\) and \(P=1\) bar, we calculate \(A\) and \(\Theta\) for various frequencies. The result of these calculations is given in Table I. The relation between \(f_0\) and \(r\) is taken in the form of equation (6).

TABLE I

\(f\) \(x\) \(A\cdot 10^8\) \(\Theta\)
75 10 0.2 \(0^\circ\)
75 20 2 \(0^\circ\)
75 25 7 \(2^\circ\)
75 30 28 \(7^\circ\)
75 31 40 \(11^\circ\)
125 10 0.2 \(0^\circ\)
125 20 2 \(1^\circ\)
125 25 8 \(8^\circ\)
125 30 71 \(33^\circ\)
125 31 126 \(90^\circ\)
250 10 0.2 \(0^\circ\)
250 20 3 \(2^\circ\)
250 25 30 \(27^\circ\)
250 25.8 67 \(62^\circ\)
250 26 75 \(90^\circ\)
250 27 28 \(157^\circ\)
250 31 8 \(172^\circ\)
500 10 0.2 \(0^\circ\)
500 15 0.8 \(1^\circ\)
500 20 8 \(9^\circ\)
500 20.5 16 \(19^\circ\)
500 21 40 \(90^\circ\)
500 21.5 17 \(158^\circ\)
500 22 9 \(168^\circ\)
500 25 3 \(176^\circ\)
500 31
1000 10 0.3 \(0^\circ\)
1000 15 2.6 \(5^\circ\)
1000 15.5 5.4 \(10^\circ\)
1000 16.0 30.4 \(90^\circ\)
1000 16.5 5.8 \(169^\circ\)
1000 17.0 3.1 \(174^\circ\)
1000 20.0 0.9 \(178^\circ\)
1000 31 0.4 \(178^\circ\)
2000 5 0.2 \(0^\circ\)
2000 10 1.0 \(2^\circ\)
2000 10.5 2.0 \(5^\circ\)
2000 11.0 22 \(90^\circ\)
2000 11.5 2.3 \(174^\circ\)
2000 12.0 1.1 \(177^\circ\)
2000 15.0 0.3 \(179^\circ\)
2000 31 0.1 \(179^\circ\)
4000 5 0.5 \(1^\circ\)
4000 5.5 1.0 \(2^\circ\)
4000 5.8 2.3 \(6^\circ\)
4000 6.0 20 \(90^\circ\)
4000 6.2 2.4 \(173^\circ\)
4000 6.5 1.0 \(177^\circ\)
4000 7.0 0.7 \(178^\circ\)
4000 20 0.04 \(180^\circ\)

The quantities \(A\) are given in \(10^{-8}\) cm; this quantity corresponds to the order of magnitude of a molecular diameter. If we calculate the values of velocity or acceleration for some frequency,

then the relative amplitudes will be the same as for \(A\). Likewise, the curvature of the membrane will vary according to a sinusoidal law with a relative amplitude differing little from that given in the table for \(A\). This tells us, in other words, that the excitation patterns on the membrane for a tone devoid of harmonics (objective or subjective) will be the same as the amplitude patterns. It is easy to see that the resonance peaks are much sharper for high frequencies than for low ones. This result agrees with the data on the masking of two tones, as we shall see below.

Fig. 7.

Fig. 7.

It is important to note that, according to the assumptions made, and also according to other acceptable assumptions that may be made, that part of the membrane which is closer to the oval window will lead the more distant parts in phase. Thus, if one schematically draws a cinematogram of the shape of the membrane at every \(1/8\) of a period, then we obtain a series of curves similar to those shown in Fig. 7; these curves are plotted from the numbers of Table 1 for the case \(f = 1000\) cycles/sec. Békésy* assumed that such a motion of the membrane leads to the formation of vortices in the fluid above the region of maximum excitation. With loud exciting tones he observed such vortices both in a model of the labyrinth and through the opening of the labyrinth in a cadaver. He considers the cause of nerve excitation to be the constant pressure exerted on the membrane by this vortex. Békésy believes that such a vortex produces motion of the fluid in the semicircular canals and leads to excitation of the organs

* Békésy, Phys. Zs., 29, p. 783, 1928.

equilibrium, as a result of which, with strong high-pitched sounds, there occurs a reflex inclination of the head toward the ear perceiving the sound. He asserts that such observations were made a sufficient number of times to prove his point of view. However, it should be thought that such vortices can hardly arise if the sound alone does not attain very great strength. Moreover, it seems difficult to explain the masking effect otherwise than by excitation of the nerve endings caused by oscillations of the basilar membrane moving in the manner we have assumed.

TABLE II

$f$ Relative reciprocal value of the pressure at the threshold of audibility (sensitivity) Relative displacement Relative velocity Relative acceleration
75 0.004 1.22 0.03 0.007
125 0.024 4.14 0.52 0.065
250 0.125 2.47 0.62 0.17
500 0.500 1.52 0.76 0.38
1 000 1.000 1.00 1.00 1.00
2 000 1.250 0.73 1.47 2.84
4 000 1.250 0.62 2.48 9.92
8 000 0.180 0.64 5.12 40.06
0.027 0.22 1.72
12 000 0.050 0.0036 0.043 0.52

In Table II are given the relative values of the maximum displacement, velocity, and acceleration produced by tones of various frequencies $f$, giving one and the same amplitude of pressure $P$ inside the cochlea; in addition, the reciprocal values of the amplitudes of the air pressures necessary to produce the sensation of a barely audible tone are given.

— For 8 000 a second series of values was calculated on the assumption that the nerve endings do not come closer than $x = 1.5$ instead of $x = 1$ mm. From the table it is evident that, for frequencies above 1 000 oscillations/sec, the data on sensitivity calculated on the basis of the threshold of audibility agree better with the variation of the maximum velocity, and for tones below 1000 oscillations/sec—with the variation of the maximum acceleration. These calculations were ma-

A New Theory of Hearing

given in the assumption that \(P\)—the pressures at the oval window—are constant for tones of different pitch. As was mentioned above, measurements of the impedance of the tympanic membrane* show that, for constant tone intensity in air at frequencies below 1000 cycles/sec., owing to the fact that the impedance of the tympanic membrane is determined in this region by its elasticity, what is obtained is a constant amplitude rather than a constant velocity of the tympanic membrane. Consequently the quantities \(P\), determined by the velocity of the tympanic membrane in this range of frequencies, will be proportional to the frequency, if one assumes that no further changes arise in the process of transmission through the inner ear. Introducing this new factor into the calculation, we find that the relative values of the velocity in the region below 1000 cycles/sec. will correspond to the values given in the column of accelerations. With this correction the velocity values over the whole range agree as well as could be expected for such calculations with the reciprocals of the pressure values at the threshold of audibility (sensitivity). The coefficient \(k_s\) becomes smaller for low frequencies, which leads to a decrease in the calculated velocities at low frequencies and gives an even better agreement with experiment. From the computational results presented one may conclude with some confidence that equal excitation of the nerves in different parts of the membrane is obtained at equal amplitudes of velocity, even though the frequencies \(n\) are different. It can hardly be expected that equations (10) and (11) give values in good agreement with the actual spatial patterns of excitation, excluding regions near the resonance maximum.

Determination of the Spatial Patterns of Excitation from Masking Data

The spatial patterns of excitation along the entire length of the basilar membrane can be determined from observations of tone masking. In Figs. 8 and 9 are presented the results

* Tröger, loc. cit., p. 34.

determination of masking for my left ear (by the method of Wegel and Lane* for exciting tones of 75, 125, 250, 500, and 1000 cycles/sec. The sound-intensity level for all these tones was the same—20 db, i.e. 0.01 μ in \(m/cm^2\). Corresponding to this, the sensation levels proved to be equal to 27, 40, 54, 65, and 72 db. The vertical solid lines represent the sensation level of the fundamental tone and of the subjective

Fig. 8 and Fig. 9: masking curves for tones 75, 125, 250, 500, and 1000 cycles/sec.

Fig. 8.          Fig. 9.

harmonics, determined by the “best beats” method. The ordinates of the curves express the degree of masking, i.e. the number of decibels by which the threshold of audibility is raised in the presence of the tone being investigated. It is interesting to note that under the given conditions each of the 5 above-mentioned tones excites the same number of audible subjective harmonics, and moreover they lie at approximately the same sound-intensity levels. For example, the sound-intensity levels of the second harmonic are 42, 41, 49, 34, and 42 db below the zero level. The third harmonic lies respectively at 56, 52, 67, 60, and 53 db, the fourth at 67, 75, 82, 85, and 86 db below the zero level.

For the tone of 75 cycles/sec. and for higher tones, the subject-

* Wegel and Lane, Phys. Rev., 19, p. 492, 1922.

subjective harmonics can be heard down to the threshold of audibility of the tone. Thus, for example, if tones of 50 and 103 cycles/sec act on the ear simultaneously, then, for the sound intensity of the tone at 50 cycles at which it will still lie only at the threshold of audibility, one can already find a suitable sound intensity for the tone at 103 cycles at which 3 beats per second will be heard. This circumstance makes it difficult to apply masking data to determine the form of vibration of the principal membrane for pure low tones, since already at the threshold of audibility strong subjective harmonics arise.

Fig. 10. Spatial pattern for pure tones. Vertical axis: excitation level—db. Horizontal axis: distance from the oval window.

Fig. 10.

If one uses equation (10), which gives the relation between \(f\) and \(x\), then the curves of excitation level can be found from consideration of the masking curves. The excitation level for the fundamental tone and harmonics is given on the curves of Figs. 8 and 9. For intermediate points, where one has to rely only on masking data, the determination is less reliable. If it is assumed that the masked tone is barely audible when it lies at the same level below the masking tone as in the case of equality of tone pitches, then one can apply the data for determining the minimal perceptible difference in sound intensity. These data show that if the masking tone lies at a level greater than 40 \(db\), then the masked tone always lies approximately at a level 20 \(db\) below the masking tone, i.e. in intensity it amounts to approximately 0.01 or 1% of the intensity of the masking tone. As the masking tone is lowered below the level of 40 \(db\), the difference between the levels of the masking and masked tones decreases from 20 \(db\) to zero. In this way the curves in Fig. 10 have been constructed.

The right-hand part of the curve for each tone is shown by a dotted line, so that the course of the curves may be seen more clearly. The curves at the top of the figure, drawn with thin lines, were calculated from the data of Table I, the amplitudes being expressed in decibels, and all the maximum values placed on a common level. It is clearly seen that over a length of 3–4 mm of the main membrane these curves have the same form as the masking curves. Toward the higher frequencies the masking curves fall off more steeply than the resonance curves. We thus see that, when the indicated tones act upon the ear, there arises on the main membrane a spatial distribution, or excitation pattern, corresponding to the sinusoidal temporal pattern of pressure in the air.

Fig. 11.

Fig. 11.

In the same way, for any sound acting upon the ear, a certain spatial excitation pattern is obtained on the main membrane of the cochlea, corresponding to the temporal pattern of sound pressure in the air. This spatial excitation pattern on the main membrane is transmitted to the brain along a nervous “cable” containing about 3 thousand separate nerve fibers. The spatial excitation pattern depends both on the character of the sound (pitch and timbre) and on the strength of its action on the ear. This latter circumstance is clearly evident from consideration of Fig. 11, which shows excitation patterns for a tone of 1000 cycles/sec at different sound intensities; the intensity of the fundamental tone is characterized by the maximum of each curve, which gives the excitation level of the given tone.

in relation to the threshold of audibility. These curves were obtained from masking observations and by the “best beats” method described above.

Mechanism of Nervous Excitation

Before considering how the pattern of excitation on the basilar membrane is related to loudness, let us examine in greater detail the mechanism of nerve excitation. The “auditory nerve”* in its structure is very reminiscent of a telephone cable. It contains about 3,000 nerve fibers, each of which consists of an “axial cylinder” surrounded by a fatty substance—myelin. The axis of the cylinder has a diameter of about 0.001 cm and amounts to only about 9% of the diameter of the fiber. Thus the nerve fibers are constructed in an extremely similar way to the insulated wires of a telephone cable. This analogy in structure led some physiologists to the conclusion that all nerve impulses are of electrical origin and that their transmission is similar to electrical transmission along a telephone cable. However, this theory proved untenable.

Most of the early experiments on nerve conduction were made with motor nerves, and the view of the mechanism of nerve conduction presented here is based on numerous investigations of that kind. However, recent work by Adrian** has shown that the action of sensory nerves is essentially the same. A nerve impulse can be evoked by a thermal, chemical, electrical, or mechanical stimulus; it can also be evoked reflexly. Touching the nerve with a red-hot wire, applying acid to it, or making a prick, in all cases causes the appearance of a nerve impulse. In physiological laboratories, induction is usually used for excitation—

* The following paragraph has been included, for greater clarity, from H. Fletcher’s book Speech and Hearing (Speech and Hearing, pp. 125–128). —Trans.

* Adrian, The Basis of Sensation*, 1928.

...by means of a coil with an interrupter. For excitation of the impulse, the force of the stimulus and the speed of its change must exceed a certain minimum. It turned out that the nerve impulse propagating along a nerve is not at all like an electric current traveling along a wire. An individual nerve fiber either does not conduct the nerve impulse at all, or else sends the excitation with full force. In other words, in a nerve fiber the strength of the impulse will be either the greatest possible, or there will be no excitation at all, and the character of the excitation plays no role. The minimum magnitude of the irritating stimulus is different for the various fibers that make up the nerve.

Physiologists often liken the process of nervous conduction to the burning of a powder train (in a shell). The speed of propagation of the burning process along such a train and the amount of heat released during burning are entirely independent of the manner in which the train is ignited at the end. From this point of view it obviously follows that the loudness of a sound is directly connected with the number of excited nerve fibers and the number of impulses sent per unit time, since it is clear that each fiber always carries only its own maximum impulse. Moreover, it should be considered that the minimum exciting force of sound for different fibers must differ by a considerable amount.

The correctness of this point of view is excellently confirmed by the experiments of Porter and Hart* . Nerves that cause muscle contraction were excited by an electric impulse. The strength of the exciting current was gradually increased. The successive contractions of the muscle increased not gradually, but in definite steps, as is seen in Fig. 12. The height of the lines in the figure shows the magnitudes of the successive muscular contractions. If the auditory nerves act in the same way, then as the strength of the tone gradually increases from the threshold of audib—

* E. L. Porter and V. W. Hart, Amer. Jour. of Physiol., Oct. 1923.

A NEW THEORY OF HEARING

...to high loudnesses, the excitation reaching the brain must increase in definite steps, and the threshold of audibility will be determined by the excitation of the first nerve fiber. When all the nerve fibers are excited and the number of impulses sent by each fiber reaches the greatest possible value, any further increase in loudness is impossible.

Fig. 12.

After the passage of an impulse along a nerve there occurs the so-called “refractory” period, during which the nerve is incapable either of being excited or of conducting excitation. This is followed by a “relative refractory” period, during which excitability, conductivity, and the velocity of propagation of the impulse gradually return from zero to their normal value. Before returning to the normal state, the nerve becomes “supernormal,” i.e., more sensitive, a better conductor, and with a higher velocity of propagation. This supernormal state gradually ceases, and the nerve returns to the normal state. Only during the relative refractory period can the nerve send impulses diminished in strength. The duration of the refractory period has been measured by many observers, and although there are rather

broad discrepancies; the best measurements at the present time estimate it at approximately 0.001 sec., while the duration of the relative refractory period is approximately 0.003 sec. According to these data it is clear that the maximum number of excitation impulses that can pass along an individual nerve fiber cannot exceed 1,000 per second. Periodic excitations with a frequency greater than 300 per second will no longer be conducted as normal impulses, since each subsequent impulse will fall within the relative refractory period of the preceding one. According to Adrian’s work, nerve endings have a longer relative refractory period than the nerve fiber itself. He found that, along sensory nerves, excitation can pass with a frequency greater than 150 per second. Consequently, if a pure tone having a frequency of 2 or 3 thousand cycles excites the ear, then it is likely that the number of nerve impulses sent to the brain during one second will be considerably less than the frequency of the exciting sound. Adrian’s work also showed that, when the strength of the exciting stimulus is increased, the number of nerve “discharges” increases.

Some of the nerve endings are very sensitive, whereas others are very insensitive; still others have a degree of sensitivity lying midway between these limits. We shall call the excitation level of a given point of the basilar membrane the ratio of the velocity of motion at that point to the velocity of motion required for excitation of the most sensitive nerve fibers in the group of nerves situated over a length of 1 mm along the membrane; we shall express the excitation level in logarithmic units, i.e. in bels or decibels. We shall suppose that the distribution of threshold values for the nerve endings in the ear obeys the same law as for the eye. Let \(Z\) be the fraction of fibers having an excitation threshold lying below some level \(\beta\); then

\[ \beta-\beta_0=\log \frac{Z}{1-Z}, \tag{12} \]

The constant \(\beta_0\) is the value (with the minus sign) of the right-hand side of the equation at the threshold of excitation; \(\beta_0\) depends on the value of \(Z\) at the threshold of excitation. As Hecht* showed, this law is valid for the nerve endings of the eye.

As the level of excitation rises above the threshold of a given nerve ending, the sending of nerve impulses becomes more frequent. Let us consider this process. We shall assume that in the process of excitation of a nerve there are two opposite phases. The first phase, which we shall call active, depends on the strength of the membrane’s motion; during this phase there is created some substance or energy which acts catalytically and causes a nerve impulse. The second phase, which we shall call reactive, consists in the reverse process, which destroys or annihilates the result of the active process. Let us call that which is created as the result of the active process the “exciter”**; this may be the energy of motion of the given element of the nerve ending or a substance arising as the result of a chemical process, or else some electrical accumulation. The active process creates the “exciter” at a rate depending on the strength of the excitation; the reactive process destroys the “exciter” at a rate proportional to the amount of it present. If the exciter is the energy of oscillation, then the frictional forces represent the reactive process. Let us denote the quantity of exciter by \(q\).

As a first approximation we shall take the rate of occurrence of \(q\) to be proportional to the energy of motion***, and the rate of its destruction to be proportional to \(q\). Formulating this mathematically, we obtain the differential equation:

\[ \frac{dq}{dt}=a v^2-bq, \tag{13} \]

* Hecht, Proc. of Michelson Mtg. of Opt. Soc. of Amer., Nov., 1928.

** The author uses the term “discharger”—discharger.

*** This assumption is arbitrary; it leads to the conclusion that pulsations of double frequency arise in the brain. Other assumptions may be made in order to obtain, as a result, pulsations of the fundamental frequency and its harmonics.

where \(a\) and \(b\) are constants independent of \(x\) and \(t\), and \(c\) is the velocity of the membrane at the point where the given nerve ending is located.

For the case of a periodic force produced by a pure tone:

\[ c = V \cos \omega t . \tag{14} \]

Substituting this expression into equation (13) and integrating it under the conditions

\[ q = 0 \quad \text{when } t = 0, \]

we find:

\[ q=\frac{aV^{2}}{2b}\left(1-e^{-bt} -\frac{1}{\sqrt{1+\left(\frac{2\omega}{b}\right)^{2}}} \left[\cos(2\omega t-\theta)-e^{-bt}\cos\theta\right]\right). \tag{15} \]

where

\[ \tan \theta=\frac{2\omega}{b}. \]

If \(b\) is small in comparison with \(\omega\), then

\[ q=\frac{aV^{2}}{2b} \left(1-e^{-bt}-\frac{b}{2\omega}\sin 2\omega t\right). \tag{16} \]

Thus we see that the quantity of the excitant at a given moment is proportional to the energy of the vibration, inversely proportional to \(b\), and directly proportional to the quantity in parentheses, which is a function of time. This process continues until the value of the excitation threshold for the given nerve ending is reached, and then a new process of excitation of the nerve begins. While this second process is taking place, equilibrium is disturbed, and the accumulation of the “excitant” ceases. This period of arrest of the first process is the reaction time (refractory period) of the nerve ending. After the reaction period the active process begins again. At high sound intensities the number of impulses per second will depend chiefly on the duration of the reaction period, whereas at low sound intensity it is determined chiefly by the time of accumulation of the “excitant.”

For illustration, equation (16) is presented graphically in Fig. 13 for frequencies of 125 and 500 oscillations/sec., at—

than was adopted, \(b=125\). This nerve fiber can receive the quantity \(q\) necessary for excitation either in a short time \(t\) at a high velocity \(V\), or in a long time \(t\) at a low velocity \(V\). Consequently, for a given value of \(V\), sensitive fibers require only a short time \(t\) for excitation, whereas insensitive fibers require a longer time. From the drawing we see that, for a tone of 125 cycles/sec., \(q\) increases almost tenfold over the time from 0.01 to 0.003 sec., whereas over the time from 0.003 to 0.005 sec. it changes almost not at all. Obviously, the number of impulses issuing from an excited bundle of nerve fibers during the time of the first period will be much greater than during the time of the second.

Fig. 13. Plot of the quantity \((1-e^{-bt}-\frac{b}{2\omega}\sin 2\omega t)\) for \(b=125\); vertical axis: number of excitations; horizontal axis: time in seconds \(\times 10^{-3}\); curves marked \(\omega/2\pi=500\) and \(\omega/2\pi=125\).

Fig. 13.

Consequently, the “bombardment” of the brain by nerve impulses will have maxima of intensity separated by intervals equal to the half-period of oscillation of the exciting sound. In other words, the temporal patterns of the oscillations occurring in the medium that transmits the sound waves will also be transmitted to the brain. It therefore seems correct to us to suppose that there exists a certain kind of relation between the frequency of the nerve impulses reaching the brain and the temporal pattern of the wave exciting the ear. As the frequency of the tone increases, the temporal pattern in the nerve wave becomes less distinct, since the ratio \(\frac{b}{2\omega}\) becomes smaller. As can be seen from the figure, for 500 cycles/sec. the pulsations become much smaller. The deviations from the exponential curve for a tone of 5000 cycles/sec. would be another 10 times smaller than for a tone of 5000 cycles/sec.

For large values of the velocity \(V\), the nerves will “dischar-”

“... repeated” after a very short interval of time; for small values of \(V\), the “discharge” will occur after a longer time, or will not occur at all. Let \(V_1\) be the value of the velocity for an excitation level \(\beta\) such that \(q\) corresponds to the threshold of stimulation for some nerve fiber excited for a sufficiently long time, so that \(e^{-bt}\) becomes small compared with unity and the expression in parentheses in equation (16) attains its maximum value \(\left(1+\dfrac{b}{2\omega}\right)\). Let, further, \(V_2\) be another velocity, greater than \(V_1\), corresponding to the level \(\alpha\), which gives the same quantity \(q\) in time \(t\). Then:

\[ V_1^2\left(1+\frac{b}{2\omega}\right) = V_2^2\left(1-e^{-bt}-\frac{b}{2\omega}\sin 2\omega t\right); \tag{17} \]

\[ \alpha-\beta = \log \frac{1+\dfrac{b}{2\omega}} {1-e^{-bt}-\dfrac{b}{2\omega}\sin \omega t}. \tag{18} \]

The quantity \(t\) in the last expression is the time required to produce a “discharge” of the nerve ending at an excitation level lying \(\alpha-\beta\) bels above the threshold of stimulation. As was indicated above, the reaction time \(\tau\) is the time required for the process of “discharge” and recovery of the nerve, after which it is again ready for excitation. Consequently, \(t+\tau\) will be the interval of time after which the process in the nerve will again be repeated. At the end of each cycle the phase of the exciting oscillation will, generally speaking, differ from the phase at the beginning. However, on the average we may neglect the influence of the term \(\dfrac{b}{2\omega}\sin \omega t\) in determining the number of “discharges” per second at some level of excitation, for the reason that this term produces only small oscillations about the value of the period obtained when this term is neglected. The number of discharges per second determined in this way will be suitable for all frequencies and depends only on the excitation level. With this understanding, the average number of discharges sent per second

nerve ending, excited at the level \((\alpha-\beta)\) bels above its threshold, will be determined by the expression:

\[ r=\frac{1}{t+\tau}. \tag{19} \]

The fraction of threshold values of excitation lying between the levels \(\beta\) and \(\beta+d\beta\) will be equal to \(\frac{dZ}{d\beta}\,d\beta\). Consequently, if \(n\) is the number of nerve fibers per \(1\) mm, then the number of nerve discharges per second, issuing from an element of length \(dx\) of the basilar membrane, will be expressed as follows:

\[ n dx \int_{0}^{\alpha} r\,\frac{dZ}{d\beta}\,d\beta . \tag{20} \]

For convenience we shall denote the number of discharges of a single nerve fiber by

\[ S_{(\alpha)}=\int_{0}^{\alpha} r\,\frac{dZ}{d\beta}\,d\beta . \tag{21} \]

The quantity \(\frac{dZ}{d\beta}\) may be substituted from equation (12), and the quantity \(r\)—from (18) and (19). The total number of nerve discharges per second \(R\), sent to the brain from all parts of the membrane excited by the given tone, will be:

\[ R=\int_{0}^{31} n S_{(\alpha)}\,dx . \tag{22} \]

The value of this integral cannot be computed analytically, but it can be found by a graphical method. In order to find the value of \(S_{(\alpha)}\), we change the variable \(\beta\) to \(Z\); then

\[ S_{(\alpha)}=\int_{Z_0}^{Z} r\,dZ, \tag{23} \]

where \(Z_0\) is the fraction of the total number of fibers that comes into excitation at the threshold of stimulation. Taking \(\tau=0.002\) sec.,

\(b=125\) and \(Z_0=10^{-4}\), we can calculate the quantities \(S_{(\alpha)}\) at various levels of excitation \(a\); the results of the calculation are shown in Fig. 14; each successive curve is plotted on a scale 10 times smaller. If \(S_{(\alpha)}\) is multiplied by the number of nerve endings at the excitation level \(a\), then we obtain the total number of discharges per second from these nerve endings. Per 1 mm of length of the basilar membrane there are about 1000 rods of Corti. Consequently, according to the curve at the threshold, i.e. at the zero excitation level, we obtain \(1000\cdot 10^{-2}=10\) discharges per second from 1 mm of length and \(1000\cdot 5\cdot 10^2=500\,000\) discharges per second at excitation with the greatest possible strength at the level of 110 db.

Comparison of the Theory of Spatio-Temporal Patterns with Experimental Studies of Hearing

We see that the change in the sensitivity of hearing as a function of frequency can be satisfactorily explained by changes in the amplitude of the velocity of the basilar membrane, with the entire organ of Corti being regarded as a purely mechanical system.

Fig. 14. Graph with vertical axis labeled “\(S_{(\alpha)}\), number of nerve discharges per second” and horizontal axis labeled “excitation level in db”; several rising curves are marked by powers of 10.

Fig. 14.

There is no doubt that loudness is closely connected with the total number of nerve discharges reaching the brain. From Fig. 14 we see that the possible range of loudness perceived by 1 mm of the basilar membrane oscillating at various amplitudes amounts to 110 or 120 db. If we take into account the action of the entire length (31 mm) of the basilar membrane, then the total possible number of impulses will be another 30 times greater, which is equivalent to an increase of the perceived loudness in the region of strong sounds by another 20–30 db. Thus the experimentally observed range of loudness of 140 db can

A NEW THEORY OF HEARING

can be fully explained by the proposed mechanism of nervous excitation. With a gradual increase in loudness from the threshold of audibility, at first only a very small portion of the membrane is excited; this portion gradually increases, then new portions begin to be excited, corresponding to the subjective harmonics, and finally, at high loudnesses, the nerve endings along the entire length of the membrane are brought into excitation. If the basilar membrane vibrated in the same way for all tones, like, for example, the membrane of Wente’s condenser microphone, then the loudness for all tones at a given level of sensation (above the threshold) would be the same. Observations show, however, that the loudness of low tones increases considerably more rapidly than that of high tones.* This difference in loudness for tones of different pitch at one and the same level of sensation is easily explained by the difference in the form (pattern) of vibration of the membrane.

The function calculated above is based on the assumption that, as the strength of excitation increases, the frequency of the nerve discharges increases and ultimately reaches a limit of 500 discharges per second. Experiments on nerves have shown that under continuous excitation a nerve becomes fatigued, after which the maximum number of impulses passing through it is considerably reduced. Consequently, if a tone of a given pitch excites the ear for some time, then immediately afterward the loudness of tones in the same frequency region should decrease. This effect has recently been investigated by Békésy,** who showed that after strong fatigue the loudness can decrease by an amount up to 30 \(db\) (by a factor of 1000).

In the same way, after strong fatigue the threshold of audibility increases considerably. Measurements in our laboratory, which will soon be published, show that for low tones a threshold shift of up to 25 \(db\) is obtained, with some shift still remaining even 2 minutes after cessation of the exciting tone. The relation, derived—

* H. Fletcher, Speech and Hearing, pp. 225—231, New-York, 1929.
* Phys. Zs.* 30, p. 115, 1929.

...stated above is applicable, evidently, only to unstimulated nerves. Observations show that the changes attain their greatest magnitude in the region lying near the fatiguing tone, and fall to zero for tones far removed from it. This provides additional proof of the fact that a simple tone excites only certain regions along the length of the basilar membrane.

We have seen that, owing to the occurrence of a large number of subjective harmonics and because the resonance peaks become less sharp, the region of strong excitation for low tones is much broader than for high ones, if both are at a common sensation level above the threshold of audibility. For this reason low tones send to the brain a much larger number of impulses and are therefore perceived as louder.

Fig. 15. Graph with vertical axis labeled “\(K\), total number of discharges per second” and horizontal axis labeled “excitation level ‘\(\alpha\)’ for 1000 cycles/sec”; curves are marked \(10^2\), \(10^3\), \(10^4\), \(10^5\), \(10^6\), \(10^7\).

Fig. 15.

The values of \(R\) according to equation (22) have been calculated for the 7 tones indicated in Fig. 10, with \(n = 1000\) per 1 mm being assumed. The results of the calculation are given in the third column of Table III. In a similar way, the values of \(R\) for the tone 1000 cycles/sec have been calculated for various levels on the basis of the curves of Fig. 11; the results are given in Fig. 15.

By the loudness of a given sound we mean* the sensation level of the tone 1000 cycles/sec, set to equal loudness with the sound under investigation. Consequently, the loudness values of each of the tones given in the column of Table III can be obtained by comparing the quantities \(R\) with the quantities \(K\) for the tone 1000 cycles/sec and noting the corresponding level

* Speech and Hearing, p. 226.

TABLE III

Frequency Sensation level $R$ Calculated loudness Observed loudness $I'$ Observed loudness $Ror$ Observed loudness $Re$ Observed loudness $G$
75 27 db 1,000 33 36 50 40 33
125 40 10,000 50 50 57
250 54 40,000 61 60 62
500 65 100,000 69 72 66
2000 75 250,000 77 79 76
4000 73 140,000 71 70

sensations according to Fig. 15. Thus the figures in the fourth column of Table III were obtained. Since all the data for constructing the curves were obtained from an investigation of my left ear, the loudness measurements were made for this same ear. The level of intensity of the 1000 cycles/sec tone was adjusted to the same loudness as each of the 6 tones sounding at the same intensity level of 20 db. The observer listened to 2 tones in alternation, so that the influence of fatigue was excluded. The results of the observations are given in the fifth column of Table III. Since these data differ somewhat from the results of Kingsbury’s measurements, I thought that my measurements had been influenced by knowledge of the expected result, and therefore asked another observer, unfamiliar with the question, to make similar measurements; his results are given in the sixth column. In view of the fact that for the tone of 75 cycles/sec too large a difference was obtained, observations for this tone were also made by two more persons (the seventh and eighth columns). The agreement between calculation and observation, within the limits of observational error, may be considered good, although for the two lowest tones the observed values turn out to be somewhat greater than the calculated ones. It is possible that the phenomenon called the “volume” of a tone influences the observer in the direction of a higher estimate of loudness at low tones. “Volume” may be one of the elements entering into the judgment of loudness.

Let us now consider how observations of the sensitivity of hearing to changes in the pitch and intensity of tones agree with our theory. These two effects are closely connected, since in both cases what is involved is a noticeable small change in the form of vibration of the basilar membrane, at least for high tones. First let us consider sensitivity to a change in the pitch of a tone.

Since the resonance peaks are sharper for high tones, it is to be expected that a smaller difference in pitch should be perceived here, which is in agreement with the experimental data. It may also be expected that, at greater intensity of the tone, the sensitivity to pitch discrimination becomes greater, since under these conditions the numerous peaks for the harmonics assist the comparison. This too is in agreement with the experimental data. The smallest change in the pitch of a tone that can be detected is about \(0.25\) centioctave

\[ \left(\frac{\Delta f}{f}=1.0017\right), \]

which corresponds to a displacement of the resonance peak by approximately \(0.01\) mm; along this extent there are about 10 nerve endings. Thus the observed effect is not in contradiction with the anatomical data.

The smallest change in the level of intensity noticeable to the ear is about \(0.25\) db. This corresponds to a change in \(R\) of approximately \(5\%\). There is every reason to think that this quantity must vary both with changes in the level of intensity and with the pitch of the tone, since in essence a change in \(R\) is connected with any change in the spatial pattern on the basilar membrane. If two tones of one and the same frequency are compared, then probably the difference in the intensity of the sound allows us to detect precisely a change in the spatial pattern. If a change in the form of the spatial pattern is the chief cause for the perception of small changes in the intensity of sound, then it is clear that the sensitivity to a change in the intensity of sound will depend on the same factors as the sensitivity to a change in the pitch of a tone. This point of view agrees with the experimental data concerning sensitivity to a change in the intensity of sound. It also agrees with the known from

experiments by the fact that equalizing the loudnesses of tones of different frequencies is very difficult. Indeed, in this case it is necessary to compare the relative quantities \(E_1^*\).

Binaural Effect

In Fig. 16 a diagram is given of the passage of the auditory nerves to the brain. The circles represent those places where certain nerve fibers end and, as it were, switch over into other fibers. Such places occur at several points along the length of the nerve pathways. Some fibers pass without interruption from the cochlea to the auditory center in the brain, whereas others are interrupted 3 or 4 times in intermediate centers. The nerves coming from each ear cross at two points, passing from the right side to the left.

Fig. 16. Diagram of the passage of the auditory nerves to the brain. Labels visible in the figure include: left cerebral hemisphere, right cerebral hemisphere, midbrain, lower part of the pons, left ear, right ear.

Fig. 16.

Most of the nerve fibers going—

* Even before Fletcher’s work appeared in print, an extremely important experimental paper by Weaver and Bray appeared (E. G. Weaver and C. W. Brei, Journ. of Gener. Psychol., XIII, p. 373, 1930), in which one can see proof of the correctness of Fletcher’s theory. Weaver and Bray performed the following experiment. The auditory nerve of a cat was severed between the cochlea and the brain, and an electrode was applied to the outer part of the nerve, which was connected to a lead from an amplifier; another lead from the amplifier was applied to the animal’s skin. It turned out that under these conditions a telephone connected to the amplifier reproduced all sounds falling on the cat’s ear; even speech was transmitted distinctly, and thus the ear appeared to be a kind of microphone converting sound vibrations into electrical ones. After the death of the animal, the transmission ceased. This experiment clearly shows that the auditory nerves carry impulses that preserve the temporal periodicity (pattern) of the sound that falls on the ear.

To be continued.

from the left ear pass, partly by different paths, to the right side of the brain; a smaller part passes to the left side of the brain; the reverse relations hold for the right ear. Thus, when sound is perceived by the left ear, in both the right and the left cerebral hemisphere there arise two excited regions similar to one another. These regions are denoted by the letter $L$. In the same way, perception of sound by the right ear causes the appearance of two excited regions, denoted by the letter $R$. If the nerve paths are cut at point $A$, then the left ear becomes absolutely deaf. If, however, the nerves are cut at point $B$ or $C$, then only some weakening of hearing on both sides will occur. Cases of brain tumors are known in which the patient had to have, by surgical means, removed a part of the brain containing the auditory center and adjacent regions in one half of the brain. After recovery, half of the body remained paralyzed, but tests of hearing showed that both ears had almost the same sensitivity as before the operation.

Although some of the nerve fibers from both ears, over part of their path, run close to one another and end close to one another at the periphery of the brain, only a very weak interaction is observed between them. Experiment shows that a tone acting on one ear produces only a very small masking effect on the perception of sound by the other ear. Wegel and Lane showed that the masking tone in the opposite ear must be $10^6$ times, or $60\,db$, stronger in order to produce the same masking effect as when acting on the same ear.

In the same way, “objective” binaural beats are not observed until the sound intensity in one ear exceeds the intensity in the other by at least $60\,db$. “Objective” beats differ from “subjective” ones in the following features: they can easily be heard by all observers, and they can be obtained at all audible frequencies, whereas “subjective” beats can be heard by only 80% of observers and only at frequencies below 1000 cycles/sec. “Objective” binaural beats undoubtedly arise—

they act only owing to the penetration of the stronger sound into the cochlea of the opposite ear by means of bone conduction. An intense tone produces the same effect as though it had been weakened and introduced directly into the same ear as the other tone.

The areas of the cerebral cortex excited by the two ears lie in close contact with one another, so that they can be compared in strength and in time of excitation. Above we have shown that the moments of maximal nervous excitation correspond to the moments of maximal displacements in the air wave. Thus the phase difference of a sound received in the two ears manifests itself as a difference in the time of the maxima of excitation of the \(R\) and \(L\) areas of the cerebral cortex in both hemispheres. In judging the direction of the source of sound, both the difference in sound intensities and the difference in phase have an influence. In the course of mental development we learn to associate the direction of a sound with a certain difference in the intensity of excitation of the \(R\) and \(L\) areas and with the difference in the times of their maximal excitation. We have seen that, according to the theory developed, the distinctness of the temporal pattern in the excitation of the nerve centers must diminish as the frequency increases, which agrees with the fact that localization of direction for sounds having a frequency of more than 1000 cycles/sec. becomes very difficult.

Conclusions

In summary, we may say that pitch is determined both by the position of the resonance maximum on the basilar membrane and by the temporal “patterns” reaching the brain. The former is probably more important for high tones, the latter for low tones. Loudness depends on the number of nerve impulses per second reaching the brain, and perhaps also in part on the size of the excited area. What psychologists call “volume” is undoubtedly connected with the length of the excited area of the basilar membrane. This magnitude is transmitted to the brain and forms an excited area of the cerebral cortex of a certain size. The size of this ...

region and determines our sensation of the “volume” of a tone. Low tones and tones of complex composition therefore have a large volume, whereas high tones have a small one.

The concept of “brightness,” introduced by psychologists, may be connected with the sharpness of the resonance peaks on the basilar membrane, as Troland* assumes. High tones give a sensation of brightness, low tones a sensation of dullness or dimness (dullness), which, as Fig. 10 shows, may correspond to a greater or lesser sharpness of the resonance peaks.

The temporal patterns of the air vibration are transformed into a spatial pattern on the basilar membrane. The nerve endings are thereby excited in such a way that these spatial patterns are transmitted to the brain and evoke two regions of excitation similar to one another, one in the right and the other in the left hemisphere of the brain. Between the temporal sequence of the vibrations of the air wave and the sequence of maximum excitations in the cerebral cortex there exists a definite connection. Thus, when listening to a sound with both ears, two excited regions are formed in each hemisphere of the brain, in each of which there is a certain periodicity of excitations with some phase shift relative to one another. The perception of the relation between the temporal pulsations in two neighboring regions of the cortex, excited by the right and left ear, explains all the peculiarities of the so-called binaural effect.

* L. T. Troland. J. of Gen. Psychol., 2, p. 28, 1929.

Submission history

A NEW THEORY OF HEARING