Abstract
In this article, I will attempt to provide an overview of contemporary research on hearing and speech that is of great importance for a range of issues in technical acoustics.
Full Text
HEARING AND SPEECH IN THE LIGHT OF CONTEMPORARY RESEARCH
C. N. Rzhevkin, Moscow.
The science of sound, which in the nineteenth century attracted the attention of a wide circle of physicists and physiologists and was brought into a finished system by the works of Helmholtz, Rayleigh, and a number of other major scholars, in subsequent years, up to the beginning of the twentieth century, remained almost outside the range of broad scientific research, yielding deservedly to more pressing questions.
If, on the one hand, the reason for the diminished interest in acoustics probably lay in the fact that the basic principles in this field had long since been firmly established, and it seemed that all that remained was to work out the details, then, on the other hand, further acoustical investigations encountered extraordinarily great experimental difficulties, and the solution of many important and disputed questions awaited further improvement of experimental technique. A not insignificant role was also played by the circumstance that, until the beginning of the twentieth century, acoustics had very few practical applications.
In the twentieth century, in connection with the rapid development of telephone communications in all countries of the globe, thanks to the development of radio broadcasting and the invention of loudspeakers, a whole series of new practical problems arose before acoustics. Let us point at least to the study and calculation of the telephone as the principal electro-acoustic instrument, and to the detailed investigation of hearing and speech as the basis of all practical applications of technical acoustics. A whole series of new and most interesting problems arose and were solved in connection with the development of underwater sound signaling; finally, an enormous number of acoustical problems were advanced by the demands of military technology.
In the present article I shall attempt to give a survey of contemporary investigations of hearing and speech that are of great importance for a whole series of questions in technical acoustics.
I. Structure of the Auditory Apparatus. Limits of Audibility of Sound.
The principal part of the human auditory apparatus is lodged deep in the right and left temporal bones, in a bony formation called the labyrinth, in which lie the nerve endings sensitive to sound. The labyrinth also serves as an organ of equilibrium (semi-
semicircular canals), while the nerve fibers sensitive to sound lie in the part of the labyrinth called the cochlea. The cochlea is a spirally twisted bony sheath (2¾ turns), inside which there runs a continuous passage from the base to the apex. This passage is about 31–33 mm long (if straightened); its diameter at the base is 9 mm and at the apex 1.8 mm. It is divided along its entire length into two parts (see Fig. 1), the boundary between which is, for the greater part, a bony partition spirally wound along the turns of the cochlea (lamina spiralis), and, for a shorter extent, a flexible membrane—the basilar membrane (membrana basilaris).
Parallel to the basilar membrane and very close to it runs a second membrane—the tectorial, or Corti’s, membrane; it begins, like the basilar membrane, at the bony partition, but extends only over an insignificant part of the width of the cochlear duct, not reaching its opposite edge. The part of the basilar membrane lying next to the tectorial membrane is considerably thickened, and it is precisely at this place that the endings of the auditory nerves branch. The endings of the cells of the basilar membrane adjacent to the tectorial membrane have the finest hairs, which can come into contact with the tectorial membrane during oscillations.
Fig. 1. Transverse section of the cochlear duct.
The basilar membrane consists of a large number of transverse fibers (from 13,000 to 24,000 fibers according to investigations by various scholars), very elastic and weakly connected with one another. The number of endings of the auditory nerve reaches 4,000. The entire cochlear duct is filled with fluid. Its two parts, separated by the basilar membrane, communicate at the apex of the cochlea through a small opening (the helicotrema). Both halves of the cochlear duct communicate with the air cavity of the middle ear through openings closed by membranes: the oval window, to whose membrane the auditory ossicle (the stapes) is attached, transmitting to the fluid of the cochlea, through the chain of other ossicles (the malleus and incus), the movements of the tympanic membrane; and the round window, connected with nothing and opening into the air cavity of the middle ear.
The structure of the ear is shown schematically in Fig. 2a; the canal of the cochlea is shown straightened and relatively wider than in reality—
sensitivity; it is shown in section, along its entire length, perpendicular to the auditory membrane and the spiral partition.
Sound waves, penetrating into the auditory canal, set the eardrum (Б.) into vibration and, through the chain of articulated ossicles (malleus, incus, stapes), transmit the oscillatory motion to the oval window (Ов) and through it to the fluid of the cochlea. Since the fluid of the cochlea is almost incompressible, the movements of the oval window cause, through the medium of the fluid, movements of the round window (Кр.) in the opposite direction (analogous to what occurs in a hydraulic press). The fluid of the cochlea, transmitting the vibrations, in its motion sets into oscillation certain parts of the basilar membrane (М.осн.), which causes
Fig. 2. a—Schematic drawing of the cochlea. b—Distribution of tones along the length of the basilar membrane.
the hairs of the auditory cells to touch the tectorial membrane (М.тект.) and leads to excitation of the auditory nerves.
The structure of the auditory membrane is such that it is highly probable that in its separate parts it is capable of almost independent oscillations with a definite natural frequency. If the oval window is subjected to slow indentation, then the fluid moves in a circle through the entire cochlear canal, through the helicotrema (Ге.); the motion is transmitted to the other half of the canal and then produces a protrusion of the round window; thus the auditory membrane does not take part in the motion. The same occurs with slow oscillations with a frequency of less than 20 per sec.
If the oval window is subjected to rapid oscillations, then the motion can be transmitted to the round window by a shorter path—through those fibers of the auditory membrane that are tuned to the frequency of the acting...
...of the corresponding oscillation; in this case only the part of the fluid nearest the base will be set in motion, while the part of it lying between the resonating fiber and the apex of the cochlea will remain motionless; thus more rapid oscillations set in motion a smaller mass of fluid and can create a greater amplitude of oscillation of the auditory membrane in the narrow region where it resonates to oscillations of the given frequency. In agreement with anatomical investigations, this leads to the conclusion that high tones are perceived by parts of the membrane lying near the base, and low tones—near the apex of the cochlea. Quite analogously, if in the path of an alternating electric current there lies a series of parallel paths consisting of a capacitance connected in series with a self-inductance, i.e. tuned to a known period, then the current will choose its path through that circuit whose period of oscillation either coincides with or in general lies closest to its own period (resonance); through circuits whose periods do not coincide with its period, the current will pass the more weakly, the greater the difference of the periods. The views set forth above were in principle developed earlier by Helmholtz and constitute the basis of the so-called resonance theory of hearing.
Experience shows ¹) that for sounds above 15,000 oscill./sec. the ear becomes sharply less sensitive and loses the ability to distinguish pitch. This argues that the auditory membrane hardly contains fibers tuned higher than 15,000 oscill./sec.
These observations are easily explained by assuming the presence of auditory fibers with tuning no higher than 15,000 oscillations. Higher sounds set these fibers into ever weaker and weaker oscillations as their pitch rises, but, by increasing the intensity of the sound, it is always possible to bring the amplitude of these fibers to the threshold of sound sensation. In view of the fact that, when the frequency is increased above 15,000, the sensation will always be obtained only through excitation of the extreme fibers, a change in the pitch of the sound will not be perceived.
The upper limit of audibility can also be explained otherwise. The point is that the eardrum stands at an angle to the auditory canal. At a tone frequency of 20,000 oscill./sec., i.e. a wavelength of about 6 mm, almost an entire wave is accommodated along the length of the obliquely standing eardrum, and therefore its separate parts will experience opposite actions, so that the motion is not transmitted to the auditory ossicles. This view is held by P. P. Lazarev. Undoubtedly, the upper boundary of hearing is conditioned by the action of both of the indicated causes together. In addition, the inertia of the fluid of the labyrinth must play a role, the influence of which increases in proportion to the square of the frequency.
¹) Lane, Phys. Rev. 19, 492, 1922.
On the basis of experiments with the damping (masking) of one sound by another, Wegel and Lane1 determined at which portions of the auditory membrane the perception of sounds of a known pitch occurs. Fig. 20 shows this distribution of frequencies along the length of the auditory membrane. As we see from the drawing, a tone of 100 cycles/sec acts almost on the very end portions of the membrane near the helicotrema. How, then, is this to be reconciled with the fact that the ear perceives much lower tones—50, 30, and even down to 16 cycles/sec.? Fletcher2 explains this fact by saying that strong sounds evoke the sensation of subjective overtones that are not present in the exciting vibration. The lower the sound, the more the ear apparently possesses this ability to create overtones even at a low sound intensity. A tone of 16 cycles/sec., probably, cannot itself be sensed by the ear for lack of the corresponding fibers of the membrane, but the overtones it creates—32, 48, etc. vibrations per second—produce the sensation of a certain low sound. From these considerations it follows that the lowest tones of the musical scale are perceived by the ear only insofar as they excite subjective overtones. This is why an exact determination of the lower limit of audibility is essentially impossible. The slight sensitivity of the ear in distinguishing the pitch of low tones is also understandable. The phenomenon of the formation of subjective overtones, as well as of combination tones, which will be discussed further on, is closely connected with the asymmetrical structure of the tympanic membrane. The tympanic membrane consists of radial fibers convex outward and has the form of a cone with its apex turned inward; therefore it yields far more readily to a force drawing it outward than to one pressing it inward. By virtue of such a structure, sound vibrations, especially those of great intensity, produce asymmetrical displacements of the tympanic membrane inward and outward, which, as Helmholtz3 showed theoretically, leads to the formation of subjective combination tones and overtones.
The resonance theory of hearing has recently received excellent confirmation in the experiments of Held and Kleinknecht4. These authors produced a minute injury in the labyrinth of a guinea pig, boring with a drill a 0.1 mm opening calculated so as not to damage the auditory nerves, but only to weaken the tension of the auditory membrane by rupturing it. After this, the animal lost the sensation of a definite pitch, which was demonstrated by the conditioned-reflex method. When the injury healed, the sensation of the tone was restored—
was observed. Drilling the cochlea closer to the base caused the loss of high tones; closer to the apex—the loss of low tones. The same result is also given by numerous experiments involving atrophy of the auditory membrane by prolonged sounds.
II. The threshold of sensitivity of the ear as a function of pitch. Acuity of hearing in the assessment of the pitch and strength of sound.
It is quite clear that the study of the sensitivity of the ear to tones of different pitch is the basis of most calculations in technical acoustics, and therefore it should not be surprising that very much attention and labor have been devoted to this question by many eminent scientists.
Fig. 3. Sensitivity of the ear according to the data of M. Wien and according to new data.
Let us agree to call the threshold of auditory perception the minimal energy of sound vibrations (in ergs) passing in one second through an area of \(1 \text{ cm}^2\), which produces the sensation of a barely audible sound. By the sensitivity of the ear we shall then mean the value reciprocal to the threshold of audibility; since the sensitivity of the ear is expressed by large numbers and varies greatly with pitch, it is more convenient to take the common logarithm of the sensitivity. Thus, for example, if the ear has a threshold of
\[ 10^{-6}\frac{\text{erg}}{\text{cm}^2\text{ sec.}}, \]
then the sensitivity will be expressed by the number \(10^6\), and its logarithm by the number 6.
Of earlier works on the determination of sensitivity we shall mention only the classical work of M. Wien \(^{1}\), who was the first to determine the absolute value of the threshold of audibility when the frequency was varied from 50 to several thousand vibrations; Wien’s data are shown in Fig. 3 on a logarithmic scale. From the curve it can be determined that sensitivity reaches a maximum at \(2300\) vibrations/sec. The increased sensitivity of the ear in the region of \(2000\)–\(3000\) vibrations is also pointed out by Helmholtz, who explains this by resonance of the auditory canal. These high sounds (the beginning of the 4th octave) really do sound especially sharp to the ear. It is also curious to note the fact that precisely sounds of about \(2500\) vibrations/sec are used for the underwater signals of the ceilon—
\(^{1}\) M. Wien. Pflüger’s Arch. 97, 1, 1903.
…Tuzems—pearl divers. Wien’s data served as the starting point for determining the most advantageous pitch height in modern installations for underwater signaling.
The maximum of sensitivity found by Wien was also found by later investigators; as for the absolute magnitude of the sensitivity, here Wien’s measurements apparently contain a serious error. New investigations1, carried out with extreme care and by entirely different methods, and agreeing excellently with one another (Fig. 3), give, at high frequencies, values of sensitivity hundreds of times smaller than those obtained by Wien. According to Wien’s data, when the frequency changes from 50 to 2,000 vibrations/sec., the sensitivity changes by a factor of 100 million; according to the newest data, however, it changes by approximately 500,000 or one million times, i.e., 100–200 times less.
Of the methods for determining the threshold of audibility we shall mention only the thermophone method2, in view of its special convincingness. The thermophone consists of a thin metal sheet or wire heated by an alternating current, which causes periodic expansions of the air and produces a sound of small intensity. The strength of the sound depends on the volume of the chamber into which the oscillations enter. If a thermophone placed in a chamber is pressed tightly to the ear or mounted in a capsule inserted into the auditory canal, then, knowing the volume of the chamber, the strength of the current passing through the thermophone, and the dimensions of the heated sheet, one can calculate the strength of the sound in the ear in absolute units even for the weakest sounds at the threshold of audibility. In measurements by Wien’s method using a telephone, the strength of the sound at the threshold of audibility could be determined only by extrapolation from stronger sounds, which led to considerable errors. The enormous difference in sensitivity between high and low tones by a factor of \(10^8\) is very difficult to explain from the standpoint of resonance theory, as Wien himself pointed out3.
He calculates (proceeding from the damping of the auditory resonators, as determined by Helmholtz) that with low tones (50 vibrations), so weak that they are still not audible, the fibers of the auditory membrane corresponding to high tones will be set into vibration, although only to a slight degree; moreover, owing to the great sensitivity of the ear in the region around 2,000 vibrations, the sensation of high sounds will arise before the low sound exceeds the threshold of audibility. But in
in reality this is not observed. Much can be objected to Wien’s reasoning: it is undoubted, for example, that sounds of different pitches that are equal in intensity will be transmitted into the fluid of the labyrinth with unequal intensity. Minton’s experiments1 show that the lower the tone, the more weakly it is transmitted into the inner ear. It is also probable that the fibers of the basilar membrane at high tones have much less damping and therefore are excited only very weakly by low tones. Fischer2 adduces still another objection: in order to produce irritation of the auditory nerves, the presence of relative motion of the basilar and Corti membranes is necessary. At low frequencies both membranes will undergo identical oscillations, and although in the region of high frequencies there will be a considerable amplitude of oscillation of the basilar membrane, there will be no relative motion of the two membranes, and no sensation of sound will result.
To explain the contradictions of the Helmholtz theory of hearing, various authors have made additional hypotheses. Thus, Lazarev3 assumes the existence in the auditory nerves of a special sound-sensitive substance, the decomposition of which proceeds considerably more energetically at high frequencies.
Fletcher4 assumes that in the brain centers the sensation of pitch is obtained in accordance with that place of the auditory membrane where the maximum amplitude of oscillations arises, independently of what the amplitude is at other points, and in this way avoids Wien’s objections. He explains the difference in sensitivity by the fact that the mechanism of the middle ear transmits oscillations into the fluid of the labyrinth the more poorly, the lower the tone. In the perception of the very lowest tones, moreover, an important role is played by the formation of subjective overtones; for the lower limit of hearing the appearance of overtones can apparently occur earlier than the fundamental tone becomes audible (if it can, generally speaking, be perceived as such).
It is interesting to point out that the threshold of sensitivity of the ear in the region of high tones (2500 oscill./sec.) is approximately \(10^{-9}\dfrac{\mathrm{erg}}{\mathrm{cm}^{2}\,\mathrm{sec}}\), which is precisely equal to the greatest sensitivity of the eye in the region of green rays. The ear and the eye are thus organs of equal absolute sensitivity. The amplitude of oscillations of air particles for a sound of 2500 oscill./sec. with an intensity of \(10^{-9}\dfrac{\mathrm{erg}}{\mathrm{cm}^{2}\,\mathrm{sec}}\) is less than \(10^{-9}\,\mathrm{cm}\), or about \(1/100\) of the diameter of an air molecule.
The most recent data on determining the threshold of audibility are collected in the graph of Fig. 4. The lower curve gives the value of the threshold of audibility
in ergs of sound energy per \(1 \text{ cm}^2\) per sec., as a function of pitch; the part of the curve in the region of tones of \(50\) cycles/sec. and below is given by a dotted line, in order to indicate the less reliable character of the measurements in this region, which is close to the lower limit of hearing.
Earlier determinations of the upper limit of audibility can hardly be regarded as indisputable, both because no measurement of the intensity of the sound was made and its pitch was not guaranteed, and because of the contradictory data of various authors.
Lane’s investigations\(^{1}\), in which he used a high-frequency generator with a cathode tube and a telephone as the source of sound, showed that the sensitivity to tones from \(4\,000\) to \(15\,000\) cycles/sec. is almost identical and close to \(10^4\). Above \(15\,000\) cycles/sec. the sensitivity at once drops sharply and at \(20\,000\) cycles/sec. is only \(1\), i.e., over an interval of \(1/3\) octave it decreases by 10 million times (Fig. 4). Determination of the threshold of audibility at still higher tones proves impossible, since before the sensation of sound occurs, a sensation of pain and pressure in the ear is obtained. At lower tones and with great sound intensity there is likewise a sensation of pressure and pain; the upper curve in Fig. 4 gives the threshold of this sensation. It is extremely interesting to note that the amplitude of the alternating sound pressures at the sensation of pain and pressure in the ear turns out to be of the same order as the threshold of the sensation of pressure by the skin, namely about \(1\,000\ \text{dyn}/\text{m}^2 \simeq 1\ \text{g}/\text{cm}^2\). In the deaf the threshold of audibility, of course, rises, but the threshold of the sensation of pressure has the same magnitude.
Fig. 4. Region of auditory sensations
In connection with the results of recent works, doubts arise as to the correctness of the former determinations of the upper limit of hearing, which estimated it at \(40\,000\) cycles/sec.
The upper limit of hearing is higher in children; in elderly people it gradually lowers in connection with the general weakening of hearing. Often
\(^{1}\) Lane. Phys. Rev. 19, 492, 1922.
old people do not hear at all the sound of the chirping of grasshoppers and crickets, or the whistle of steam being discharged from a boiler. In animals the upper limit of hearing is, generally speaking, considerably higher. This was proved by I. P. Pavlov by the method of conditioned reflexes. The upper and lower curves in Fig. 4 enclose a certain area, which includes all the sounds that the human ear perceives. This area is naturally called the region of auditory perception.
The data presented in Figs. 3 and 4 on the sensitivity of the ear are the result of a statistical investigation of a large number of people. The investigations of Minton1 and Kranz2 showed that individual sensitivity may differ extraordinarily from the average. Cases are observed of sharp minima or maxima of sensitivity over an interval of one musical tone, i.e. with a change of frequency of only 12%. This phenomenon was observed qualitatively already by Bezold (1900).
The sensitivity then changes throughout by hundreds of times. It has been found that certain diseases of the ear cause a lowering of hearing sensitivity in a definite range of pitches. In sclerosis of the ear, when the membrane of the oval window, owing to deposits of calcareous salts, becomes hard and inelastic, a weakening of hearing at low tones is usually noted, as well as sharp drops in sensitivity (by hundreds and thousands of times) in the region of higher tones. Diseases of the inner ear for the most part first entail the loss of the very highest tones. Thus an investigation of the acuity of hearing may serve the physician both for making a diagnosis of the disease and for determining the degree of deafness.
An objective investigation of the hearing of persons in certain professions, such as: radio listeners, telegraphists receiving by ear (klopferists), musicians, chauffeurs, physicians, and various listeners in military affairs, has now become an entirely practicable task and should play an important practical role. The American firm Western Electric Co. manufactures apparatus intended for this purpose, called audiometers. The author of the present article has developed an apparatus for the same purpose, which is now installed at the central telephone station in Moscow.
The second important question in the study of hearing is the question of the fineness of hearing, i.e. of the ability to distinguish changes in pitch, and also changes in the strength of sound. A detailed investigation by Knudsen3 showed that the fineness of hearing at high tones from 500 to 2,000 cycles/sec. (from \(C^2\) to \(C^4\)) proves to be almost the same and equal to about 0.3%, or \(1/40\) of a musical tone. Thus, for example, at 1,000 cycles/sec. the ear dis—
recognizes a change in pitch at 3 vibrations. Toward low tones the acuity of hearing gradually decreases and at 50 cycles/sec. reaches \(1^\circ\) or \(1/12\) of a whole tone. Persons with especially acute hearing distinguish still smaller changes in pitch.
The acuity of hearing is probably dependent on the structure of the auditory membrane. A membrane whose fibers are weakly connected and can vibrate independently of one another will give fine musical hearing; conversely, a coarse and unyielding auditory membrane will not allow subtle changes of form to arise and will enter into vibration at once over a large area and excite several nerve endings, which will correspond to an unmusical ear, incapable of distinguishing even sounds separated by large intervals. Practice makes the auditory membrane more flexible, apparently weakening the connection between its fibers, as a result of which hearing improves; here a process takes place that may be likened to the breaking-in of a violin.
Fig. 5. Characteristics of deaf ears
Knudsen’s data made it possible to calculate that, at an average sound intensity, the ear can distinguish about 1,500 gradations in pitch in the interval from 32 to 5,000 cycles/sec. The same investigator determined the acuity of hearing with respect to estimating sound intensity. For sounds in the middle part of the musical scale, the threshold of sensation of a change in sound intensity corresponds to a change in intensity of 10%; with weak sounds the ear senses only much larger changes of intensity, up to 30% and more. From these data one may calculate that at 1,000 cycles/sec. the ear will be able to perceive 270 gradations of sound intensity from the threshold of audibility to the threshold of pressure sensation. If one takes into account the ear’s capacity to perceive changes in both the intensity and the pitch of sound, then it may be calculated that within the entire range of auditory perception the ear perceives about 300,000 different tones. Thus the ear provides an enormous variety in the perception of sounds.
In view of the fact that in the ear there are only about 4,000 nerve fibers, each of which is capable of transmitting only a definite effect to the brain (“all or nothing”), without any gradations,—the explanation
of the whole diversity of auditory perception, in pitch and in strength of sound, presents great difficulties.
If the ear has some defect, then the perception of part of the tones by pitch or by strength drops out; part of the auditory region falls out of perception, and the ear perceives a number of tones smaller than normal. Having determined the threshold of sensitivity of the diseased ear and plotted it in the form of a curve (Fig. 5) on the same sheet with the normal region of auditory perception, we immediately form an idea of what percentage of the total hearing has been preserved in the ear under examination. This percentage of preserved hearing will be equal to the ratio of the area between the sensitivity curve and the pressure-sensation curve to the entire area of normal auditory perception. In the drawings presented we have the characteristics of two ears: in one, 74% of hearing has been preserved; in the other, only 12%.
Persons with good hearing have (in the region of speech) a threshold of audibility of the order of \(0.001\) dynes/cm\(^2\) (the magnitude of the variable sound pressure corresponding to an intensity of \(2.5 \cdot 10^{-8}\ \frac{\text{erg}}{\text{cm}^2\text{sec}}\)); when the threshold of audibility is raised to \(0.1\) dynes/cm\(^2\), we are dealing with slight deafness; severe deafness, still permitting conversation, corresponds to a threshold of \(1\) dyne/cm\(^2\). With a threshold of \(10\) dynes/cm\(^2\), conversation can be understood only with amplifying devices. Pressures of \(1000\) dynes/cm\(^2\) already give a sensation of pain; therefore deafness corresponding to a threshold of more than \(100\) dynes/cm\(^2\) can no longer be corrected by any sound amplifiers.
III. Loudness of Sound. Perception of Consonances.
The strength or intensity of sound is the quantity of sound energy (in ergs) flowing through an area of \(1\ \text{cm}^2\) in \(1\) sec. In a freely propagating wave, a given strength of sound corresponds to a definite magnitude of pressure oscillations (above and below normal). The relation between the strength of sound \(J\) in \(\frac{\text{erg}}{\text{cm}^2\text{sec}}\) and the effective (or mean square) magnitude of the pressure \(P\) is expressed by the following formula:
\[ P\ \text{dynes}/\text{cm}^2 = 6.4 \sqrt{J\ \frac{\text{erg}}{\text{cm}^2\text{sec}}} \]
and is explained by the table on p. 243.
Suppose that we have a series of sounds of equal intensity and different pitches. Will these sounds be equally loud subjectively? It is clear without lengthy reasoning that their loudness for the ear will not be the same. Sounds having a pitch below 16 or above 20,000 vibrations/sec will not be heard at all, however great their energy may be. But even sounds lying within the limits of audibility, at the same intensity, may have entirely different loudness, which
| Sound intensity in $\dfrac{\mathrm{erg}}{\mathrm{cm}^2\mathrm{sec}}$ | Variable sound pressure in $\mathrm{dyn}/\mathrm{cm}^2$ |
|---|---|
| 1 000 000 | 6400 |
| 10 000 | 640 |
| 100 | 64 |
| 1 | 6,4 |
| 0,01 | 0,64 |
| 0,0001 | 0,064 |
| 0,000001 | 0,0064 |
is due to the different sensitivity of the ear depending on pitch. Thus, for example, a sound with an intensity of $0{,}0001 \dfrac{\mathrm{erg}}{\mathrm{cm}^2\mathrm{sec}}$ at $50\ \mathrm{vib./sec.}$ will lie below the threshold of audibility (see Fig. 4), whereas a sound at $200\ \mathrm{vib./sec.}$ with the same intensity will be heard by the ear as fairly loud.
The measure of the subjective loudness of a simple tone may be taken to be the ratio of its intensity to the intensity of the sound at the threshold of audibility of a tone of the same pitch.
Since for sounds of ordinary loudness (speech) these ratios are expressed by very large numbers, it is more convenient to take their common logarithm. Thus, for example, a tone 1 000 000 times more intense than at the threshold of audibility will have a subjective loudness expressed by the number 6. American authors express loudness by ten times the logarithm of the indicated ratio. On this scale the loudness of a barely audible tone at $1 000\ \mathrm{vib./sec.}$ will be expressed by the number 0 ($\lg 1 = 0$), and the loudness of a tone producing a sensation of pressure by the number 130. This scale of loudnesses is in practice very convenient1.
An experimental basis for the concept of loudness set forth above is provided by the investigation of Mackenzie2, who found (with the aid of a special type of phonometer, in which two different sounds are brought to the ear in rapid alternation) at what ratios of intensities sounds of different pitches prove to be equally loud. It follows from this work that if pure tones of different pitches, equally loud, have a certain ratio of intensities, then, when amplified by the same number of times, they will also remain equally loud to the ear.
In Fig. 4 a series of curves is drawn with dotted lines, parallel to the curve of the threshold of audibility, at equal distances from one another. Since in the drawing the sound intensity is plotted on a logarithmic scale, these curves correspond to sounds a certain number of times (100, 10 000,
1,000,000, etc.) stronger than at the threshold, i.e. equally loud subjectively. The loudness numbers (on the right) are set down according to the American scale. Sounds that produce a sensation of pressure have, at 50 cycles/sec., a loudness of about 50 TU; at 1,000 cycles/sec. this is obtained at a loudness of 130 TU, i.e. with a sound 10 billion times stronger than the threshold of audibility. The region of speech lies within the limits from 50 to 80 TU; the region of speech is hatched in the drawing; it occupies, as we see, only a small part of the whole region of auditory perception. In a recently published work1 Kingsbury found that, as the strength of the sound is increased, loudness grows more slowly for low tones than for high ones. Thus if a tone of 60 cycles/sec. at the threshold of audibility is \(10^5\) times stronger than a tone of 2000 cycles/sec., then at the level of conversational speech, under the condition of equal loudness, it will be only 10 times stronger. To attain equal loudness, the tone of 60 cycles/sec. must in this case be intensified \(10^4\) times, and the tone of 2000 cycles/sec. \(10^8\) times, relative to the threshold of audibility. According to these data of Kingsbury we would have had to draw, in Fig. 4, the curves of equal loudness so that they would lie closer together at low tones and diverge at high ones.
The data of Kingsbury and Mackenzie are in obvious contradiction, and it is possible that the simple definition of loudness given above will have to be abandoned and replaced by a more complicated one.
The energy carried even by loud sounds is very small in magnitude. For example, ordinary speech corresponds to the radiation of sound energy at 125 ergs/sec. = 12.5 microwatts; in order to heat a cup of tea with such sound, a million people would have to speak in the room for \(1 \tfrac{1}{2}\) hours!
When two or several tones sound simultaneously, auditory perception proves to be entirely different in character depending on the ratio of the numbers of vibrations of the component tones. I shall not dwell here on the general questions of consonance, dissonance, musical harmony, and the construction of scales. These questions are no longer new and have been sufficiently fully illuminated by science. I shall touch in more detail only on the question of the occurrence of combination tones.
Undoubtedly, for hearing the picture of the consonance of several tones is not simply a superposition of the sensations of the individual tones. Even superficial observation shows that every consonance or chord represents a more complex combination than a simple sum of tones. When two strong tones sound together, one usually clearly hears a certain low accompanying sound, or difference tone (this discovery was made as early as the 18th century by the violinist Tartini), whose number of vibrations is equal to the difference of the numbers of vibrations of the primary tones, as well as still other accompanying sounds, lying lower than the two sounding tones.
Helmholtz showed that on some instruments (for example, on the organ and on the double siren) upper partial tones (summation tones) are also obtained, which are not present in the original sounds; the summation tones are rather weak and do not play the same role as the difference tones. All these additional partial tones are called combination tones. In some cases the formation of difference and summation tones takes place in the very source of the sound; this happens, for example, in the organ or the harmonium. But at present we are interested only in subjective combination tones.
Helmholtz showed theoretically that the occurrence of combination tones is possible in the case when some part of the auditory apparatus possesses asymmetric elasticity, i.e., yields more easily in one direction than in the other; such a property is possessed precisely by the tympanic membrane, as we have already had occasion to point out.
According to Helmholtz’s theory, the strength of combination tones is proportional to the square of the strength of the primary tones, and therefore they are especially noticeable when the primary tones are strong. Combination tones of the first order will have vibration numbers:
\[ 2n_1;\ 2n_2;\ n_1+n_2;\ n_1-n_2, \]
where \(n_1\) and \(n_2\) are the vibration numbers of the primary tones; combination tones of the second order will have vibration numbers:
\[ 3n_1;\ 3n_2;\ 2n_1+n_2;\ 2n_2+n_1;\ 2n_1-n_2;\ 2n_2-n_1; \]
in the same way, combination tones of higher orders will also be obtained.
The low combination tones stand out most strongly; theory shows that the first-order difference tone \(n_1-n_2\) must have great strength, which is also confirmed by experiment.
If the ratio of the vibration numbers of the primary tones is expressed by the irreducible fraction \(N_1:N_2\), then it is easy to see that the combination tones will have relative vibration numbers expressed by the series of natural numbers: 1, 2, 3, 4, and so on; this series will also include the primary tones, expressed by the numbers \(N_1\) and \(N_2\), and the difference tone \(N_1-N_2\), and all the summation tones. Thus, for example, if tones of 1200 and 700 vibrations/sec. sound, then the ratio of their vibration numbers (the interval) can be expressed by the irreducible fraction 12:7; the combination tones will be:
\[ 1,\ 2,\ 3,\ 4,\ 5,\ 6,\ \mathbf{7},\ 8,\ 9,\ 10,\ 11,\ \mathbf{12},\ 13 \text{, etc.} \]
In this series the difference tone \(12-7=5\) will have 500 vibrations/sec. Combination tone 1 is usually very clearly heard. If two tones sound whose relative number of vibrations is expressed by numbers differing by one, then combination tone 1 will at the same time also be the difference tone; it will have an increased
loudness. Since the harmonic overtones of any sound are expressed by a series of integers, it is clear that each pair of neighboring overtones will give the combination tone 1, i.e. will repeat the fundamental tone. Consequently, the combination tones of harmonic overtones reinforce the fundamental tone. The octave, as well as the following overtones, will be reinforced, though to a lesser degree, by the combination of non-neighboring overtones. These considerations explain many facts known from telephone practice, which we shall still have occasion to mention later.
In the case of simple intervals, the diversity of combination tones is greatly simplified. Thus, the interval of a fifth \(3:2\) gives only one low combination tone 1, always possessing considerable strength. The fourth \(4:3\) gives a strong tone 1 and a weaker tone 2.
In the case of intervals expressed by a ratio of large integers, the picture of combination tones is so varied and complex that even an experienced musical ear cannot make sense of it. Wegel and Lane\(^1\) in 1923 proposed a new method for detecting and measuring the strength of combination tones, based on the phenomenon of beats. If, to the consonance of two tones producing combination tones, a third sound of small intensity is admixed, which can be varied smoothly in pitch and intensity (the sound of a cathode generator), then, when this sound is not quite coincident with any one of the combination tones, we shall hear beats, which will be most distinct when the intensities of these two tones are equal. Thus, by retuning the third auxiliary sound over the whole scale of pitches, one can detect and measure the strength of all the combination tones. The authors investigated in this way the consonance of 1200 and 700 cycles/sec., produced by two cathode generators, i.e. the interval \(12:7\) (somewhat less than a minor seventh), and found the following combination tones:
\(1900\ (n_1+n_2),\ 500\ (n_1-n_2),\ 2400\ (2n_1),\ 1400\ (2n_2),\ 3600\ (3n_1),\ 2100\ (3n_2),\)
\(3100\ (2n_1+n_2),\ 1700\ (2n_1-n_2),\ 2600\ (2n_2+n_1),\ 200\ (2n_2-n_1),\ 2400\ (4n_1),\)
\(3800\ (2n_1+2n_2),\ 1000\ (2n_1-2n_2),\ 4300\ (3n_1+n_2),\ 2900\ (3n_1-n_2),\ 3300\ (3n_2+n_1),\)
\(900\ (3n_2-n_1).\)
Investigating the question of consonance and dissonance\(^2\) with the sounds of cathode generators, the author found that the phenomenon of consonance and the euphony of a chord are much more closely connected with the presence of combination tones than had previously been thought. The consonance of a fifth with a strong difference tone is transformed, in essence, into a single sound with harmonic overtones 1, 2, 3, etc., and under known conditions can almost not be decomposed by the ear into its component parts.
\(^1\) Wegel and Lane. Phys. Rev., 23, 266, 1924.
\(^2\) S. N. Rzhevkin. Izv. Physic. Institute under the Moscow Scientific Institute, 1, issue 2, 1926.
Investigating by the same method consonances and combination tones, Orlov1 found that combination tones are the principal factor determining the harmony of a consonance. Orlov discovered, moreover, for intervals larger than an octave, the existence of difference tones lying within the interval.
At the present time musical theorists are beginning to approach questions of harmony precisely from the point of view of combination tones. In recent works Garbuzov2 has shown, for example, that the concept of major and minor is by no means something immutable, inherent in definite intervals and consonances. Minor in its pure form can exist only in the region of low tones, where combination tones are not heard; in the region of high tones, minor consonances inevitably acquire the character of major because of the admixture of lower combination tones. Composers, of course, had felt this circumstance unconsciously even earlier, since the majority of works expressing sorrow are written in low keys.
The influence of combination tones on the character of a consonance in the case of more than two sounds in a chord had, until recently, been very little studied. Fletcher3 (Western Electric Co. laboratory) established that combination tones here play a far more important role than in the case of two sounds. Studying the distortions obtained under different conditions in the transmission of speech by telephone, Fletcher discovered the most interesting fact that the exclusion from the sound of the fundamental tone and the first overtones introduces almost no changes in the timbre and character of the vowels, although, as we shall see below, the chief part of the sound energy of speech lies precisely in the region of the fundamental tone and the first overtones. This exclusion of individual harmonics was carried out by means of electrical filters, which make it possible to pass through the telephone line any region of sound frequencies and to retain the rest.
Further Fletcher showed that appreciable distortions are obtained only when the high overtones, beginning with 1,000 cycles/sec. \((C^5)\) and higher, are excluded. Investigating the question in detail, Fletcher found that the character of the vowels does not change thanks to the subjective restoration of the excluded tones through the formation of low combination tones. Indeed, in the sound of all vowels there is always a region of strong high overtones, harmonic with the fundamental tone; for example, in the sound of the vowel a the harmonics are usually strengthened in the region from 800 to 1,100 cycles/sec.
and, as indicated above, each pair of such harmonics will give a combination tone equal in pitch either to the fundamental or to one of the low overtones, as a result of which a great strengthening of the low harmonics will be obtained. From these considerations it is clear why the exclusion of some low tones does not change the timbre: the low tones excluded from the composition of the sound arise again subjectively, with a loudness only slightly less than in the original sound. It is otherwise when the high harmonics characterizing the vowels are excluded, since they cannot be replaced; then the timbre of the sounds at once changes sharply and to the point of unrecognizability, and speech becomes completely unintelligible. The situation is approximately the same with the transmission of low sounds of musical instruments.
Fletcher showed, moreover, in a special experiment that, by composing a combination of three or more independent sounds, whose vibration numbers are related as a series of natural numbers, we always obtain the formation of powerful combination tones with relative vibration numbers 1, 2, 3, etc., of which the lowest tone 1 is the strongest. As a result of combining tones of 700, 800, 900, and 1000 cycles/sec., for example, there is obtained a strong common difference tone of 100 cycles/sec., as well as tones of 200, 300, etc., and the consonance of the four tones acquires the character of a single sound, with a beautiful timbral coloring and with a pitch of 100 cycles/sec. The addition of an objective tone of 100 cycles/sec., equal in strength to the high tones, changes nothing in the timbre of the sound.
A single sound of complex timbre with harmonic overtones cannot give any new combination tones lying below the fundamental. The formation of undertones, on which many musical theorists base their constructions, is impossible. The excitation by a given sound of fibers of the auditory membrane corresponding to tones 2, 3, etc. times lower (undertones), whereby in their vibrations they would be divided into 2, 3, and more parts, like a string when its overtones are excited, is hardly possible in view of the fact that the fibers of the auditory membrane have a rather massive structure and are hardly capable of making such sharp bends during vibrations. Thus, from this point of view as well, the occurrence of undertones is impossible.
The theory of combination tones makes it possible to foresee one very important point, little emphasized by Helmholtz himself—namely, the possibility of the occurrence of subjective overtones in the case when the ear is acted upon by only one simple sinusoidal tone. In a symmetrical membrane the elastic force may be considered proportional to the magnitude of the displacement from the position of equilibrium (Hooke’s law). In an asymmetrical membrane the elastic force is not the same for positive and negative displacements of equal magnitude. Under these conditions a sinusoidal air vibration with frequency \(n\) ev—
sets the membrane into vibration with a series of overtones \(2n, 3n, 4n\), etc. The presence of these overtones is very difficult for the ear to distinguish, since they merge into the sensation of a single sound. With the aid of the beat method described above, subjective overtones are easily detected. Indeed, if, against the background of one strong tone that is purely sinusoidal (which is easy to check by recording the curve), one causes another sinusoidal tone of variable pitch to sound, then clear beats are always found when the variable tone is 2, 3, 4, etc. times higher than the one under investigation. By varying the intensity of the variable tone, one can obtain the sharpest beats; in this case the intensity of the variable tone will be equal to that of the subjective overtone. Measurements1 show that subjective overtones at a loudness of 80 \(TU\) can already reach the intensity of the fundamental tone and even (at low tones) be stronger than it. Subjective overtones increase sharply with an increase in the primary tone, as was also to be expected theoretically. At loudnesses below 40 \(TU\) (10,000 times above threshold), subjective overtones no longer arise. Thus, the sensation of a pure tone is possible only at low loudnesses; strong tones inevitably acquire a timbral coloration, and the stronger they are, the greater this coloration.
The results of the experiments of Wegel and Lane on the question of the damping or masking of one tone by another are extremely interesting. If the threshold of audibility of tones of different pitches is considered known, then the question arises: will this threshold change when some tone of another pitch is sounding? Experiment shows that the threshold rises. The magnitude of the masking effect is therefore conveniently measured by the increase or shift of the threshold of perception of a tone in the presence of a second, masking tone, in comparison with its threshold in silence; thus the amount of masking will be measured in units of loudness.
The magnitude of the masking effect is greater the closer the masking and masked tones are to one another in pitch (although, with a small difference in frequency, beats occur, which facilitate the perception of one tone against the background of another, as a result of which the masking effect decreases). If the masking effect is represented graphically, plotting along the abscissas the pitch of the masked tone, and along the ordinates the increase in the threshold of sound perception against the background of the masking tone (in units of loudness \(TU\)), then we obtain something like a resonance curve with a dip at the top (see the lower curves of all the graphs in Fig. 6).
With an increase in the sound intensity of the masking tone, an entirely different phenomenon appears: tones lying below the masking tone are damped comparatively very little, while tones lying above the masking tone are damped—
weaken, on the contrary, very strongly, even at a great distance in pitch. The masking effect is especially intensified in the region close to the overtones of the masking sound, which is clearly visible from the curves of Fig. 6 for high loudnesses at frequencies of the masking tone of 800 and 1,200 cycles/sec. This observation should naturally be connected with the occurrence of subjective overtones.
Each of the graphs in Fig. 6 shows the masking effect of one of the tones: 200, 400, 800, 1,200, 2,400, 3,500 cycles/sec., at various loudnesses (20, 40, 60, 80, and 100 TU), on the tones of the entire musical scale from the very lowest up to 4,000 cycles/sec.
We see from the curves that loud tones mask all higher tones extremely strongly, and, to a much lesser degree, lower tones. A tone of 200 cycles/sec. masks almost the whole musical scale, whereas a tone of 3,500 cycles/sec., although it masks higher sounds, has almost no effect on the threshold of audibility of tones below 1,000 cycles/sec.
The correctness of these conclusions is easily verified on the sounds of an orchestra: a melody in low tones always stands out clearly from the general mass of sounds; the low sounds of the organ (the organ point) stand out especially well and, conversely, high instruments must produce a very great strength of sound for them to be heard against the background of low tones. In large orchestras, for dozens of violins there are 5–6 double basses, and nevertheless they are always clearly audible.
Observation of the masking effect makes it possible, by a method quite different from the former one, to approach the study of the mechanism of sound perception by the auditory membrane of the cochlea. It may be supposed with a certain degree of probability, as Wegel and Lane do, that the tone being investigated will be heard against the background of the masker when the amplitude of the vibrations of the auditory membrane caused by it becomes close in magnitude to the amplitude of the vibrations of the same membrane produced by the masking tone. Under this assumption we may consider that the masking curves in each graph of Fig. 6 characterize nothing other than the amplitude of the vibrations of the auditory membrane under the action of the masking tones, at their different loudnesses.
Examination of the curves shows that each tone produces vibrations of the membrane over a fairly wide interval; the sensation of pitch is apparently determined by the place of maximum amplitude. It should not be forgotten that the scale of the curves is logarithmic and that, consequently, on a linear scale the rise and fall of the curves would be much sharper and the maxima would stand out more strongly.
Special attention is drawn to the masking curve of a tone of 200 cycles/sec. at a loudness of 80 TU; this tone masks in the region around 1,000 cycles/sec. approximately 10 times more strongly than in the region of 200 cycles/sec.; this shows that the subjective harmonics of low
tones may be stronger than these tones themselves, as we pointed out in connection with the question of the lower limit of audibility.
The study of masking curves made it possible to calculate the hypothetical distribution of frequency perception along the auditory membrane, which is given in Fig. 2b. The curves of Fig. 6 provide rich material for the theory of hearing and make it possible to approach the question of the acuity of hearing and of the damping of the fibers of the auditory membrane somewhat differently than Helmholtz did when studying at what vibration frequency beats cease to be distinguished. Further investigation of this question will probably make it possible to arrive at an exact quantitative calculation of auditory perception and of the emergence of combination tones.
Measurement of the masking effect when the masking tone acted on the opposite ear made it possible to determine that sound, passing through the head to the other ear, is weakened approximately 100-fold, which coincides with measurements made on persons completely deaf in one ear.
We have seen that the formation of subjective overtones and combination tones in the case of polyphonic sounds occurs, at great sound strength, to a very considerable degree. This, of course, entails distortion of timbre. With an enormous increase in the intensity of speech, as occurs in modern loudspeakers, distortion of the timbre of vowels and consonants must inevitably result, entailing unintelligibility of speech. Therefore, for example, testing powerful loudspeakers in a small room is completely meaningless—they will necessarily give unintelligible speech, whereas in the open air they will work excellently. The same reasoning must also be applied to amplifying apparatus for the aid of the deaf. Strong sounds produced by these apparatuses will inevitably produce distortions of speech in the ear and reduce intelligibility. That is why the correction of severe deafness, although it would seem to be technically possible, since modern amplifiers provide almost unlimited amplification, in practice does not lead to the desired results because of the peculiarity of the mechanism of our ear.
IV. Determination by Hearing of the Direction of Sound.
Experience shows that the auditory apparatus is capable of determining the direction of an arriving sound wave with an accuracy of up to \(3\text{--}4^\circ\). This ability is a consequence of the fact that a person has two ears, for people deaf in one ear cannot determine the direction of sound.
Research in recent years has elucidated this question in considerable detail, both from the theoretical and from the experimental side. To explain the ability to determine directions, three assumptions are possible:
-
Judgment of direction is formed thanks to the fact that the ear turned toward the source of sound receives a more intense sound than the opposite ear. If the sounds perceived by both ears are equally loud, then the sound is localized by us in the median plane, that is, it seems to come from in front or from behind.
-
Judgment of direction is produced thanks to the ability to perceive the difference of phases arising because the sound reaches one ear earlier than the other.
-
Judgment of direction is formed thanks to the ability to perceive the interval of time between the arrival of the sound at one ear and at the other.
The theory of the diffraction of sound around a sphere, developed by Lord Rayleigh1, shows that the ratio of the intensities of the sound in front of and behind a sphere of the size of a human head is, generally speaking, very small and, moreover, rapidly decreases as the wavelength of the sound wave diminishes. This cause could explain the ability to perceive direction only at the very lowest frequencies.
Stewart2 showed experimentally that the difference of intensities is in fact not an essential factor in judging direction.
If the second assumption is correct, then for one and the same difference of phases of two sounds the same judgment of direction should result, independently of the pitch of the sound. Let us suppose that we perceive at one and the same angle two sounds of 100 and 1,000 cycles/sec. Since the velocity of sound and the distance between the ears along the line of propagation of the wave are constant, for both sounds the interval of time \(t\) between the moment at which the sounds reach the two ears will be the same. The phase difference, equal to \(2\pi Nt\), will obviously be 10 times greater for the tone of 1,000 cycles/sec. than for the tone of 100 cycles/sec. From these considerations it is clear that, if the second hypothesis were correct, the acuity of determination of direction would increase with frequency.
Experience shows, however, that the acuity of determining direction is the same at different frequencies and is equal to approximately \(3—4^\circ\); for large animals, for example elephants, there is reason to consider it higher. Thus the third assumption becomes the more probable.
Final clarity on this question was provided by the experiments of Hartley, Frey, and Stewart3. The experimental arrangement is as follows: two independent sounds of equal strength are supplied to the two ears, the phase difference of which can be measured exactly, and it is determined in what direction
is localized—the “apparent” source of sound—depending on the phase difference.
With a phase difference of 0, the sound always seems to be coming exactly from the front. When there is a phase difference, the sound “image” shifts toward the sound that leads in phase. With a phase difference close to \(180^\circ\), the displacement of the image is greatest; at low tones (\(100\) cycles) it is about \(180^\circ\), while at \(1000\) cycles/sec. it falls to \(40^\circ\). Thus, the position of the “apparent source of sound,” for a given phase difference, depends on the pitch of the sounds. As the phase difference is increased to \(180^\circ\), the sound image rapidly approaches, as it were enters the head, jumps to the other side as the phase passes through \(180^\circ\), and then follows a symmetrical path on the other side up to the median plane: with a phase difference of \(\varphi^\circ\) the position of the “image” is the same as with \((360-\varphi^\circ)\) or with \((-\varphi^\circ)\). At high tones, with considerable phase differences, not one sound “image” is observed, but several, which may be connected with the fact that several sound waves fit into the distance between the ears. Above \(1500\) cycles/sec. a judgment of direction by phase difference already becomes difficult. Stewart showed that between the phase difference \(\varphi\) of the sounds in the ears and the angle of displacement \(a\) of the “image” from the median plane there exists the relation:
\[ \frac{\varphi}{a}=kN+\frac{1}{2}, \]
where \(k\) is a constant.
At high frequencies this relation turns out to be practically a direct proportionality, since the first term is much greater than \(1/2\). Since, as a first approximation, it may be assumed that with a phase difference \(\varphi\) there is obtained a certain difference in the arrival times of identical phases at the two ears:
\[ \Delta t=\frac{\varphi}{2\pi N}, \]
then the dependence written above for high frequencies may be approximately expressed in the following form:
\[ \frac{a}{\Delta t}=\frac{2\pi}{k}=\mathrm{const}. \]
Thus, from Stewart’s experiments it follows approximately that the angle of displacement of the sound “image” is proportional to the difference in the arrival times of identical phases of the sound wave at the two ears. The value \(2\pi/k\) proves to be equal to about \(1570\), that is:
\[ a=1570\,\Delta t, \]
where \(\alpha\) is measured in angular measure. This conclusion argues for the correctness of the third hypothesis.
Assuming that a person estimates the deviation of a sound from the median plane at \(3^\circ\), we shall find that this corresponds to the perception of a time interval of about 0.00003 sec. Hornbostel and Wertheimer1, observing short sound impulses, proved that judgment of direction depends precisely on the difference in the times at which the sound arrives at the two ears. They determined the minimal perceptible difference in times to be just 0.00003 sec. With a difference in times of 0.0006 sec., there is already the sensation of a sound arriving at an angle of \(90^\circ\).
The totality of all the experimental data thus speaks for the correctness of the third of the proposed assumptions.
The regularities established in this field have made it possible to carry out a number of practical applications, making use of the human ability to determine the direction of a sound. If the sound base is increased and the sound is brought to the ears from two horns situated at a great distance (the same from both ears), then the minimal perceptible time interval of 0.00003 sec. will be obtained not at a deviation of \(3^\circ\) from the median line, but at a much smaller angle. Instruments built on this principle—acoustic direction finders—make it possible to determine direction with an accuracy of up to \(1/2^\circ\) and less. These instruments find application in military affairs.
Sensitivity in estimating direction is greatest in the median plane. This circumstance can be used to estimate a small change in phase difference introduced by the passage of sound through a certain medium (for example, a gas), whose properties it is desired to study; by compensating the resulting phase difference, forcing the second sound to pass a known path in air, and thereby bringing the sound back into the median plane, one can easily compare the velocity of sound in air and in the substance under investigation.2
V. Study of the Human Vocal Apparatus from the Acoustic Point of View.
I do not consider it necessary, within the scope of this article, to dwell on a description of the structure of the vocal apparatus, even as briefly as I did with respect to the auditory apparatus. The structure of the vocal apparatus is anatomically considerably simpler and much better studied, owing to the possibility of studying it while it is functioning in living people by means of laryngoscopy and X-ray photographs. Recently
a Russian translation of Musehold’s book, The Acoustics and Mechanics of the Vocal Apparatus, has appeared, in which one may find a detailed exposition of questions concerning the anatomy and physiology of the voice. In what follows I shall deal chiefly with the physical phenomena involved in voice production and with the question of the objective analysis of the sound of speech.
The vocal apparatus consists of three principal parts: 1) the lungs, with the system of inspiratory and expiratory muscles; 2) the larynx, with the vocal cords or “lips,” in Musehold’s terminology; and 3) the system of air cavities that play the role of resonators and sound radiators.
As a physical instrument the vocal apparatus is quite fundamentally compared with a reed organ pipe. Of its three parts, the cords are, of course, the least like the reed of an organ pipe, but in principle their role is the same. During voice production the cords are two muscular ridges, standing parallel and pressed against one another; the degree of tension of these ridges determines the pitch of the tone produced. The voice may vary in pitch over a range of more than two octaves, but alongside the normal—chest—voice a person can produce special, higher sounds—this is the so-called falsetto, or fistula. I shall dwell somewhat more fully on the mechanism of voice production (phonation) in the chest voice and in falsetto, in view of the great interest of this question, which has been elucidated in detail only recently by Musehold with the aid of laryngoscopic photography; moreover, by combining it with a stroboscope he succeeded in following the separate phases of the opening of the glottis.
The tension of the vocal cords may occur either through the contraction of the muscular fibers that make up the cords themselves—this is the so-called internal tension—or by the action of other muscles of the larynx, which stretch the cords passively—this is external tension. The character of the functioning of the cords in these two cases is entirely different. In instantaneous photographs of the glottis in the chest voice, the cords appear like two thick, tense muscular ridges resembling lips, pressed tightly against one another; this corresponds to contraction and to “internal” tension of the cords. During phonation the opening of the glottis occurs only for a very brief moment during a small part of the period, and in this time a strong burst of air passes through it. The periodic succession of such impulses gives a sound rich in overtones, with a metallic timbre. With this sound the anterior chest wall gives a strong vibration (fremitus pectoralis), perceptible by the hand, from which this type of voice has received the name chest register.
In falsetto the cords appear flat and strongly stretched, and between them in the middle a gap is formed, as between two tightly stretched rubber strips. The edges of the cords are thinned out.
and, apparently, do not have the same elasticity as in the chest voice; during phonation they vibrate, moving upward and sideways, chiefly at the edge of the cord; complete closure of the glottis is not obtained even at the phase of the greatest convergence of the cords. Thus there is no complete interruption of the air stream; only its weakening and strengthening take place. As a result, the falsetto voice is not rich in overtones; it sounds very soft, does not possess a metallic shade (high overtones), and has no strength. There is also no trembling of the chest wall.
Dayton Miller’s investigations in America, completed by 1914 and described in his extremely interesting book, The Science of Musical Sounds1, represent a substantial step forward in the study of the sound of the voice. To record sound Miller uses an instrument called by him a “phonodeik,” consisting of a thin glass membrane with a horn, which transmits its vibrations to a small mirror rotating on an axis. The instrument is very sensitive and registers sounds up to 10,000 vibrations/sec.; in principle this instrument is not new, but it is very well constructed.
Fig. 7. “Spectra” of the voice of a soprano and a bass on the vowel a.
For calibrating the instrument for sound intensity Miller uses a set of closed organ pipes from 129 to 4,138 vibrations/sec., of equal loudness to the ear. Such calibration proved to be an extremely painstaking and difficult task and, according to the author’s calculation, took in all 3½ years of daily 8-hour work by one investigator. As a result of this work, however, it became possible to take into account the distortions introduced by the various horns and membranes, and to calculate the relative intensity of the harmonic overtones of different sounds. By using pipes of equal loudness (subjective), Miller, of course, simplified the problem, but, by not taking into account the decrease in the sensitivity of the ear in the region of low tones,
introduced a considerable error. It should be thought that all of Miller’s analyses give a sound strength for low tones diminished by tens of times; the error becomes smaller the higher the tones.
Miller presented all his analyses in a very convenient graphical form, which later received the name “sound spectra.” Fig. 7 represents, for example, the “spectrum” of the vowel a (as in the English word father), sung by a soprano at a pitch of 488 cycles/sec. and by a bass at a pitch of 91 cycles/sec.; along the abscissa are laid off, on a logarithmic scale, the pitches of the tones, and along the ordinate—the intensity of the individual harmonics in the form of little pyramids. From the spectrum of the vowel a for the bass it is evident that the 2nd, 8th, 9th, and 11th harmonics are comparatively strong; the 8th harmonic—732 cycles/sec.—stands out especially. The intensity of the fundamental tone and of all harmonics up to the 6th, according to this analysis, is so small that it cannot be indicated in the drawing. For the soprano the 2nd harmonic, 977 cycles/sec., is especially strong. As we shall see further on, Miller’s data concerning the strength of the low harmonics are erroneous.
Fig. 8. Position of the formants of various vowels.
In Fig. 8 is shown the general result of the analysis of various vowels, flowing from Miller’s work: the vowels a, oa (intermediate between a and o), o and u, at whatever pitch they may be pronounced, always have one definite region of strengthening of the overtones; for a it is the highest, about 1000 cycles/sec. (c³), for u—the lowest—about 300 cycles/sec. Each of the vowels ae, e, eu, i has two characteristic regions—one low and one high; for i, for example, they lie around 300 and 3000 cycles/sec. The presence of characteristic regions for the vowels is in direct dependence on the shape and volume of the upper resonant cavities, which was firmly established by Helmholtz as early as the sixties of the last century.
The sound produced by the vibration of the vocal cords, as well as by the vibration of the tongue of an organ pipe, is completely modified by the presence of resonant cavities. The resonant cavities of the pharynx, nasopharynx, and mouth play above all the role of amplifiers of those overtones of the sound of the cords which lie close to the natural tone of the cavity. In view of the fact that the walls of these cavities are soft, their resonance is not confined to a narrow
region of tones; a deviation from the cavity tone by a major third (by 25%) reduces the resonance only by half. As a result, such a resonant cavity amplifies the overtones of the voice over the extent of almost an entire octave. In the example of Fig. 7 we saw that the 7th, 8th, 9th, and 11th overtones are amplified. Thus it is clear that the timbre of the sound of the voice will depend substantially on the tuning of the resonator cavities, chiefly the cavities of the mouth and pharynx. The oral cavity is, moreover, a resonator with a large opening into the external space and therefore is a sound radiator, or, using a term from radiotelegraphy, a sound antenna.
The cavity of the nasopharynx, since it lies to the side, may be a kind of sound filter1: it may absorb certain tones and not let them pass outward; this role of the nasopharynx is especially important in singing for imparting a certain coloration to the sound of the voice.
The form of the cavity of the pharynx and mouth in the pronunciation of various vowels is shown in Fig. 9.
Fig. 9. Form of the oral cavity for various vowels.
The theory of vowel formation set forth belongs to Helmholtz2. The physiologist Hermann3 views the matter somewhat differently: he assumes that each puff of air produced by the cords excites oscillations of the oral resonator at its characteristic frequency, and then these oscillations rapidly die away until they are again excited by the next puff of air. Thus the vowel curve must consist of an alternation of series of damped oscillations having a pitch determined by the proper period of the resonator, and following one another with the frequency of oscillation of the cords. According to Hermann’s view, the tone of the resonator excited in this way may also fail to be a harmonic over-
connection tone. For the sound of each vowel there are characteristic only rhythmic alternations of a tone of a definite pitch for that vowel. Hermann calls the proper tone of the resonator cavity the formant of the vowel.
From the curves of recordings of the sound of vowels, taken by Hermann himself and, in recent years—by improved methods, which will be discussed below—by Miller, Trendelenburg, and others, it is evident that the separate series of damped vibrations do indeed follow one another, and moreover follow after exactly equal intervals of time (the period of the fundamental tone of the ligaments), and are absolutely identical with one another. In recordings of vowels one never obtains periods unlike one another. This already serves as an undoubted indication that in the composition of the sound of a vowel there are no inharmonic overtones, although the resonator cavity may well oscillate with its own period, inharmonic to the fundamental tone.
Fig. 10. Spectra of the vowel a in the singing and speaking registers of the voice.
If we recall that, as a result of the superposition of two neighboring harmonic overtones, there will always be obtained a curve of beats with a beat frequency equal to the frequency of the fundamental tone, and if we take into account that pulsations of a more complex form can always be composed as a result of adding a larger number of harmonics, and if we take into account the identity of the separate periods, then we shall come to the conclusion that Hermann’s view and Helmholtz’s view are essentially identical. Whether one speaks of a formant pulsating with the frequency of the fundamental tone (Hermann), or of a combination of a series of harmonics which in sum give the same curve of one form (Helmholtz), is mathematically immaterial. In exactly the same way, the beat curve of a tone of a definite pitch is identical with the curve obtained from the addition of two undamped tones close in pitch.
The term “formant,” introduced by Hermann, quickly took root in science, but it is usually applied not as a designation of a definite—
…tone characteristic of the given vowel, but as a characteristic region of resonance, in which a series of harmonics of the sound produced by the vocal cords is amplified.
As a result of a detailed study of the sound of the voice, carried out by me jointly with V. S. Kazansky by means of Kazansky’s “sound oscillograph”1, it proved possible to obtain some new data on the composition of the sound of the singing voice. The first undoubtedly noted feature of singing sound is the extraordinary sharpness of the resonance of overtones in a definite region. In Fig. 10 a and b are given the spectra of the vowel a for the singing and speaking voice at a pitch of 129 vibrations/sec.; in Fig. 10 c and d are given analogous spectra of the vowel a at a pitch of 259 vibrations/sec. The sharpness of the resonance can perhaps be explained by an increase in the rigidity of the walls of the resonating cavities (pharynx, fauces, mouth). It is interesting to note that for singing sound, for almost all vowels, there is usually observed a sharp reinforcement of one definite overtone; for \(c = 259\) vibrations/sec. this will be, for example, the second
Fig. 11. Recording of the voice of artist Vorob’eva at the note \(e'\)—325 vibrations/sec.
overtone—517. Thus it becomes clear why well-trained voices have a similar “singing” timbre on all vowels.
The second characteristic feature of singing sound is the extraordinary variability of the curve, the presence of a whole series of rapid and irregular pulsations. Without these pulsations the voice acquires a lifeless character. Fig. 11 gives the sound curve of the voice for a high note (\(e'\)—325 vibrations/sec.) of a baritone. In recordings of the singing voice one very often encounters curves of a nonperiodic character, or, more precisely, gradually changing periodic curves, which already indicates the undoubted presence of nonharmonic overtones. Finally, the presence is noted, for the vowels a and o, of two low resonance regions, and not one, as had hitherto been accepted. I shall not dwell here on a number of other, less interesting data on the singing voice.
VI. Analysis of Speech
At the present time we have a whole series of highly important investigations on the question of the analysis of speech, its transmission by wires and radio, and reproduction with great loudness. Not having the possibility here to touch upon
in the scope of this article, of the last two questions, as the more specialized ones, I shall dwell only on the first, namely on the analysis of speech.
The American physicist Wente1 found a perfect method for perceiving sound and converting it into electrical form without distortion. The instrument constructed by him is a condenser consisting of a fixed massive cylindrical electrode, 5 cm in diameter, and of a taut, very thin steel membrane lying at a distance of several hundredths of a millimeter from it and electrically insulated from it. If such a “condenser microphone” is connected into a battery circuit, then, when the membrane vibrates, the charge of the condenser will change, and these charge oscillations can be amplified to any degree by means of vacuum-tube amplifiers and recorded by any method, for example by an oscillograph or a string galvanometer.
The membrane of Wente’s condenser microphone had a very high natural period—up to 12,000 vibrations/sec., which is increased still further, owing to friction in the air layer between the membrane and the back electrode, to 17,000 vibrations/sec. In the region of the sounds of speech and music such a membrane has no resonant properties and makes it possible to perceive sounds from 25 to 8,000 vibrations/sec. without distortion. The steel membrane can successfully be replaced by thin silk, onto which aluminum foil several microns thick is pasted2.
The starting point for the study of speech as a physical phenomenon was provided by investigations of the sensitivity of hearing, which made it possible to establish, in connection with the analysis of the sound of speech, the distribution (average statistical) of speech energy by frequencies and the relative loudness of speech in relation to the threshold of audibility. A summary of the results of these works is given in Fig. 4, where the region of speech is the shaded part of the entire region of auditory perception.
Determining the average distribution of speech energy by frequencies was extremely important for telephone practice, in order to decide the question of the importance of particular frequencies in the transmission of speech. To solve this problem, before the condenser microphone there was pronounced a series of syllables consisting of vowels and consonants in the same proportion as they occur in actual speech; the resulting electrical oscillations of sound frequency were amplified, recorded, and, after analysis, represented in the form of sound spectra. The total result of the study of the distribution of energy in the “spectrum” of speech, carried out by Crandall and Mackenzie3, is shown in Fig. 12, curve A. This curve refers to the average human voice;
it has a sharp maximum in the region of 150 cycles/sec. In the spectrum of female speech the maximum lies at 200 cycles/sec., and in the spectrum of male speech—at 100 cycles/sec. Since vowels constitute the principal component of speech, the distribution of speech energy can be constructed on the basis of Miller’s analyses. In this way one obtains curve \(B\), with a sharp maximum in the region of 700 cycles/sec.; the discrepancy between these curves is explained by the fact that in all of Miller’s analyses the energy of the low frequencies is underestimated.
From the data given it would seem natural to conclude that, for the perfection of speech transmission by telephone, one must be concerned chiefly with ensuring that the region of low frequencies from 100 to 200 cycles/sec. is transmitted without attenuation, and with this in mind construct the transmitting and receiving instruments. However, experience has shown that the matter is entirely different.
Fig. 12. Distribution of energy in the spectrum of conversational speech, \(A\)—according to Crandall and Mackenzie, \(B\)—according to Miller’s data.
At the laboratory of the Western Electric Co. the intelligibility of speech transmission was investigated under the condition that a certain frequency range was absorbed by means of electric filters. The experiment gave a remarkable result: it turned out that, when all frequencies up to 500 cycles/sec. are excluded, the intelligibility or articulation of speech1 falls by only 5% in comparison with normal; the exclusion of all frequencies up to 1,000 cycles/sec., while 82% of the total energy of speech is absorbed, reduces intelligibility by only 15%. Absorption of the frequency region above 1,000 cycles/sec. immediately entails a sharp decrease in intelligibility—
...of intelligibility. It is interesting to note that the timbre of speech, when low frequencies are excluded from it, becomes cutting, with “metallic” overtones. When all high frequencies are excluded, the timbre acquires a dull, “dark” character. Excluding all frequencies above 1,000 cycles/sec., which carry only 18% of the energy, reduces intelligibility to 40% of normal.
From these experiments it is easy to calculate the relative importance of the various frequencies for the intelligibility of speech; from the curve in Fig. 131 we see that the most important region is that of the frequencies in the vicinity of 1,000 cycles/sec., and that in general high frequencies are more important than low ones. Here one must recall with surprise that the resonance of the majority of iron telephone diaphragms lies precisely close to 1,000 cycles/sec. The empirical selection of thickness and diameter evidently led precisely to the diaphragm dimensions best suited for clarity of transmission.
Fig. 13. Relative importance of different frequencies for the intelligibility of speech.
The observations indicated concern only the intelligibility of speech. If we turn to the transmission of music or singing, then we encounter entirely different characteristics of timbre, connected with the beauty and fullness of the sound, and here the task of designing receiving and reproducing apparatus faces new difficulties, not yet overcome. I shall point out, as an example, that the loudspeaker of the weak-current trust familiar to us, with a paper diaphragm—the so-called “diffuser”—when investigated in the acoustics laboratory of the State Experimental Electrotechnical Institute, showed excellent articulation, but for the transmission of music and singing it is of very little use because of its unpleasant “paper” timbre, which produces an inartistic impression.
The reason for the insignificance of timbre distortions when low frequencies are excluded was clarified in his investigations, already mentioned above, by Fletcher. The point is that the mechanism of the ear has the ability to combine low tones as a result of the sounding together of several high overtones. As a result of this property of the ear, the ear is able subjectively to restore excluded low tones as a result of the combination of upper harmonics.
Research on sound by improved methods has led to the discovery of certain new features of vowels, previously unknown, and also to a clarification of the nature of consonants, which had previously been little known in general. Trendelenburg1 developed a highly sensitive method of recording sound without distortion by means of an oscillograph and a condenser microphone. As a result of his work he found that the vowels a, o, and u always possess, in addition to a low formant, further characteristic overtones in the high region above 3,000 cycles/sec. These overtones are characteristic not of the vowel, but of the individual timbre of the voice; they are especially strong in the vowel a. The earlier recording methods were apparently insufficiently sensitive in the region of high tones to detect these features. Trendelenburg also determined anew the regions of the formants of all vowels, and in this he found complete agreement with earlier work; with Miller’s data there is a divergence in the determination of the intensity of low frequencies, which was to be expected for the reasons already indicated by us. Low frequencies in vowel sounds have considerable intensity; it is interesting to note that the voice of a trained singer gives a stronger fundamental tone and second harmonic (octave), but a less sharply expressed vowel character, which is in full agreement with my investigations and those of Kazansky.
Stumpf,2 investigating the timbre of vowels by ear, also found high characteristic overtones in the same regions as Trendelenburg.
An entirely new method of recording sound was proposed by the well-known physiologist Einthoven.3 This method is apparently entirely free from distortions and permits the recording of the highest sounds up to 15,000 cycles/sec., but it does not possess great sensitivity. An exceedingly thin quartz filament less than 1 μ in diameter is suspended, in a slightly stretched state, in the mouth of a horn; it is so light that it is freely carried along by the most rapid oscillations of the air. The oscillations of the filament are observed, as in Einthoven’s string galvanometer, through a microscope under strong illumination and are photographed on a moving strip. In the photographs obtained of the sound of the vowels u and a, the presence of high overtones in the region above 2,000 cycles/sec., discovered by Trendelenburg, is noted.
An extraordinarily complete analysis of speech sounds was carried out by Crandall4 in the laboratory of Western Electric. He recorded several hundred sounds of vowels and consonants with the aid of a high-frequency oscillograph, and all the recordings were subjected to analysis and presented in the form of spectra.
The most important result of this work consists in the fact that, for all vowels, 2 resonance regions (formants) have been found. A summary of these Krandall data for various vowels (in English pronunciation) is given in Fig. 14. Earlier investigators, beginning with Helmholtz and ending with Miller, found for the vowels u, o, and a only one resonance region each.
Fig. 14. Position of the formants of various vowels according to Krandall.
The calculation made by Krandall shows that the presence of two resonance regions can be explained by the presence of two resonating cavities: the posterior one—the pharyngeal cavity—and the anterior one—the oral cavity, separated by the narrowing formed by the root of the tongue and the soft palate. The result of calculating the vibrations of such a coupled system of two resonators agrees very well with the experimental material for all vowels.
Richard Paget1 in 1923 quite successfully reproduced, with the aid of double resonators and an artificial larynx, the sound of the majority of vowels. All this evidently speaks in favor of the correctness of the theory of double resonance and definitively clarifies the question of the origin of the various vowels.
In concluding the survey of methods for analyzing the sound of vowels, it is necessary to note the low sensitivity of the method of recording curves in the frequency region above 3000. This is due both to the inertia of the recording devices and to the difficulty of accurately decomposing the curve obtained. The curves obtained must be decomposed into a Fourier series in order to obtain the sound spectrum. The procedure of harmonic analysis, even when analyzers are used, is very complicated and requires a great deal of time, which must especially be kept in mind given that tens and hundreds of curves have to be subjected to analysis.
In 1923 I proposed2 the design of an apparatus that makes it possible to record sound spectra entirely automatically.
The idea of the instrument is analogous to the method of Helmholtz’s sound analyzer, based on detecting the sound of harmonics by means of resonance, but with the replacement of acoustic resonators tuned to different frequencies by a single electrical resonant circuit with smooth tuning over a wide range. This apparatus, which may be called an acoustic wavemeter, makes it possible to reduce the time needed for a complete recording of a sound spectrum to a few seconds. The sensitivity of this method in the region of high frequencies is much greater than that of the curve-recording method.
It remains for us to touch only on the analysis of the sound of consonants. The recording of curves with the aid of a condenser microphone enabled Trendelenburg1 to solve this more complex problem as well. The curves obtained show a clearly expressed period corresponding to the tone of the vocal cords that underlies these consonants, but in addition they contain oscillations of a high period, obviously non-harmonic with the fundamental tone, since within each period the pattern of small oscillations on the curve is entirely different. It proves possible to pronounce the consonants л, м, and н in such a way that they give a completely periodic curve, but then the character of the consonant changes considerably: л, м, and н may be called “semivowels.” Thus it may be considered that consonants are a mixture of several sounds that are not harmonic with one another.
Fig. 15. Characteristic regions of the consonants Т, М, Н, Р, С, Ш according to Mumphu.
Where, then, can another sound arise besides the sound of the cords? In the case of the consonants л, м, and н, this question is comparatively easy to answer. The air emerging through the glottis and setting the cords into vibration then exits not through the wide opening of the mouth, as in the case of vowels, but through a more or less narrow opening (in the case of л—between the tongue and the teeth; in the case of м—through the nose; in the case of н—through the clenched teeth) and therefore can set into vibration, by blowing, the cavities through which it passes. Thus the resonating cavities participate in sound formation, on the one hand amplifying certain harmonics of the fundamental tone of the cords, and on the other hand—they come
into independent oscillations and produce sounds inharmonic with the tone of the vocal cords.
Hermann conceived the origin of vowels in precisely this way, but in the curves recording vowels each period is similar to another, and consequently there are no inharmonic overtones in them. To verify Hermann’s theory, Trendelenburg recorded the sound of the vowel a with an inordinately large air leak and found in the curve a certain inequality of the individual periods, i.e., the presence of inharmonic overtones. Thus, excitation of independent—and therefore in some cases inharmonic—oscillations of the air resonators is also possible in the case of vowels, provided that the air stream is poorly used for exciting the sound of the vocal cords.
Fig. 15 presents Stumpf’s table1, which shows the position of the regions of characteristic resonance (formants) of certain consonants. Trendelenburg’s data are generally in agreement with this table, but in addition he finds in these consonants still higher overtones than Stumpf was able to detect by ear. As a result of numerous recordings and analyses it has been established that, for perfect transmission of speech, there must be no distortions in the region from 50 to 5,000 cycles/sec. Most important in this respect is the region from 1,000 to 3,000 cycles/sec., distortion of which makes speech completely unintelligible. In recordings of the sound of the consonant p, beats of the sound with a frequency of from 20 to 40 per second are characteristic. The consonant ш has a high characteristic overtone in the region of about 4,000 cycles/sec.; the consonant с has overtones still higher, up to 5,500 cycles/sec., and according to some studies even up to 8,000 cycles/sec.