STEADY-STATE PROCESSES IN ACOUSTICS[^1]
G. Backhaus
Submitted 1938 | SovietRxiv: ru-193801.92522 | Translated from Russian

Full Text

STEADY-STATE PROCESSES IN ACOUSTICS1

H. Backhaus

Contents

I. Introduction.
II. Theory of nonstationary processes:
1. Systems with one degree of freedom, 2. Systems with many degrees of freedom, 3. Symbolic methods, 4. Approximate calculations of steady-state processes.
III. Auditory perception of nonstationary sound processes.
5. Model of the ear and steady-state processes, 6. Perception of pitch, 7. Auditory perception of steady-state processes, 8. Modulated tones.

I. Introduction

The human ear has the ability to distinguish stationary sounds from one another by their intensity, pitch, and timbre. The intensity of a sound is determined by the amplitudes of the fundamental tone and overtones; pitch is identified with the frequency of the fundamental tone; and timbre is the result of the superposition upon one another of harmonic overtones with different amplitudes that enter into the composition of a complex sound. This conception of timbre, called by Helmholtz2 “musical timbre,” was for a long time almost the sole subject of investigation in musical acoustics. With the aid of increasingly refined electroacoustic apparatus it became possible to discover ever finer details characterizing stationary sounds. In the theory of the formation of vowels, considerable progress was made in establishing the characteristics of different vowels as functions of the presence of clearly expressed regions of overtone amplification (formants).

Considerable difficulties were first encountered in attempts to establish the distinguishing features of the sounds of musical instruments. Only for very few instruments did it prove possible to establish definite frequency regions whose amplification

characteristic of the corresponding musical instrument. From this it followed that the concept of “musical timbre” by no means exhausts the acoustic characterization of various musical sounds. Of particular importance are the investigations of Stumpf ^20, from which it follows that the acoustic features characterizing each individual musical instrument are determined not so much by stationary sounds as by nonstationary phenomena.

This point of view is also of very great importance for technical acoustics. Previously, in the manufacture of electroacoustic receiving and transmitting devices and apparatus, discussion was limited to the question of the frequency dependence of transmission only for steady-state processes. Taking into account the difficulties involved in eliminating resonance in large mechanical oscillatory systems—for example, in loudspeakers—it is usually considered sufficiently satisfactory if it is possible to achieve a more or less uniform frequency characteristic of the apparatus covering a sufficiently wide range of frequencies. Within this range it is considered possible to be satisfied with slight irregularities of the frequency characteristic, all the more so since the sensitivity of the ear to amplitude distortions is very small.

In doing so, however, no attention was paid to the fact that such slight peaks of the resonance curve, indicating weakly damped natural frequencies of the apparatus, although they have no practical significance in the steady-state mode of transmission, nevertheless in non-steady-state processes give rise to free oscillations which decay the more slowly the more sharply the corresponding peaks of the resonance curve are expressed. As a consequence of this, under ordinary operating conditions, as a result of repeated impulses there arise extraneous frequencies characteristic only of the given apparatus and significantly distorting the correctness of transmission.

For these reasons, nonstationary processes have already for several years been studied in acoustics. The numerous substantial results obtained in this field will be set forth below. However, the very extensive field of architectural acoustics, in which nonstationary processes play an exceptionally large role, will not be included in the present review, since it requires special consideration.

II. Theory of Nonstationary Processes

1. Systems with one degree of freedom.

The motion of a system with one degree of freedom is described by the following differential equation:

\[ m\frac{d^2x}{dt^2}+r\frac{dx}{dt}+cx=P\sin\omega t, \]

where \(m\) denotes the mass of the system, \(r\) the coefficient of friction, and \(c\) the coefficient of elasticity.

The general solution of this equation has the following form:

\[ x=Ae^{-\delta t}\sin(\omega_0 t+\psi)+B\sin(\omega t-\varphi), \tag{1} \]

where

\[ B=\frac{P}{m\sqrt{(\Omega_0^2-\omega^2)^2+\frac{\omega^2 r^2}{m^2}}} =\frac{P}{m\rho};\qquad \Omega_0=\sqrt{\frac{c}{m}}; \]

\[ \varphi=\operatorname{arc\,tg}\frac{\frac{\omega r}{m}}{\Omega_0^2-\omega^2}; \qquad \delta=\frac{r}{2m}; \qquad \omega_0=\sqrt{\frac{c}{m}-\frac{r^2}{4m^2}}. \]

\(A\) and \(\psi\) denote arbitrary constants depending on the initial conditions; their calculation is usually rather difficult. If it is assumed that the system was initially at rest and that the applied sinusoidal force began to act suddenly at the time \(t=0\), then the arbitrary constants \(A\) and \(\psi\) take the following particular values:

\[ A=\frac{P\omega}{m\rho^2\omega_0} \sqrt{4\delta^2\omega_0^2+(\Omega_0^2-\omega^2-2\delta^2)^2}; \qquad \operatorname{tg}\psi= \frac{2\delta\omega_0}{2\delta^2-(\Omega_0^2-\omega^2)}. \]

The actual motion of the system in the transient period is the result of the superposition of two oscillatory processes: an undamped one with amplitude \(B\) and a damped process with initial amplitude \(A\).

If the frequency of the exciting force \(\omega\) is close to \(\omega_0\), the natural frequency of the system, then beats arise, damped with damping coefficient \(\delta\). The exponential factor may be represented in the following form:

\[ e^{-\delta t}=e^{-\frac{r}{2m}t}=e^{-d\frac{t}{T_0}}, \]

where \(d=\dfrac{\pi r}{\omega_0 m}\) denotes the logarithmic decrement, and \(T_0\) is the period of the natural oscillations of the system. Hence it is clear that the rate of damping is determined by the magnitude of the logarithmic decrement, or damping\(^{1}\) of the oscillatory system. The stronger the damping, the sooner the stationary oscillations are established. After a time \(\dfrac{T_0}{d}\), the process being established decreases to \(e^{-1}=37\%\) of its initial value.

The relative magnitude of the process being established can

\(^{1}\) By damping the author evidently means the quantity sometimes used in the scientific literature, denoted by \(D\) and defined in terms of the logarithmic decrement \(d\) by means of the relation \(D=\dfrac{d}{\pi}\).

Translator’s note.

to define as the ratio of the amplitude \(B\) to the initial amplitude \(A\). This quantity depends on the relation between the natural frequency of the system and the frequency of the disturbing force. If the natural frequency of the system is very small in comparison with the frequency of the disturbing force, i.e., if \(\Omega_0 \ll \omega\), then

\[ B=\frac{P}{m\omega}\frac{1}{\omega}; \qquad A=\frac{P}{m\omega}\frac{1}{\Omega_0}, \quad \text{i.e. } B \ll A \]

In the case where \(\Omega_0=\omega\), then \(B=\frac{P}{\omega r}\), \(A=\frac{P}{\omega r}\frac{\Omega_0}{\omega_0}\); consequently,

with small damping \(B \simeq A\).

Finally, if the natural circular frequency of the system is considerably greater than the circular frequency of the forcing force, i.e., if

\[ \Omega_0 \gg \omega, \quad \text{then } B=-\frac{P}{m\Omega_0^2} \quad \text{and} \quad A=-\frac{P}{m\Omega_0^2}\frac{\omega}{\omega_0}, \quad \text{i.e. } B \gg A. \]

It follows from the above analysis that the process of establishing oscillations causes the greater distortions, the higher the frequency of the exciting force in comparison with the natural frequency of the excited system.

The general integral (1) also makes it possible to obtain a solution for the process of establishing equilibrium when the action of the exciting force ceases. Namely, if the applied force ceases to act at the moment when its magnitude is equal to zero, then the process of establishing equilibrium can be represented by the following equation: \(x=Ae^{-\delta t}\sin(\omega t+\psi)\), i.e., by means of the first term on the right-hand side of the general integral (1).

2. Systems with several degrees of freedom

For calculating transient processes in systems with two degrees of freedom, it is necessary to determine the arbitrary constants from the initial conditions; however, this determination is so cumbersome that in practice it is almost never carried out. In this case, in order to obtain a rigorous solution for transient processes, a special method is used, consisting in the fact that the nonperiodically applied force, as a nonperiodic function, is expanded into a Fourier integral series consisting of an infinite number of sinusoidal oscillations with suitably chosen amplitudes, and then, on the basis of the superposition principle, the final solution is obtained as the result of summing the individual particular solutions.

The sudden application at the time \(t=0\) of a constant force with maximum value \(K\) can be represented in the form of the following expression:

\[ k=K\left[\frac{1}{2}+\frac{1}{\pi}\int \frac{\sin \omega t}{\omega}\,d\omega\right], \tag{2} \]

where the binomial standing in parentheses takes the following values: for \(t<0\), \(0\); for \(t=0\), \(\dfrac{1}{2}\); for \(t>0\), \(1\).

Formula (2) can be represented in the following complex form:

\[ k=\frac{K}{2\pi j}\int_{-\infty}^{+\infty}\frac{e^{j\omega t}}{\omega}\,d\omega . \tag{3} \]

Let the constant force \(K\) exert on the system the stationary action \(\dfrac{K}{W}\), where \(W\) denotes a certain function of \(\omega\). The equation of motion of the system may be represented as follows:

\[ x=\frac{K}{2\pi j}\int_{-\infty}^{+\infty}\frac{e^{j\omega t}}{W(\omega)\omega}\,d\omega . \tag{4} \]

The integrand has singular points at \(\omega=0\) and \(\omega=p_\nu\), where \(p_\nu\) is a root of the characteristic equation \(W(\omega)=0\). The general case, when the equation has multiple roots, was analyzed by K. W. Wagner\(^7\). Here we shall restrict ourselves to the case in which the roots of the equation are finite and simple. Further, it is not difficult to see that for passive oscillatory systems, i.e., for those which are not themselves sources of energy, the imaginary parts of the roots must be positive. If one considers the path of integration in the complex plane \(\omega=Re^{j\theta}\) (Fig. 1), then one can verify, after substituting the integrand for the upper semicircle, that the integral taken along this semicircle tends to zero as \(R\to\infty\). Instead of the desired integral (4), taken along the real axis, one may use the integral taken along the path according to Fig. 1, and then, on the basis of the theory of functions of a complex variable, one obtains

Fig. 1. Path of integration of the function expressed by formula (4)

Fig. 1. Path of integration of the function expressed by formula (4)

\[ x=K\left[\frac{1}{W(0)}+\sum \frac{e^{jp_\nu t}}{p_\nu\left(\dfrac{\partial W}{\partial\omega}\right)_{\omega=p_\nu}}\right]=K\cdot\varphi(t). \tag{5} \]

This formula, first given by Heaviside\(^3\) without proof, was subsequently proved by K. W. Wagner. It makes it possible to compute quite rigorously the behavior of a linear passive system on which a constant force has suddenly begun to act. For this computation it is necessary to solve an algebraic or trans-

cendental equation

$$ W(\omega)=0. $$

\(\varphi(t)\) is called the “transition function.”

In exactly the same way one can also compute the effect of any force varying in time; for this it is necessary to decompose the applied force in time into a series of infinitely small forces constant in magnitude, for which their transition functions (5) are known. If the applied force, which begins to act at the moment of time \(t=0\), is denoted by \(f(t)\), then the behavior of the system may be described, following Carson\(^{11}\), by means of the following expression:

$$ x=\frac{d}{dt}\int_0^t f(t)\varphi(t-\lambda)\,dt. \tag{6} $$

The most important case is that of the action of a force varying sinusoidally, beginning from the moment of time \(t=0\). In this case one must put \(f(t)=K\sin \omega t\), or \(f(t)=K e^{j\omega t}\). In the latter case, from the final result one must discard the imaginary part in order to obtain sinusoidal motion. Applying Laplace’s interpolation formula for the computation, we obtain the following result:

$$ x=K\left[\frac{e^{j\omega t}}{W(\omega)}+\sum_{\nu=1}^{n}\frac{e^{jp_\nu t}}{(p_\nu-\omega)W'(p_\nu)}\right]. \tag{7} $$

The first term on the right-hand side represents undamped oscillations, while the second term expresses the sum of superposed establishing oscillations. Only this second term represents the establishing process, computed under initial conditions corresponding to the switching-on conditions. If all the individual systems can oscillate, then all the roots \(p_\nu\) are complex. The imaginary parts of all the roots, as established above, are positive. Therefore each pair of terms of the sum corresponding to conjugate complex roots represents a damped oscillation. Consequently, the entire establishing process is a collection of separate damped oscillations whose frequencies are equal to the natural frequencies of the system.

Using formula (7), the author\(^{19}\) calculated the establishing process arising in a system with two and three degrees of freedom. The calculation referred to the case of two coupled electrical oscillatory circuits, to one of which a sinusoidal electromotive force is applied from outside, and the value of the current in the second circuit is computed as a function of time. This corresponds to the velocity of the second system in the transition to the mechanical analogy. The value of the quantity \(W\) here is, correspondingly, different. In the case when both systems have one and the same natural frequency (in the absence of coupling) \(\omega_0\), coinciding with the exciting frequency,

and equal damping, then for the values of the velocity of the second system under elastic coupling the following formula is obtained:

\[ v_2=-\frac{P}{\sqrt{r_1 r_2}}\frac{m}{1+m^2} \left[ \cos \omega_0 t -\frac{\sqrt{1+m}}{2m}\,e^{-\omega_0\frac{D}{2}t} \left\{ \frac{\sqrt{1+k^2}}{\nu_1}\cos(\omega_0\nu_1 t-\varphi_1) + \frac{\sqrt{1-k^2}}{\nu_2}\cos(\omega_0\nu_2 t-\varphi_2) \right\} \right]. \tag{8} \]

Here the following notation has been used: \(P\) is the amplitude of the forcing force, \(k\) is the coupling coefficient, \(r_1\) and \(r_2\) are the corresponding friction coefficients of both systems, \(D\) is the damping, identical for both systems,

\[ m=\frac{k}{D};\qquad \nu_1=\sqrt{1+\frac{D^2}{4}+k};\qquad \nu_2=\sqrt{1-\frac{D^2}{4}-k}; \]

\[ \varphi_1=\operatorname{arc\,tg}\frac{2+k}{2\nu_1 m};\qquad \varphi_2=\operatorname{arc\,tg}\frac{2-k}{2\nu_2 m}. \]

In practically interesting cases the coupling coefficient and the damping have values of the same order, \(10^{-2}\). In these cases, neglecting higher powers of \(D\), one may simplify the exact formula (8) and obtain the following approximate formula:

\[ v_2= \]

\[ =-\frac{P}{\sqrt{r_1 r_2}}\frac{m}{1+m^2}\cos\omega_0 t \left[ 1-\frac{\sqrt{1+m^2}}{m}\,e^{-\omega_0\frac{D}{2}t} \cos\left(\omega_0\frac{k}{2}t-\varphi\right) \right], \]

where

\[ \varphi=\operatorname{arc\,tg}\frac{1}{m}. \]

If, for the process of dying out of the oscillations, it is assumed that the forcing force ceases to act at the moment when its value passes through zero, then the process of dying out of the oscillations will be represented by the following formula:

\[ v_2=\frac{P}{\sqrt{r_1 r_2}}\frac{e^{-\omega_0\frac{D}{2}t}}{\sqrt{1+m^2}} \cos\omega_0 t\cdot \cos\left(\omega_0\frac{k}{2}t-\varphi\right). \tag{9} \]

The process of dying out of the oscillations consists in the superposition upon one another of two coupled oscillations with angular frequencies

\[ \omega_{1,2}=\omega_0\left(1\pm\frac{k}{2}\right), \]

which form beats. Since the amplitudes of both coupled oscillations may, with great approximation, be considered equal to one another, then in the beats that arise the amplitude falls to zero. The frequency

beatings is proportional to the coupling coefficient \(k\); consequently the beatings occur the faster, the stronger the coupling. The oscillations in a decaying process occur between two pairs of envelope curves; one pair of curves is expressed by the equation

\[ \pm e^{-\omega_0 \frac{D}{2} t}\cos\left(\omega_0 \frac{k}{2}t-\varphi\right), \]

and the other pair of exponential curves is expressed by the equation

\[ \pm e^{-\omega_0 \frac{D}{2} t} \]

(Fig. 2). Of great practical importance is mainly the question of the duration of the decaying process, which may be defined as the time during which the oscillations decay to \(10\%\) of their initial value. On the basis of Fig. 2 one may formulate the entirely general proposition that the decaying process ceases the sooner, the greater the damping. In practical cases the coupling is usually so weak that the envelope curve of the beatings in its initial part does not differ from an exponential curve. On this basis it may be considered that the duration of the decay decreases with increasing coupling or with the diameter of the opening, which is what is also arrived at in the theory of filters.

Fig. 2. Decay of oscillations of a system with two degrees of freedom

Fig. 2. Decay of oscillations of a system with two degrees of freedom

Formulas (8) and (9), expressing a rigorous calculation of the course of the decaying process, may be applied in order, on the basis of experimentally obtained oscillograms of damped oscillations, to obtain the magnitude of the coupling and of the damping. First of all, by counting the number of oscillations falling within one beat, one can in the general case obtain the value of the coupling coefficient for systems with two degrees of freedom. Since the natural frequencies in the presence of coupling are expressed by the formulas

\[ \omega_1=\omega_0\left(1+\frac{k}{2}\right);\qquad \omega_2=\omega_0\left(1-\frac{k}{2}\right), \]

then, obviously, the beat frequency is equal to \(\omega_0 k\). If by \(n\) we denote the number of oscillations falling within one period of beats, then \(k=\frac{1}{n}\). Further, in the case of such clearly expressed beats, which are obtained with elastic coupling under imposed conditions, it is possible to obtain the value of the quantity \(m=\frac{k}{D}\) from the ratio of the first max-

the ratio of the beat amplitude \(v_{m_1}\) to the amplitude of the undamped oscillations \(v_s\). On the basis of (8) one obtains \(v_{m_1}=v_s e^{-\frac{\pi}{m}}\), whence

\[ m=\frac{\pi}{\ln \frac{v_s}{v_{m_1}}}. \]

These formulas were applied by the author of the present article\({}^{94}\) in order, on the basis of oscillograms of the damping of the natural vibrations of a violin in the region of its principal resonance, to compute both the damping of the separate parts of this system, consisting of the string and the body, and the coefficient of coupling between them. In these experiments the conditions imposed in deriving formula (9) were fulfilled. The oscillations of the sound pressure in the sound field, recorded on the oscillograms, are determined chiefly by the velocities of the directly excited system, namely the body of the violin. With the appropriate pressure of the string, it is tuned to the principal resonance of the instrument.

The coupling effected through the bridge is mainly elastic. It must be concluded that both systems, the string and the body, have approximately the same natural damping. The string, which in itself has considerably less damping than the body, receives additional damping as a consequence of the deformation of the bridge.

3. Symbolic methods. Important auxiliary means for the calculation of transient processes, especially in complicated cases when the usual computational procedures become very unclear, are provided by the symbolic methods proposed by Heaviside\({}^{3}\) and developed and rigorously substantiated by K. W. Wagner\({}^{7}\), Bromwich\({}^{9}\), and Carson\({}^{11}\). The main application of these methods was formerly confined to electrical engineering. At the present time Prager\({}^{40}\) has applied these methods to the solution of mechanical problems. The basic procedures of investigation by means of these methods must be set forth here.

Let the motion of a system be described by the linear differential equation

\[ a_n \frac{d^n y}{dt^n}+a_{n-1}\frac{d^{n-1}y}{dt^{n-1}}+\cdots+a_1\frac{dy}{dt}+a_0 y=z(t). \tag{10} \]

Instead of the unknown function \(f(t)\), a new function \(F(p)\) of the complex variable \(p=c+j\omega\) is introduced, by means of the following relation:

\[ f(t)=\frac{1}{2\pi j}\int_{c-j\infty}^{c+j\infty} k(t,p)\,F(p)\,dp. \tag{11} \]

Here the kernel of the transformation has the following form:

\[ k(t,p)=\frac{e^{pt}}{p}. \tag{12} \]

The path of integration lies in the plane of the complex variable \(p\), parallel to the imaginary axis. The real positive quantity \(c\) is chosen so that all singular points of the integrand lie to the left of the path of integration. If differentiation under the integral sign is considered permissible, then the following formulas are obtained for the derivatives of the function \(f(t)\):

\[ \frac{d^n}{dt^n}[f(t)] = \frac{1}{2\pi j} \int_{c-j\infty}^{c+j\infty} k(t,p)\,p^n F(p)\,dp. \tag{13} \]

We assume that both the perturbing function \(z(t)\) and the desired function can be expressed in the form (11):

\[ \left\{ \begin{aligned} z(t)&=\frac{1}{2\pi j} \int_{c-j\infty}^{c+j\infty} k(t,p)\,Z(p)\,dp, \tag{14} \\[6pt] y(t)&=\frac{1}{2\pi j} \int_{c-j\infty}^{c+j\infty} k(t,p)\,Y(p)\,dp. \tag{15} \end{aligned} \right. \]

Substituting the functions from (14) and (15) into (10) and taking (13) into account, we obtain

\[ \frac{1}{2\pi j} \int_{c-j\infty}^{c+j\infty} dp\,k(t,p) \bigl[ (a_n p^n+a_{n-1}p^{n-1}+\cdots \]

\[ \cdots+a_1p+a_0)Y(p)-Z(p) \bigr]=0. \]

It follows that

\[ Y(p)= \frac{Z(p)} {a_n p^n+a_{n-1}p^{n-1}+\cdots+a_1p+a_0}. \]

Consequently, the solution of the differential equation is expressed by relation (15). If (12) is substituted into (11) and the complex integration is carried out, one can verify that the function \(f(t)\) vanishes for negative values of the time \(t\). It is precisely in this that the special advantages lie of choosing the function \(F(p)\) in the form (12) when computing the effect of suddenly applied disturbing forces on a system initially at rest.

The symbolic method of calculation proposed by Heaviside consists in leaving in the right-hand side of formula (11) only

\(F(p)\). The formula is then rewritten as:

\[ F(p)\div\to \frac{1}{2\pi j}\int_{c-j\infty}^{c+j\infty}\frac{e^{pt}}{p}F(p)\,dp=f(t). \]

\(F(p)\) is an abbreviated symbolic notation for the function \(f(t)\). From (13) it follows that

\[ \frac{d^n}{dt^n}[f(t)] \leftarrow\to p^n F(p). \tag{16} \]

Calculations are performed with operators as with numbers, and then, by means of formulas (14) and (15), after carrying out the integration one passes to the actual functions. It makes sense to find the values of the simplest functions of the operator \(p\).

For \(F(p)=p^{-n}\), the integrand takes the following form:

\[ e^{pt}\cdot p^{-n-1}=p^{-n-1}+tp^n+\ldots+\frac{t^n}{n!}p^{-1}+\ldots \]

It is then easy to see that \(f(p)\) vanishes for \(t<0\); for \(t>0\), at \(p=0\) there is an \((n-1)\)-fold pole, which gives \(\dfrac{t^n}{n!}\).

Consequently,

\[ p^{-n}\div\to \begin{cases} 0 & \text{for } t<0,\\[4pt] \dfrac{t^n}{n!} & \text{for } t>0. \end{cases} \]

In the particular case

\[ p^0\div\to \frac{1}{2\pi j}\int_{c-j\infty}^{c+j\infty}\frac{e^{pt}}{p}\,dp=1. \]

This holds for \(t>0\). On the basis of this example, in analogous cases one may assign to the function the value \(0\) for \(t<0\).

Similarly one may find, if \(\alpha\) is complex,

\[ \left\{ \begin{aligned} \frac{p}{p-\alpha} &\div\to e^{\alpha t}\\[6pt] \frac{p}{(p-\alpha)^{n+1}} &\div\to \frac{t^n}{n!}e^{\alpha t}\\[6pt] \frac{p^2}{p^2+\alpha^2} &\div\to \cos \alpha t\\[6pt] \frac{\alpha p}{p^2+\alpha^2} &\div\to \sin \alpha t\\[6pt] \frac{p^2}{p^2-\alpha^2} &\div\to \cos \operatorname{hyp}\alpha t\\[6pt] \frac{\alpha p}{p^2-\alpha^2} &\div\to \sin \operatorname{hyp}\alpha t \end{aligned} \right. \tag{17} \]

Next, from (11) and (12) it follows that

\[ F\left(\frac{p}{a}\right)\div\!\!\longrightarrow f(at). \tag{18} \]

Consequently,

\[ e^{-ap}\cdot F(p)\div\!\!\longrightarrow f(t-a). \tag{19} \]

The operator \(e^{-ap}\) expresses a function which for \(t<a\) is identically equal to zero, and for \(t>a\) assumes the value equal to 1. This function analytically expresses a suddenly applied impulse; graphically the function is shown in Fig. 3a. If two such functions are added algebraically, then a function is obtained which is represented graphically in Fig. 3b and analytically expressed by means of the operator

\[ A\left(e^{-ap}-e^{-(a+h)p}\right). \]

Integration of this function over the limits from \(-\infty\) to \(+\infty\) gives \(Ah\). If, as \(h\) approaches zero, \(A\) correspondingly increases so that the product \(Ah=1\), then in the limit a special function is obtained which for \(t\gtrless 0\) is equal to zero, and for \(t=0\) assumes an infinite value, so that its integral from \(-\infty\) to \(+\infty\) has the value equal to 1. Analytically this function may be represented by means of the operator:

\[ \lim_{h=0}\frac{e^{-ap}-e^{-(a+h)p}}{h}=p\cdot e^{-ap}. \]

Fig. 3a and 3b. Graphical representation of the function \(S_1(t)\)

Fig. 4. Graphical representation of the function \(S_2(t)\)

According to (19), for the systems \(S_1(t)\) one may compose the following symbolic equalities:

\[ \left\{ \begin{aligned} p&\div\!\!\longrightarrow S_1(t)\\ p\cdot e^{-ap}&\div\!\!\longrightarrow S_1(t-a) \end{aligned} \right\} \tag{20} \]

It is easy to see that

\[ \int_{-\infty}^{+\infty} S_1(t)\,dt=1. \]

The function \(S_1(t)\) may be regarded as a suddenly applied impulse of force which imparts to a resting mass, at the time \(t=0\), the velocity

\[ v=\frac{1}{m}. \]

With the aid of entirely analogous reasoning, applied to the limiting case as \(h\) tends to 0, one can determine the function \(S_2(t)\), graphically represented in Fig. 4 and expressing two equal in magnitude and oppositely directed impulses, following one another infinitely rapidly. Such two impulses suddenly impart to a mass at rest, at the instant \(t=0\), a deviation from the initial position \(s=\dfrac{1}{m}\). On the basis of the definition of the function set forth, we obtain the following relations:

\[ \int_{-\infty}^{+\infty} tS_2(t)\,dt=0 \quad \text{and} \quad S_2(t) \leftrightarrow p^2 . \tag{21} \]

With the aid of these representations one may additionally determine one more function

\[ p^0 \leftrightarrow S_0(t). \]

This symbolic method makes it possible to solve complicated problems very elegantly. Let us show its application in two simple examples. A certain mass \(m\) is under the action of an elastic force characterized by the coefficient of elasticity \(c\). Suppose that at the initial instant of time \(t=0\) the mass is displaced relative to the initial position by an amount \(s_0\) and has velocity \(v_0\). At this initial instant of time a constant force \(P\) is suddenly applied to it. With the aid of the notation introduced above, the following differential equation of motion may be written:

\[ m\frac{d^2s}{dt^2}+cs=PS_0(t)+v_0mS_1(t)+s_0mS_2(t). \]

On the basis of (16), (20), and (21), we obtain

\[ (p^2+\omega_0^2)S=\frac{P}{m}+v_0p+s_0p^2, \]

where \(\dfrac{c}{m}=\omega_0^2\). Solving the last equation with respect to \(S\), we obtain

\[ S(p)=\frac{P}{m\omega_0^2} \left(1-\frac{p^2}{p^2+\omega_0^2}\right) +\frac{v_0p}{p^2+\omega_0^2} +\frac{s_0p^2}{p^2+\omega_0^2}. \]

Replacing the operators on the basis of (17), we obtain

\[ s(t)=\frac{P}{m\omega_0^2}(1-\cos \omega_0 t) +\frac{v_0}{\omega_0}\sin \omega_0 t +s_0\cos \omega_0 t. \]

Another example concerns the case of an oscillating system with continuously distributed constants, namely a homogeneous string, to the middle of which at the initial instant of time \(t=0\) a constant force \(P\) is suddenly applied. If the tension of the string is denoted by \(S\), and by \(\mu\) the mass per unit length, then the differen-

cial equation of motion of the string can be written as follows:

\[ S \frac{\partial^2 y}{\partial x^2} - \mu \frac{\partial^2 y}{\partial t^2}=0. \]

Let us denote \(\frac{S}{\mu}=a^2\). On the basis of (18) one can obtain the following differential equation for the function \(Y(x,p)\):

\[ \frac{d^2Y}{dx^2}-\frac{p^2}{a^2}Y=0. \]

The solution of this latter equation, which becomes zero for \(x=0\), is

\[ Y=A\sin \operatorname{hyp}\frac{p}{a}x. \]

The constant of integration \(A\) may be determined from the equilibrium condition for the middle of the string\(^1\):

\[ 2S\left(\frac{dy}{dx}\right)_{x=\frac{l}{2}}=-P. \]

Hence we obtain

\[ A=-\frac{P}{2S\frac{p}{a}\cos \operatorname{hyp}\frac{pl}{2a}}. \]

The symbolic solution has the following form:

\[ y\left(\frac{l}{2},t\right)\leftarrow\div \frac{Pl}{4S} \frac{\sin \operatorname{hyp}\frac{pl}{2a}} {\frac{pl}{2a}\cos \operatorname{hyp}\frac{pl}{2a}}. \]

The following relation holds:

\[ \frac{\sin \operatorname{hyp} p}{p\cos \operatorname{hyp} p} = \frac{8}{\pi^2}\left(1-\frac{p^2}{p^2+\frac{\pi^2}{4}}\right) + \frac{8}{9\pi^2}\left(1-\frac{p^2}{p^2+\frac{9\pi^2}{4}}\right) +\cdots . \]

From this relation, by means of some transformations and on the basis of (17) and (18), we obtain the required solution:

\[ y\left(\frac{l}{2},t\right)= \frac{Pl}{4S} \left[ 1-\sum_{n=0}^{\infty} \frac{8\cos(2n+1)\pi \frac{at}{l}} {(2n+1)^2\pi^2} \right]. \]

Another form of the solution, which in some respects proves to be more intuitive, can be obtained on the basis of the fol-

giving the expansion into a series:

\[ \frac{\sin \operatorname{hyp} p}{p \cos \operatorname{hyp} p} = \frac{1-e^{-2p}}{p(1+e^{-2p})} = \frac{1}{p}\left(1-2e^{-2p}+2e^{-4p}-\ldots\right). \]

This series converges uniformly, since along the path of integration for (II) \(\left|e^{-2p}\right|<1\), on the basis of the fact that \(\operatorname{Re}(2p)>0\).

This expansion makes it possible to integrate each term of the series and to investigate the solution in greater detail:

\[ \left\{ \begin{aligned} \frac{4S}{Pl}\cdot y\left(\frac{l}{2},t\right) &=\frac{2at}{l} &&\text{for } 0<\frac{2at}{l}<2,\\ &=4-\frac{2at}{l} &&\text{for } 2<\frac{2at}{l}<4,\\ &=\frac{2at}{l}-4 &&\text{for } 4<\frac{2at}{l}<6 \end{aligned} \right. \]

and so on.

Other methods of symbolic computation, in particular those proposed by Schmidt \(^{24}\), are set forth in Prager’s article \(^{40}\), which also gives the literature on this question.

4. Approximate method of calculating nonstationary processes

In the rigorous calculation of establishing processes it is necessary to solve algebraic equations. Even for a system with two degrees of freedom, equations of the fourth order are obtained. It is possible to present the results of the calculation in a visual form only in certain very simple cases. With systems having a larger number of degrees of freedom this is much more difficult to do. An exact solution in these cases can be obtained only by imposing such conditions as are practically not fulfilled. However, even in these cases, by considering the resonance properties of systems, it is possible to relate them to the duration of the establishing process. This possibility was first pointed out by Caughey-Müller \(^{16}\).

Caughey-Müller also first pointed out the advantages obtained by considering Fourier integrals under various modes of switching on. The sudden switching-on of a constant force \(K\) at the initial moment of time \(t=0\) may, on the basis of (2), be represented in the form of the following formula:

\[ k = K\left[ \frac{1}{2} + \frac{1}{\pi}\int_{0}^{\infty}\frac{\sin \omega t}{\omega}\,d\omega \right] = \frac{K}{2} + \int_{0}^{\infty} A(\omega)\sin \omega t\,d\omega. \]

This increasing process may be regarded as caused by the action of a constant force \(\frac{K}{2}\) and of an infinite series of variable forces \(\frac{K}{\pi\omega}\sin \omega t\), lying in the frequency interval \(d\omega\) and treated as strictly stationary, i.e. as applied during the entire

of time from \(-\infty\) to \(+\infty\). The multiplier \(A(\omega)=\dfrac{K}{\pi\omega}\) expresses the weight which the corresponding oscillation has in the formation of the rising process. Consequently, by means of the function \(A(\omega)\) the spectrum of the rising process can be represented. In the case under consideration the amplitudes of the individual oscillations are inversely proportional to the corresponding frequencies. Graphically this amplitude spectrum is represented in the form of a hyperbola.

If the constant force is switched on not instantaneously, but varies according to an exponential law with the time constant \(T\), according to the formula

\[ k=K\left(1-e^{-\frac{t}{T}}\right), \]

then for the amplitude spectrum Bjórk, Kotowski, and Lichte \({}^{67}\) obtained the following expression:

\[ A(\omega)=\frac{K}{\pi\omega}\frac{1}{\sqrt{1+\omega^2T^2}} . \tag{22} \]

The falling-off in the region of high frequencies occurs in this case considerably faster than with instantaneous switching-on of a constant force, and the difference is the greater, the larger the time constant \(T\). Acoustically this means that, with such a “smoothed” switching-on, the switching-on process is accompanied by less noise, which is due chiefly to the high overtones.

A short impulse, switched on instantaneously and, after reaching its maximum value, decaying exponentially with the time constant \(T\), has the following amplitude spectrum (67):

\[ A(\omega)=\frac{T}{\pi\sqrt{1+\omega^2T^2}} . \tag{22a} \]

In conclusion, it is of interest to cite the behavior of an electrical oscillatory circuit with parameters \(R, L, C\), in the limiting case of an aperiodic regime, to which at the instant \(t=0\) a constant electromotive force \(E\) is instantaneously applied. As is known, in this case

\[ i=\frac{2E}{R}\cdot\frac{t}{T}e^{-\frac{t}{T}}, \]

where \(T=\dfrac{2L}{R}\). If one puts \(\dfrac{2E}{R}=1\), then for the amplitude spectrum of the current impulse one obtains

\[ A(\omega)=\frac{2}{T}\frac{1}{\frac{1}{T^2}+\omega^2}. \tag{22b} \]

If the force \(K\) is switched on at the instant \(t_1\), and switched off at the instant \(t_2=t_1+T\), then one must add two processes expressed by formula (2),

but with opposite signs. The constant forces \(\frac{K}{2}\) cancel, and the following amplitude spectrum is obtained:

\[ A(\omega)=\frac{2K}{\pi\omega}\sin\frac{\omega T}{2}. \tag{23} \]

In this case, too, the amplitudes of the individual overtones decrease as their frequency increases; however, they do not do so monotonically, but oscillatory. For all frequencies whose periods contain an integer number of \(T\), \(A\omega=0\).

Of particular interest is the case of the instantaneous switching-on of a sinusoidal force with angular frequency \(\Omega\). In this case, in formula (2), instead of \(K\) one must substitute \(K\sin\Omega t\). For the amplitude spectrum the following expression is obtained:

\[ A(\omega)=\frac{K\Omega}{\pi(\Omega^2-\omega^2)}. \tag{24} \]

For \(\omega=\Omega\) the function \(A(\omega)\) becomes infinite. The curve \(A(\omega)\) has a form analogous to the resonance curve of a system without damping with natural frequency \(\Omega\).

If the sinusoidal force \(K(\lambda)=a\sin\Omega\lambda+b\cos\Omega\lambda\), with \(\sqrt{a^2+b^2}=1\), is switched on at the instant \(t=-\frac{T}{2}\) and continues to act until the instant \(t=+\frac{T}{2}\), then the amplitude spectrum according to Bjork, Kotowski, and Lichte \(^{68}\) has the following form:

\[ \left\{ \begin{aligned} A(\omega)&=\frac{1}{\pi}\sqrt{F^2+G^2+2FG(1-2a^2)},\\ \text{with}\qquad F&=\frac{\sin(\Omega-\omega)\frac{T}{2}}{\Omega-\omega},\qquad G=\frac{\sin(\Omega+\omega)\frac{T}{2}}{\Omega+\omega}. \end{aligned} \right. \tag{25} \]

If we restrict ourselves to frequencies that are not too low and to switching times that are not too short, then all terms involving \(G\) may be neglected as small in comparison with \(F^2\), and the following approximate expression is obtained:

\[ A(\omega)=\frac{1}{\pi}\frac{\sin(\Omega-\omega)\frac{T}{2}}{\Omega-\omega}. \]

One can obtain another form for the amplitude spectrum if one takes into account that the phase at switching-on, as has been established experimentally, has only a secondary influence on the course of the transient process. On this basis one may put \(a=1\) and \(b=0\). If, further, one assumes that the duration of switching-on \(T\) constitutes an integer number of periods \(\frac{2\pi}{\Omega}\), without limiting

by this condition the generality of the treatment is substantially limited, then one obtains

\[ A(\omega)=\frac{2\Omega}{\pi(\Omega^{2}-\omega^{2})}\sin\omega\frac{T}{2}. \tag{26} \]

From this it is easy to visualize the form of the curve \(A(\omega)\). The envelope curve has the same form as in formula (24). The values (26) oscillate between the envelope curve (24) and the axis of angular frequencies \(\omega\), in such a way that for all frequencies \(\omega=\dfrac{2\pi n}{T}\), where \(n\) is an integer, the amplitude of the corresponding overtone becomes zero. The oscillations occur the more rapidly, the greater the duration \(T\) of the switching-on. In this case the amplitude at \(\omega=\Omega\) no longer becomes infinite, but assumes the value

\[ A(\omega)_{\omega=\Omega}=\frac{T}{2\pi}, \tag{27} \]

i.e., it increases with an increase in the duration of switching-on.

In calculating transient processes arising from causes similar to those considered above, one proceeds as follows: for each overtone of the amplitude spectrum the action is calculated according to the rules derived for undamped oscillations, and then these separate actions are summed. A separate action can be related to the corresponding cause producing it by means of the relation \(s=B\cdot A(\omega)\), where the transmission coefficient \(B\) in the general case must be taken in complex form \(B=B(\omega)e^{i\beta(\omega)}\), in order to represent both amplitude distortions and phase distortions. \(B\) and \(\beta\) are functions of frequency. The integral result can be obtained by summing the separate actions in the form of a Fourier integral, expressed in the present case, just like the function \(B\), in complex form:

\[ S=\int_{-\infty}^{+\infty}s\,d\omega =\int_{-\infty}^{+\infty} A(\omega)\,B(\omega)\,e^{i(\omega t+\beta)}\,d\omega . \]

In the simplest cases it is possible to compute the exact value of this integral, as was done by Küpfmüller \(^{16}\) for a system with one degree of freedom. The advantage of such a method of treatment consists in the fact that the transient processes, by introducing a transmission coefficient by means of which the frequency characteristic of the system under consideration is represented, are connected with its resonance properties. In more complicated cases one proceeds by choosing for the transmission coefficient the simplest forms that approximate the actual behavior of the system, and carrying out the integration.

The simplest case is that in which \(B\) is constant up to some limiting frequency \(\omega_{0}\), and beyond this boundary is assumed to become vanishingly small. This case is realized fairly closely

in high-quality microphones. As for the angle of the transmission coefficient, it may, to a good approximation, be taken as proportional to the frequency, i.e. \(\beta=\omega t_0\). If at the moment \(t=0\) a constant force \(K\) is applied, then on the basis of (2) we obtain

\[ x=BK\left[\frac{1}{2}+\frac{1}{\pi}\int_0^{\omega_0}\frac{\sin\omega(t-t_0)}{\omega}\,d\omega\right]. \]

Integration gives a result expressed with the aid of the sine integral function (Fig. 5)

\[ x=BK\left[\frac{1}{2}+\frac{1}{\pi}\operatorname{Si}\omega_0(t-t_0)\right]. \]

This formula does not give an exact solution, since at the initial instant of time \(t=0\) the displacement \(x\) is not equal to zero, as it should be according to the conditions of the problem, but becomes zero only at \(t=-\infty\). The reason for this discrepancy lies in the fact that the assumed form for the transmission factor is too idealized and, strictly speaking, can be applied only to systems with an infinitely large number of degrees of freedom and with an infinitely long build-up time. The rise time \(\tau\) may be obtained as the projection of the tangent at the point \(t=t_0\) onto the abscissa axis. On this basis we obtain \(\tau=\dfrac{\pi}{\omega_0}\), or \(\tau\omega_0=\pi\). This basic relation, first obtained by Küpfmüller \(^{16}\), establishes that the rise time for such idealized systems is inversely proportional to the limiting frequency. In order that this result agree with the practically obtained resonance curves, it is necessary to choose the limiting frequency \(\omega_0\) so that the following relation holds:

\[ \omega_0=\frac{1}{B(0)}\int_0^{\infty} B(\omega)\,d(\omega). \]

Fig. 5. Switching-on of a low-frequency filter (after Küpfmüller)

Fig. 5. Switching-on of a low-frequency filter (after Küpfmüller)

Fig. 6. Ideal filter (after Küpfmüller)

Fig. 6. Ideal filter (after Küpfmüller)

In practice this means the frequency at which the curve \(B(\omega)\) has the greatest curvature. Of particular importance are the establi—

developing processes in selective systems, which respond especially well only to a definite band of frequencies and respond very little to frequencies lying outside this band. The simplest example is a vibratory system with one degree of freedom, already considered in II, 1. A rigorous solution of the problem for a system with two degrees of freedom under certain imposed conditions has likewise already been given (II, 2). In more complicated cases one may regard the resonance curve as having the form of a rectangle, and take the transmission factor in this range to increase linearly with frequency (Fig. 6).

If at the instant of time \(t=0\) a sinusoidal force with frequency \(\Omega=\dfrac{\omega_1+\omega_2}{2}\), i.e. equal to the mean frequency of the transmitted band, begins to act instantaneously, then a calculation similar to that set forth above leads, for the rise time, to the same Kopfmüller rule

\[ \tau=\frac{2\pi}{\omega_1-\omega_2}=\frac{2\pi}{\Delta\omega};\quad \text{or}\quad \tau\Delta\omega=2\pi. \]

Consequently, the rise time is inversely proportional to the transmitted frequency width. In practical cases, in which the resonance curve does not have the form of an exact rectangle, the limiting frequencies are taken to be those at which the resonance curve falls from its maximum value \(B_{\max}\) to

\[ \frac{B_{\max}}{\sqrt{e}}=0.607\,B_{\max}. \]

With the aid of amplitude spectra, by the method of Bjørk, Kotowski, and Lichte \(^{83}\), one can obtain a number of qualitative conclusions concerning the nature and duration of developing processes. The curve of the amplitude spectrum obtained under the action of a sinusoidal force during a time \(T\) has, according to (27), the maximum \(A(\omega)_{\max}=\dfrac{T}{2\pi}\), corresponding to the frequency of the applied force. The longer the tone lasts, the more strongly the frequency of the sinusoidal force stands out in comparison with other frequencies, which produce noises during switching on and switching off. If the force is not switched on immediately at its maximum value, but gradually with a certain time constant, then in this case too, just as when a constant force is switched on, according to (22) the extraneous frequencies are expressed, in comparison with the frequency of the forcing force, the more weakly the longer the rise time is, i.e. the softer the switching on is made.

If the frequency of the sinusoidal force \(\Omega\), applied to the vibratory system, is chosen so that it corresponds to the maximum of the resonance curve, then this frequency is distinguished still more strongly in the amplitude spectrum in comparison with extraneous frequencies. The resulting oscillation begins to grow softly. The time constant of the rise of oscillations in a system with one degree of freedom

is determined by the formula \(T=\dfrac{1}{fd}\), where \(f\) denotes the frequency and \(d\) the logarithmic decrement. If the switching-on is performed on the steep side of the resonance curve, then the frequency of the driving force is less sharply expressed, while the frequencies situated nearby are amplified owing to resonance. As a result the switching-on process gives an amplitude spectrum similar to that obtained with the instantaneous switching-on of a sinusoidal force; the oscillations arise very rapidly, almost at once, and the resulting time constant is practically equal to zero. Finally, if the frequency of the driving force differs considerably from the resonance frequency, then in the amplitude spectrum the region corresponding to resonance stands out, while the frequency of the driving force is expressed very weakly. In this case the noises at switching-on bring out the frequencies of the driving force better. In this case the process corresponds to a negative time constant. The cases discussed are explained in Fig. 7.

Fig. 7. Oscillograms and simplified diagrams of certain typical processes occurring during switching-on and switching-off (after Björk, Kotowski, and Lichte)

Fig. 7. Oscillograms and simplified diagrams of certain typical processes occurring during switching-on and switching-off (after Björk, Kotowski, and Lichte)

III. Auditory Perception of Nonstationary Sound Processes

5. Model of the Ear and Transient Processes.

According to Helmholtz’s conceptions\(^6\), the basilar membrane of the human ear consists of a large number of fibers tuned to different tones within the range from the lower to the upper limit of hearing. When sound acts upon it, individual fibers whose natural frequencies coincide with the frequencies of the acting sound begin to vibrate and excite the endings of the auditory nerves. This theory explains well the resolving power of the ear, i.e., the ability to distinguish tones of different frequency. However, difficulties immediately arise when one attempts to explain the perception of nonstationary sounds. Helmholtz determined the damping of the membrane fibers by observing how rapidly a trill ceases to be distinguished by the ear as separate tones passing into a merged sound, and from this he de...

cluded that the logarithmic decrement of damping for all the resonator elements of the ear is one and the same and is equal to 0.23. Weitmann ^4^ measured the speed of trills at which the fusion of two different tones is observed at different frequencies, and came to the conclusion that the frequency of trills required for this is the same at all pitches. From this fact it follows that all resonators damp in one and the same time to the same fraction of their initial amplitude, i.e. that the logarithmic decrement decreases with increasing frequency.

The damping of the resonator elements of the ear can also be determined by another method, namely from the recorded resonance curve. Determining the sharpness of the resonance curve gives directly the value of the logarithmic decrement. Obviously, such measurements can be carried out only in the case where it is possible to exclude the vibrations of neighboring fibers and to observe the result due to the vibrations of an isolated fiber. Such conditions are observed in certain diseases and injuries of the auditory organ, when some definite frequency regions are eliminated from the entire range of auditory perception, owing to the destruction of the corresponding fibers. In these cases it is possible to record at least half of the resonance curve. The resonance curve proves to be strikingly sharp. It corresponds to a logarithmic decrement \(d = 0.0006\). Such small damping should condition a very prolonged process of growth of the vibrations; however, this is never observed. This discrepancy leads, in Békésy’s ^22^ opinion, to the conclusion that Helmholtz’s ideas about the role of the basilar membrane in the process of auditory perception are not confirmed experimentally. Békésy showed, on the model of the cochlea that he constructed, that under the action of sounds of medium frequencies the basilar membrane vibrates uniformly along its entire length, beginning from the stapes, and that only in the region of the helicotrema does it remain at rest. Békésy’s theory, substantiated by him in experiments with a model of the cochlea, is a further development and deepening of the resonance theory of hearing proposed by Helmholtz. The selectivity of auditory perception, which absolutely cannot be explained under the assumptions made about the character of the vibrations of the basilar membrane, must be ascribed to other causes that were experimentally established in experiments with a model of the cochlea. According to Békésy’s theory, when sound acts, vortices arise on both sides of the basilar membrane; their position is determined by the frequency of the sound and shifts in the direction of the helicotrema as the frequency decreases.

To elucidate the mechanism of action of nonstationary sounds on the ear, it is necessary to determine its damping and its natural frequency. Both of these characteristics can be determined from the ear’s own obtained damped vibrations. Békésy ^37,66^ succeeded in carrying out these determinations by listening with the aid of a rubber tube (by another observer) or by recording with the aid of-

by means of a contact microphone of short sound impulses (clicks), which arise in the ear during swallowing movements. From the curves obtained in this way, the values of the logarithmic decrement were determined; they proved to be of the order \(d = 1.1—1.8\), and the natural frequency of the ear of the order of \(1200—1500\) Hz. Hence it follows that the oscillatory processes actually occurring in the middle ear can be correctly described on the assumption that they are in the limiting case of an aperiodic regime. Békésy was able to elucidate on a model (with the aid of stroboscopy) the mechanism of action on the ear of aperiodic impulses. In this case a piston-like motion is observed in the region close to the stirrup; the basilar membrane is instantaneously deflected in the part adjacent to the stirrup, and remains at rest in the part adjacent to the helicotrema (Fig. 8, solid curve). Then the deflected part of the basilar membrane aperiodically returns to the position of rest, while along the membrane in the direction of the helicotrema there propagates a wave which, after 0.05 sec., is leveled out.

Fig. 8. Character of the motion of the basilar membrane (according to Békésy)

Fig. 8. Character of the motion of the basilar membrane (according to Békésy)

If processes actually occur in the ear which are observed on models, then one should expect that the “center of gravity” of the impulse in time will occur, upon excitation of the ear, the earlier the faster the wave running toward the helicotrema dies out. Since these regions of the basilar membrane serve for the perception of tones of low pitch, then by superposing on the impulse a sufficiently strong constant low tone it is possible completely to mask the impulse in the distant parts of the basilar membrane. If simultaneous sound impulses are delivered to both ears by means of two telephones, and in one ear the sound impulse is masked with the aid of a low tone, then in that ear the “center of gravity” of the impulse in time will occur earlier than in the other, where the wave propagates along an undisturbed basilar membrane. These processes will cause a subjective displacement of the “sound image.” On the basis of the binaural effect, simultaneous and identical sounds produce a “sound image” lying in the median plane of the head. When a masking low tone is added to one sound, one should expect a displacement of the “sound image.” Békésy\(^{36,66}\) succeeded in observing this displacement experimentally and thereby confirmed the conclusions obtained by him on the ear model.

  1. Perception of pitch. The question of how long a tone must sound for the ear to be able to determine its pitch clearly was subjected to investigation comparatively long ago. As early as 1873, Mach\(^{1}\) established that the recognition time for a tone of 128 Hz is 35 msec. For tones of 500, 1000, and 2000 Hz, Lübcke\(^{14}\) determined the minimum sounding times, respectively, as 10, 7, 8, 6, and 7.3 msec. Investigations by other authors on this

on the question were carried out with an insufficiently perfected methodology and led to incorrect results. A very detailed investigation of this question was performed by Bürk, Kotowski, and Lichte \(^{68}\) with the aid of a method previously developed by Laimbach \(^{5}\). All necessary switchings of contacts during the measurements were carried out automatically with the aid of a Helmholtz pendulum. The values obtained by them for the minimum time necessary for a distinct perception of pitch, on the basis of experiments with many observers, performed at different frequencies, are presented graphically in Fig. 9, where curves \(a\) and \(b\) bound the regions of definite values.

Fig. 9. Time of recognition of the pitch of a tone, represented as a function of frequency: \(a\)—upper limit, established experimentally; \(b\)—lower limit, established experimentally; \(c\)—computed curve (after Bürk, Kotowski, and Lichte)

With a very short sound impulse the ear hears a sound without a definite tonal character. As the duration of sounding is increased, a tonal character begins to be perceived, which, however, for tones both very high and very low does not correspond to the true frequency of the sounding impulse. In the region of middle frequencies, where the smallest values are observed for the minimum times necessary for establishing the frequency, the auditory sensations of pitch from the very beginning acquire the correct tonal character. The pitch perceived by the ear for short sound impulses seems higher in the region of low frequencies and lower in the region of high frequencies. To explain this result the authors \(^{68}\) proposed the following theory. If the ear is regarded as a system with one degree of freedom, possessing a natural frequency \(\Delta\) (in the absence of damping) and operating at the boundary of the aperiodic regime, and if it is assumed that the sound impulse perceived by the ear is resolved into an amplitude spectrum represented by formula (25), then the following expression is obtained for the magnitude of the energy acting on the ear and expressed as a function of frequency:

\[ i_\omega^{2} = c\, \frac{4\Delta^{2}\omega^{2}}{\left(\omega^{2}+\Delta^{2}\right)^{2}}\, \frac{\sin^{2}(\Omega-\omega)\,\dfrac{T}{2}}{(\Omega-\omega)^{2}} . \tag{28} \]

Assuming that the maximum of this expression occurs at the frequency \(\Omega + x\), we obtain, by means of an approximate calculation,

\[ \frac{x}{\Omega} = \frac{12}{\Omega^{2}T^{2}}\, \frac{\Delta^{2}-\Omega^{2}}{\Delta^{2}+\Omega^{2}} . \]

It follows from this that the longer the duration \(T\) of the pulse, the smaller the relative displacement of the maximum of the sound power. Further, for \(\Omega < \Delta\) this displacement is positive, i.e., the maximum shifts toward higher frequencies, and conversely, for \(\Omega > \Delta\) the relative displacement is negative, i.e., the maximum shifts toward lower frequencies. If it is assumed that the perception of pitch is determined mainly by the frequency at which the maximum of the sound power occurs, then the theoretical considerations presented satisfactorily explain the experimental facts.

If two tones follow one another, separated by some interval of time, then the smallest interval at which the ear is still able to ascertain the separate succession of the tones, as distinct from their fused perception, is, according to the investigations of Björk, Kotowski, and Lichte \(^{83}\), equal to the above-defined minimum time necessary for clear perception of pitch.

In their other work, Björk, Kotowski, and Lichte \(^{79}\) investigated the question of the degree of concentration of power in the sound spectrum necessary for clear perceptions of pitch. The total energy \(E_g\), distributed over the entire sound spectrum, can be calculated by integrating expression (28) with respect to frequency over the limits from \(0\) to \(\infty\). This total energy may be regarded as consisting of three parts: \(E_g = E_e + E_a + E_0\), each of which can be calculated by means of the corresponding integration. Here \(E_e\) denotes the energy falling on the frequencies in the range surrounding the perceived frequency, i.e., bounded by the frequencies \((1-p)\Omega\) and \((1+p)\Omega\), respectively. \(E_a\) denotes the energy falling on the lower frequencies of the range, starting from \((1-p)\Omega\), and \(E_0\) denotes the energy falling on the higher frequencies, starting from \((1+p)\Omega\). Consequently, \(E_e\) denotes the proper tonal energy necessary for the perception of a definite pitch, and \(E_a + E_0\) the extraneous energy. For \(p\) the value \(\pm 0.05\) is taken, and the proper frequency of the ear \(\Delta\) is taken equal to \(2\pi \cdot 1300\). The ratio

\[ \frac{E_e}{E_g} = \mu \]

expresses the relative amount of energy falling on the frequency range lying within \(\pm 5\%\) of the frequency being determined. The numerical value of the quantity \(\mu\) is determined for selected \(T\) and different frequencies \(\Omega\) so as to approach as closely as possible the experimentally obtained curves \(a\) and \(b\) (Fig. 9). This will be achieved if \(\mu = 0.70\) is adopted. Under this assumption the curve \(c\) in Fig. 9 is constructed. It further follows from this that if, with a continuous distribution of energy over the entire sound spectrum, a tone of definite pitch is clearly perceived, then 70% of all the sound energy must be concentrated around the tone being determined within the limits of \(\pm 5\%\). By means of further calculations one can obtain that if, for

\[ \frac{\Omega}{2\pi} = 65 \text{ Hz} \]

one chooses the switching-on time \(T\), equal to \(\frac{1}{8}\) of the minimum time necessary for the clear establishment of pitch,

then \(\frac{E_u}{E_g}=0.31\), and \(\frac{E_o}{E_g}=0.3725\). This means that a larger amount of energy falls in the range of frequencies lying above the frequency of the sound impulse than in the lower-lying range. With similar assumptions for \(\frac{\Omega}{2\pi}=9100\ \mathrm{Hz}\), the following ratios are obtained, respectively: \(\frac{E_u}{E_g}=0.504\) and \(\frac{E_o}{E_g}=0.009\).

In this case the “center of gravity” of auditory perception is shifted toward lower frequencies. These theoretical considerations lead to the conclusion that very short sound impulses of low pitch should seem higher to the ear, while very short sound impulses of high pitch should seem lower.

The assumption made concerning the value \(\mu=0.70\) was confirmed experimentally. The function represented by formula (28), giving the frequency spectrum of a short tone sounding for a time \(T\), has a form similar to the resonance curve of a simple oscillatory system, with a maximum corresponding to the frequency \(\Omega\), and lying the higher the longer the duration of sounding \(T\). When the duration of sounding \(T\) is decreased, the curve becomes flatter. A frequency spectrum of the same character can also be obtained in another way. The amplitude spectrum of a short impulse can be represented in the form of a segment of a horizontal straight line. If, by means of such an impulse, a simple oscillatory circuit is excited, then its oscillations can be represented in the form of a frequency spectrum corresponding to its resonance curve, i.e. in a form essentially similar to the frequency spectrum of a very short tone. The broadening of the resonance curve that occurs as a consequence of decreasing the duration of sounding can here be achieved, with the same result, by increasing the damping of the oscillatory circuit. The experiments were arranged so that short impulses, by means of an appropriate circuit, were applied to excite a suitable oscillatory circuit with adjustable damping. The damping was selected of such magnitude that the impulse being listened to would still possess a definite tonal character, the pitch of which could be established with certainty. After squaring the frequency spectrum obtained with this damping and calculating the area by means of a planimeter, it was found that in the region located about the resonance frequency within \(\pm 5\%\) of it, \(68.6\%\) of all the energy was concentrated. This result is in excellent agreement with that obtained earlier.

An approximate value of the time for recognizing the pitch of a tone can be obtained on the basis of a principle formally resembling the uncertainty principle, as was first done by Stewart[^31]. This principle can be formulated by the inequality: \(\Delta\nu\cdot \Delta t \geq 1\). This means that the frequency of a tone sounding during a time interval \(\Delta t\) can be determined with an error not less than \(\Delta\nu\). The curve in Fig. 9 in the region of low

up to 1000 Hz corresponds to this result, if it is assumed that the frequency is determined with an accuracy of a whole tone.

Additional data concerning the functioning of the human organ of hearing were obtained by Bouman^89 in studying the question of the hearing’s discrimination of rapid, brief changes in the pitch of a continuously sounding tone. Changes in pitch of very short duration are perceived by the ear as a sound having no tonal character. Perception of pitch occurs only beginning with some minimal duration of sounding of the intermediate tone, the magnitude of which is inversely proportional to its frequency.

The number of periods required for perception of the pitch of the intermediate tone proved to be independent of the pitch and approximately equal to 8. It follows from this that to each time interval \(d\) there corresponds a quite definite limiting pitch of the intermediate tone, \(\frac{8}{d}\). This result obtained by Bouman in the frequency range from 200 to 1400 Hz is generally in agreement with the data presented in Fig. 9. Further, it was found that, in the case of very short intermediate tones, which are perceived by the ear as a sound without a definite tonal character, the presence of about a 40 Hz difference between the frequencies of the continuously sounding tone and the short intermediate tone is required in order for the intermediate tone to be perceived at all by the ear. If, however, the intermediate tone continues to sound for so long an interval of time that it is perceived by the ear as such, then the normal values of the threshold of pitch perception are valid for it. From these data the author concludes that the mechanism for sound analysis is not located in the central nervous system, which is not capable of reacting to such short time intervals. When listening to two intermediate tones of the same frequency, following one another after definite intervals of time, the limiting pitch of the tone is obtained from twice the magnitude of the sounding interval of the intermediate tone if subjective fusion, or “summation,” of the intermediate tones is observed. Otherwise, in the absence of subjective fusion, the limiting pitch of the tone is determined from the sounding interval of the individual intermediate tone. For example, at a continuously sounding tone frequency of 100 Hz and an intermediate tone of 530 Hz, it was established that, with a separating interval of 100 msec and less, summation of the two intermediate tones is observed. This summation occurs according to the ordinary psychological laws: it does not depend on the duration of the individual impulse, nor on the number and intensity of the individual impulses; consequently, summation is a function of the nervous system. Summation is also observed at different pitches of the intermediate tones. The greater their loudness, the farther apart may be the frequencies of the two intermediate tones that give the subjective impression of summation. From these data it follows that

the physiologically active width of the basilar membrane increases with increasing loudness. In conclusion, the author investigated a large number of sounds of different frequencies, following one another not systematically, but at such large intervals of time that summation did not take place. It turned out that in this case a sensation of noise arises. On the basis of this observation the author concludes that noises, like tones, are perceived by the organ of Corti. In order for the impression of noise to arise, the entire process must last from 200 to 400 msec, depending on the intensity.

Noises with a continuous frequency spectrum were investigated by Tiede ^84.

To obtain noises, discharges in an $\alpha$-particle counter, the Barkhausen effect, microphone noises, and the shot effect were used. The frequency spectrum of the shot effect has the special feature that its individual components over the whole range have one and the same magnitude. With these noises, about 200 msec is also required to obtain the impression of noise.

7. Auditory perception of transient processes. According to Békésy’s observations ^59, when a tone ceases to sound in accordance with a regular exponential law, the ear subjectively perceives a monotonic decrease of the sound as long as the time of decrease to the threshold of auditory perception does not exceed 0.8 sec. With a longer time of decrease, which proceeds monotonically throughout, the ear subjectively perceives a minimum 0.8 sec after the beginning of the process of decrease, independently of the initial loudness. The reason for this lies in the fact that a continuous process lasting longer than a certain interval of time is no longer perceived as a single whole in the sense in which this is defined in psychology. Such an interval is called the “time of presence.” This property of hearing is reflected also in language: words with more than four syllables are very rare. Since on average the time of pronouncing each syllable is 0.2 sec, the time of pronouncing a four-syllable word is just the “time of presence.”

In later works by the same author ^36, the question of the influence of transient processes on auditory perception was studied in detail. The time of decay (or ringing time, Ausschwingdauer) is the time during which a tone decays from its initial value to the threshold of auditory perception. The subjective perception of a sound that is ceasing, depending, besides the decay time, also on the loudness level ^1) and on the frequency, can be defined as the strength of the ringing (Ausschwingstärke). When the ringing time is decreased, there arises such a sensation in which the ear

^1) The loudness level of any sound is called the level of intensity of an equally loud tone of 1000 Hz. Per unit of loudness level the phon is taken, corresponding to the previously established term decibel.
Translator’s note.

subjectively no longer able to distinguish between an instantaneous cessation of sound and a decaying sound. The Sabine reverberation time corresponding to this sensation (i.e., the time during which the amplitude of the sounding tone decreases to 0.001 of its initial value) may be defined as the physiological reverberation time, and the auditory sensation corresponding to it as the physiological strength of reverberation. This terminology may also be applied to processes of sound build-up (Nazvuk) and, in general, to all transient processes. The measurement of the physiological reverberation time is carried out in such a way that, with a gradual reduction of the decay time, a sensation arises in which the ear can no longer distinguish the attained decay process from the same one characterized by a decay time 10 times smaller. In these measurements an increase in the physiological reverberation time of a sound with decreasing loudness level was established (Fig. 10).

Fig. 10. Physiological reverberation time at various loudness levels (after Bekesy)

Fig. 10. Physiological reverberation time at various loudness levels (after Bekesy)

For the physiological time of sound build-up, one and the same value, 0.07 sec, was obtained at different loudness levels. In more complex processes, as Ashov noted \(^{93}\), one may observe a difference between substantially shorter transient processes. Evidently, in perception here it is not only the duration of the subjective impression that matters, but also the shape of the envelope curve, as well as the character of the change in the amplitudes of the individual overtones with the passage of time. Bekesy found that the magnitude of the difference threshold for the reverberation time of instantaneously beginning transient processes is approximately \(10\%\). If, however, the transient process is followed by a tone sounding for a longer time, then in short processes an increase in the difference threshold is observed. If the duration of the build-up process is chosen as 1.5 sec, then at a loudness level of 60 phons the difference threshold proved to be approximately \(16\%\) for frequencies of 800 Hz and above; at lower frequencies the magnitude of the difference threshold proved greater and at 100 Hz amounted to about \(27\%\). With respect to the influence of changes in the course of the decay process on auditory perception, it was established that if during the decay process the decay time is changed by 10—\(20\%\), then this is subjectively perceived as a jump in the course of the decay curve. This is observed with longer decay times—on the order of 1—2 sec. With shorter decay times, an increase is observed in the percentage change of the decay time still distinguishable by the ear. In sound build-up processes, the magnitude of the percentage change in the build-up time that is barely perceived by the ear is greater and amounts to 20—\(30\%\) in slow build-up processes.

In conclusion of these experiments, by interrupting the build-up process

until it reaches its limiting amplitude value, with preservation of the value attained at the moment of interruption, it was established that a rising process, upon reaching 75% of its stationary amplitude, is perceived by the ear as already completed.

The works of Steinel\({}^{38}\) and Bürck, Kotovsky and Lichte\({}^{67,80}\) are devoted to the investigation of the question of the loudness level of sound impulses. The latter authors showed, moreover, that the loudness level can be calculated with good approximation on the basis of the form of the sound impulse, proceeding from certain plausible assumptions concerning the mechanism of action of human hearing. In this, the above-stated conceptions of the mechanism of action of hearing were very convincingly confirmed.

The process of increase of a direct current, occurring according to the law

\[ i=i_0\left(1-e^{-\frac{t}{T}}\right), \]

can be represented on the basis of (22) in the form of the following amplitude spectrum:

\[ \frac{A(\omega)}{i_0}=\frac{1}{\pi\omega}\frac{1}{\sqrt{1+\omega^2T^2}}. \]

In this case, for an unsmoothed impulse with amplitude \(a<1\),

\[ A(\omega)=\frac{a}{\pi\omega} \]

according to (2). To calculate the loudness level it is necessary to sum the energy of the separate parts of the spectrum. In a first approximation one may discard the highest and the lowest frequencies and assume a constant sensitivity of the ear in the range from \(\omega_1\) to \(\omega_2\). In this case the loudness level of an impulse of smoothed form is obtained as equal to

\[ L_A=\sqrt{\pi\int_{\omega_1}^{\omega_2}\left[\frac{A(\omega)}{i_0}\right]^2d\omega}= \]

\[ =\frac{1}{\pi}\sqrt{\frac{1}{\omega_1}-\frac{1}{\omega_2}+T\left(\operatorname{arc\,tg}\omega_1T-\operatorname{arc\,tg}\omega_2T\right)}, \]

and of the unsmoothed form

\[ L_P=\frac{a}{\sqrt{\pi}}\sqrt{\frac{1}{\omega_1}-\frac{1}{\omega_2}}. \]

The measurements consisted in the loudness levels of the smoothed and unsmoothed impulses being compared with one another and equalized by changing the magnitude of the unsmoothed impulse. From the equation \(L_A=L_P\) the quantity \(a\) can be determined:

\[ a^2=1+\frac{T}{\frac{1}{\omega_1}-\frac{1}{\omega_2}}\left(\operatorname{arc\,tg}\omega_1T-\operatorname{arc\,tg}\omega_2T\right). \]

Experimental and theoretical results are presented in Fig. 11, in which the loudness levels are given as a function of the time constant \(T\). Curve \(a\) was obtained experimentally. Curve \(c\) was calculated on the basis of the considerations given above in the frequency range used in telephone equipment, namely between 300 and 2400 Hz. Still greater agreement can be obtained if the limits of integration are chosen in accordance with the loudness differences on the basis of equal-loudness curves, so that the sensitivity of the ear varies within \(\pm 6\) phons. In this way curve \(d\) of Fig. 11 was obtained. On the basis of the excellent agreement of the theoretical conclusions with the experimental data it follows that, in the perception of sound

Fig. 11. Loudness levels of sound impulses

Fig. 11. Loudness levels of sound impulses

impulses, the ear may be regarded as a linear receiver. Steidel \(^{38}\) measured the loudness levels of impulses in a telephone, arising upon the sudden switching off of a direct current decaying according to an exponential law, for various time constants. A calculation carried out on the basis of the amplitude spectrum (22a) leads to the following expression \(^{67}\):

\[ L = \sqrt{\frac{T}{\pi^{3}} \operatorname{arc\,tg} \frac{T(\omega_{2}-\omega_{1})}{1+\omega_{1}\omega_{2}T^{2}} } . \tag{29} \]

If, in the frequency range from \(\omega_{1}\) to \(\omega_{2}\), over which the integration is performed, the same considerations are used as were applied in obtaining curve \(d\) of Fig. 11, then the dashed curves \(a'\), \(b'\), \(c'\) of Fig. 12 are obtained, calculated for various maximum values of the current strength. The curves \(a'\), \(b'\), \(c'\) are drawn so that, for the time constant \(T = 10^{-3}\) sec., they coincide with the curves \(a\), \(b\), \(c\), experimentally obtained by Steidel. The agreement of the theoretical and experimental results is quite satisfactory, except in the region of very low loudness levels. The reason for this discrepancy evidently lies in the fact that,

that the adopted upper limit of the range, 7500 Hz, is too high for the telephone equipment used in the measurements.

The perception of loudness does not follow the stimulus instantaneously, but a certain inertia of auditory perception is observed. The corresponding measurements were carried out by Steudel\(^{88}\), who determined the subjective decay of loudness of a periodically interrupted tone by masking it with a constant tone varying in intensity, and found that the decay occurs according to an exponential law with an average time constant of 50 msec. Further measurements clarifying this question were carried out by Békésy\(^{25}\), who observed the increase in loudness of a tone switched on for a short time as a function of the time of switching on. From these measurements a value of the time constant of 130 msec was derived. The difference between the values obtained by Békésy and Steudel is rather considerable and must in part be attributed to the individual peculiarities of the observers in these very difficult measurements. In addition, it is possible that the results of Békésy’s measurements were affected by noise when the tone was switched off. In any case, on the basis of these measurements one can derive a differential equation for the loudness level\(^{67}\), with the aid of which the course of the loudness level can be calculated.

Fig. 12

Fig. 12. Measured \((a, b, c)\) and calculated \((a', b', c')\) values of the loudness level of sound impulses with different initial amplitudes and with different decay times (according to Bürck, Kotowski, and Lichte)

As was already indicated above, in the perception of sound impulses the ear may be regarded as a ballistic instrument whose readings depend on the energy imparted to it, in analogy with the readings of a thermal instrument. If the sound pressure is denoted by \(p\), then the estimated loudness is proportional to

\[ L=\sqrt{\int_{-\infty}^{t} p^{2}\,dt}. \]

Hence, assuming that the losses are proportional to the energy of the system, one can obtain the following differential equation:

\[ a\,\frac{d(L^{2})}{dt}=p^{2}-cL^{2}. \]

It follows from the equation that after the instantaneous switching off of a sinusoidal tone the ear still continues to preserve the sensation of sound, decreasing according to the exponential law

\[ L=L_{0}e^{-\frac{c}{2a}t}=L_{0}e^{-\frac{t}{T}}, \]

where \(T=\dfrac{2a}{c}\) denotes the time constant of hearing. The onset of a sinusoidal tone is perceived as occurring according to the following law:

\[ L^2=\frac{p_0^2}{4a}\,T(1-e^{-\frac{2t}{T}}). \tag{30} \]

Steidel \(^{38}\) compared the loudness level \(L\) of a pure tone with the loudness level \(L'\) of a sound impulse [according to (29)], with the initial value of the pressure of the exponentially decaying sound impulse set equal to the amplitude of the pressure of the sinusoidal tone. From formulas (29) and (30) one obtains the following relation of loudness levels:

\[ V=\frac{L}{L'}= \sqrt{\frac{T\pi}{4T\,\operatorname{arc\,tg}T\,\dfrac{\omega_2-\omega_1}{1+\omega_1\omega_2T^2}}}. \]

The values calculated by this formula agree well with those measured experimentally by Steidel (Fig. 13).

To elucidate the mechanism of the action of hearing, experiments with short clicks proved especially fruitful: the clicks were listened to in a telephone connected to an oscillatory circuit with damping corresponding to the limiting aperiodic regime and excited by short current impulses. If it is assumed that the ear functions like such an oscillatory circuit, and that the curve of equal loudness at 50 phon corresponds to an oscillatory circuit with resistance \(3500\ \Omega\), inductance \(0.1\ \mathrm{H}\), and capacitance \(33400\ \mu\mu\mathrm{F}\), then, taking into account the amplitude spectrum expressed by formula (22), the following value of the loudness level is obtained:

\[ L=\int_{0}^{\infty} \frac{1}{T^4} \frac{1}{\left(\dfrac{1}{T^2}+\omega^2\right)^2} \frac{1}{4+\left(\dfrac{\omega}{T'}-\dfrac{T'}{\omega}\right)} \,d\omega, \]

where the damping coefficient of the circuit equivalent to the ear is taken to be \(T'=1.75\cdot 10^4\ \mathrm{sec}\).

After integration, the following formula is obtained:

\[ L=\frac{\pi}{4T}\, \frac{T'}{\left(\dfrac{1}{T}+T'\right)^3}. \]

In Fig. 14, for comparison, are presented the values of the loudness level calculated by this formula (curve \(b\)) and the values obtained experimentally by Steidel \(^{38}\) (curve \(a\)). If one takes into account that in the experimental determinations of the frequency

up to 300 Hz are damped considerably more strongly, then excluding the computed energy falling at these frequencies (curve c) from the total amount of energy computed by the formula gives curve d, which agrees even more closely with curve a.

The results presented show that, in calculating the loudness level of sound pulses, the ear may be regarded as a linear receiver with a frequency characteristic corresponding to the ear-sensitivity curve and with a time constant approximately equal to 50 msec. In these experiments no

Fig. 13

Fig. 13. Comparison of the loudness level of sound pulses and prolonged tones: a—curve obtained experimentally by Steidel; b—calculated; c—loudness level of a prolonged tone (according to Bjork, Kotowski, and Lichte)

Fig. 14

Fig. 14. Loudness level of sound pulses produced by current pulses in an aperiodic oscillatory circuit (according to Bjork, Kotowski, and Lichte)

indications were obtained of any noticeable influence of the nonlinearity of the organ of hearing on the magnitude of loudness in the sense of the Weber–Fechner (logarithmic) law.

The question of distortions of auditory perception introduced by transient phenomena was investigated in detail by Bjork, Kotowski, and Lichte83. By means of a measuring installation consisting of a condenser microphone, a three-stage resistance amplifier, a power amplifier, and a reproducer, time constants were measured under various loads and at various frequencies. Auditory perception caused by a sinusoidal tone suddenly switched on by means of the appropriate circuit was compared with the auditory impressions produced in a telephone set connected into an oscillatory circuit with a known and adjustable time constant. In this case, in accordance with the data presented in Fig. 7, positive time constants are observed in the regions of resonance, and negative ones in the dips of the frequency characteristic of the installation. The difference between the extreme

with the values of the time constant proved to be independent of the power. However, with an increase in the transmitted power, a displacement of the time constant during the build-up toward negative values is observed, i.e., an increase in the impulsiveness of the sound. The maximum difference in the time constant was, for the loudspeaker of low quality under investigation, about 7 msec and caused clearly noticeable distortions during the build-up of the oscillations. For loudspeakers of better quality and with a more even frequency characteristic, the maximum difference proved to be only 3.5 msec.

A very convenient circuit for objective measurements of the processes of build-up of oscillations was developed by Burke, Kotowski, and Lichte^[83]. The oscillatory process at the input and at the output of the transmitting device was rectified and fed to a ballistic galvanometer in opposition to one another. The apparatus was calibrated with the aid of a Helmholtz pendulum. The deflection of the galvanometer served as a measure of the distortions during the process of build-up of oscillations. The direction of the deflection depended on whether switching-on was carried out in resonant regions or in the valleys of the frequency characteristic.

  1. Modulated tones. A prolonged steady sound, even one rich in overtones, produces a monotonous impression. Special techniques are used to enliven the sound. This is done in many musical instruments, especially those preferred precisely because of this possibility of modulation afforded by them, and also in the singing voice. Obata, Hirose, and Tezima^[56] report on an extremely peculiar form of modulation that is used in the Japanese school of singing and consists in a rapid change of pitch by a whole tone, occurring with a modulation frequency of 4 Hz. In his extensive works, Bartolomeo^[54] investigated 40 singing voices and found that in singers with well-trained voices, during vibrato the pitch, intensity, and timbre change simultaneously. The modulation frequency in this case proves to be very constant and amounts to 6–7 Hz for different voices. The maximum of the intensity, falling on the overtones, is thereby displaced within limits exceeding an octave.

The theory of modulation is set forth in the work of Tolmie^[75], who developed previously known calculations and applied them to the solution of certain problems of acoustics. Pure frequency modulation can be represented by the following equation:

\[ f(t)=\sin(\omega t+m\sin \mu t); \tag{31} \]

here \(\frac{\omega}{2\pi}\) denotes the frequency of the modulated tone, the so-called fundamental frequency, \(\frac{\mu}{2\pi}\) denotes the modulation frequency, \(m=\frac{\Delta\omega}{\mu}\) is the modulation depth. In this case additional conditions are imposed:

\[ \Delta\omega \ll \omega,\quad \mu \ll \omega. \]

The function expressed by (31) can be expanded in a Fourier series

\[ f(t)=\sum_{\nu=-\infty}^{+\infty} J_\nu(m)\sin(\omega t+\nu\mu t). \tag{32} \]

Here \(J_\nu\) denotes the Bessel function of order \(\nu\).

In accordance with the practical case, the modulation frequency \(\mu\) may be regarded as constant. Further, the width of the modulation band must be the same, for example equal to \(1/4\) tone. In this case \(\Delta\omega\) is proportional to \(\omega\), and the coefficient of the series \(J_\nu(m)\) tends to 0, as \(\nu\) increases, the faster the lower the modulated tone is.

Fig. 15. Spectra of frequency-modulated tones (after Tolman)

Fig. 15. Spectra of frequency-modulated tones (after Tolman)

The overtones of the expansion are arranged symmetrically at equal frequency intervals about the fundamental tone, as follows from Fig. 15 for various fundamental tones.

If (32) is written in complex form,

\[ f_1(t)=\sum_{-\infty}^{+\infty} J_\nu(m)e^{i(\omega t+\nu\mu t)} \]

and the vector method of representing oscillations and the known rules of Bessel functions are applied to the angular velocity \(\omega\) of the rotating vector, then we obtain

\[ f_2(t)=e^{im\sin\mu t}. \]

This vector oscillates within the limits of an angle \(2m\), expressed in radians, with angular velocity \(m\mu\cos\mu t\) about a vector rotating with angular velocity \(\omega\).

If, along with frequency modulation, there is also amplitude modulation with modulation depth \(k\), then such a process can be represented by the following function:

\[ F(t)=[1+k\cos(\mu t+\varphi)]\sum_{-\infty}^{+\infty} J_\nu(m)\sin(\omega t+\nu\mu t), \]

or, in complex form, dropping the factor \(e^{i\omega t}\),

\[ F_1(t)=\left\{1+\frac{\gamma k}{2}\left[e^{i(\mu t+\varphi)}+e^{-i(\mu t+\varphi)}\right]\right\}\sum_{-\infty}^{+\infty}J_\nu(m)e^{i\nu\mu t}. \]

After simple calculations, based on applying the recurrence formula for Bessel functions, we obtain

\[ F_1(t)=\sum_{-\infty}^{+\infty}\sin(\omega t+\nu\mu t+\psi)\times \]

\[ \times\sqrt{\left[1+\frac{\nu k}{m}\cos\varphi\right]^2J_\nu^2(m)+k^2\sin^2\varphi J_\nu'^2(m)} \tag{33} \]

where

\[ \psi=\operatorname{arc\,tg}\frac{k\sin\varphi\,J_\nu'(m)} {\left(1+\dfrac{\nu k}{m}\cos\varphi\right)J_\nu(m)}. \]

In artistic performance, modulation in frequency and in amplitude occurs either with the same or with the opposite phase, i.e. \(\varphi=0\) or \(\varphi=\pi\).

In this case (33) takes the following form:

\[ F_1(t)=\sum_{-\infty}^{+\infty}\left(1\pm\frac{\nu k}{m}\right)J_\nu(m)\sin(\omega t+\nu\mu t), \]

where the positive sign is taken when \(\varphi=0\), and the negative sign when \(\varphi=\pi\). The last formula differs from the formula given earlier for purely frequency modulation only by the factor \(\left(1\pm\dfrac{\nu k}{m}\right)\). This produces an asymmetric arrangement of the spectrum; Fig. 16 shows that for \(\varphi=0\) the higher overtones are intensified, while for \(\varphi=\pi\) the lower ones are intensified. This entails a discrepancy between the “center of gravity” of auditory perception and the fundamental frequency, and such a combination of frequency and amplitude modulations causes a subjectively perceived noticeable shift of the fundamental frequency. This asymmetry is eliminated when \(\varphi=\dfrac{\pi}{2}\).

Fig. 16. Spectra of tones modulated in frequency and in amplitude (after Tolmie)

Fig. 16. Spectra of tones modulated in frequency and in amplitude (after Tolmie)

In the musical instruments in use such a case does not occur. Tolmie points out that on electromusical instru-

instruments, a similar effect can be realized under certain circumstances.

Applying the uncertainty principle in the form \(\Delta \nu \cdot \Delta t \geq 1\), Kok\(^{22}\) found that a certain frequency modulation with a given frequency and depth of modulation, perceived by the ear as such at a sufficiently high fundamental frequency, can in the low range be perceived only when the modulation frequency is lowered or when the modulation depth is increased, provided that the modulation frequency remains constant.

Bartholomew believes that one of the chief requirements imposed on the singing voice is good “vibrato.” According to experimental investigations, it was established that purely frequency modulation, produced with the aid of electrical devices, does in fact begin to be perceived only when a certain magnitude of modulation depth is reached. Before this point, the tone seems unsteady; the auditory perception resembles amplitude modulation.

(To be continued)

  1. Ergebnisse d. Exakt. Naturwiss. 16, 237, 1937. Translation by B. G. Shpakovsky. 

  2. The literature will be given at the end of the article. 

Submission history

STEADY-STATE PROCESSES IN ACOUSTICS[^1]