Timbre and acoustics

A vowel is two resonances

The vowel in "heed" is the same vowel sung high or low, and nothing about it is a property of the note. It is two peaks in the response of the mouth, sitting at fixed frequencies while the partials of the voice slide underneath them.

Assumes: The ear hears the list, not the shape · A string does everything at once

Sing “ee” on a low note and then on a high one. The pitch changes; the vowel does not. That is unremarkable until it is put beside what a note actually consists of, at which point it becomes a genuine puzzle: the whole spectrum has moved, every partial is at a different frequency, and the ear reports the same vowel.

The resolution is that the vowel was never in the partials. It is in something that did not move — which makes it the same kind of fact as a pitch that survives having its fundamental removed.

The vowel in "hod", sung at 110 HzThe partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.F1730 HzF21090 Hz050010001500200025003000hertzamplitudethe filter — the shape of the mouththe partials — the pitch being sungPeterson & Barney, 1952 — mean adult male values11 partials carry most of the identity
Fig. 1 The partials of a 110 Hz note, each drawn at the amplitude the vocal tract’s resonances give it, with the resonance curve behind them. The peaks are the formants; they belong to the shape of the mouth and have nothing to do with the note being sung.

Two things, multiplied

The standard account is the source–filter model, and it separates the voice into two independent parts.

The source is the vocal folds. They chop the airflow into a pulse train at whatever rate they are vibrating, which produces a harmonic series: energy at the fundamental and at every whole-number multiple of it, falling off with frequency. That is the pitch, and only the pitch: the harmonic series a vibrating thing produces, at whatever rate the folds are running.

The filter is everything above the folds — the pharynx, the mouth, the position of the tongue and jaw and lips. It is a tube of a particular shape, and like any tube it resonates at particular frequencies. Those resonances are the formants, and they boost whatever partials happen to fall near them.

The output is the source spectrum multiplied by the filter’s response, and the two are independent — which is why a spectrum is a list rather than a shape. Change the pitch and the source moves while the filter stays. Change the vowel and the filter moves while the source stays.

That independence is why a language can carry pitch and vowel at the same time without them interfering, and it is why singing works at all.

Which computation produced the number

The filter curve is not drawn freehand. Each formant is modelled as a two-pole resonator, whose amplitude response at frequency ff is

H(f)=1(1(f/fc)2)2+(f/(Qfc))2,Q=fcbandwidth,|H(f)| = \frac{1}{\sqrt{\left(1 - (f/f_c)^2\right)^2 + \left(f/(Q f_c)\right)^2}}, \qquad Q = \frac{f_c}{\text{bandwidth}},

and the bank of three is combined in power. The source is taken as falling at 6 decibels per octave — the usual simplification, since a glottal pulse train falls at about 12 and radiation from the lips adds back about 6.

Each partial’s amplitude is then the filter’s gain at that partial’s frequency divided by the partial number, and those amplitudes go straight to the synthesiser. So the buttons under these figures play the spectrum that is drawn, not an approximation of it: one list of numbers produces both.

The formant frequencies themselves are measurements. They come from Peterson and Barney’s 1952 survey of American English vowels, averaged over their adult male speakers — which is a claim about a language, a dialect and a population, and not about vowels in general.

The vowel in "heed", sung at 196 Hz. The partials of a 196 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 270 Hz and 2290 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.
Fig. 2 The vowel in “heed” at 196 Hz, with the curve for “who’d” drawn behind it for comparison. The two have first formants within thirty hertz of each other and second formants fourteen hundred hertz apart, which is the entire difference between them.

The map of vowels

Plot every vowel by its first two formants and the result is a map, and the map turns out to be a picture of the mouth.

Every vowel, placed by its first two formants. The nine vowels of the Peterson and Barney survey plotted by their first two resonances, with both axes reversed so the layout matches the position of the tongue. The horizontal lines are sung pitches: above one of them, the first formant is below the fundamental, and no partial of the note can excite it.
Fig. 3 The nine vowels of the Peterson and Barney survey placed by their first two resonances, with both axes reversed so the layout matches the position of the tongue. The horizontal lines are sung pitches, and they matter for reasons the next section is about.

The first formant tracks how open the mouth is: an open vowel like the one in “hod” has a high first formant, a close vowel like “heed” a low one. The second formant tracks how far forward the tongue is: front vowels have high second formants, back vowels low ones.

Reversing both axes therefore produces a diagram in which each vowel appears roughly where the tongue is when making it, which is why the vowel quadrilateral in every phonetics textbook looks the way it does. It is not a convention; it is a measurement drawn with the axes flipped.

Two formants are enough to distinguish the vowels of most languages. The third contributes to individual voice quality and to a few distinctions — the American English “r” vowel has a very low third formant, which is its main acoustic signature — and above the third the formants are largely speaker identity rather than phoneme.

The whispered proof

There is a demonstration that settles the source–filter claim more sharply than any figure, and it costs nothing.

A whisper has no vocal-fold vibration at all. The source is turbulent noise from a partly closed glottis — broadband, with no harmonic structure and no pitch. Every partial in the figures above is absent.

Vowels in a whisper are completely identifiable. Whispered speech is intelligible across a room.

That is decisive. If vowel identity lived in the partials, removing the partials would remove the vowel, and it does not. What survives is the filter: the noise source has energy everywhere, so the resonances have something to boost at every frequency, and the peaks appear in the output as bands of noise rather than as boosted harmonics.

The same argument runs the other way. A pure sine tone passed through a vocal tract produces no vowel at all, because there is nothing at the formant frequencies for the tract to work on. It is the same failure as the soprano case, taken to its limit.

So the source contributes pitch and nothing else, and the filter contributes vowel and nothing else, and each can be removed without touching the other. It is unusually clean, as separations of this kind go.

What a glissando does to the loudness

The independence has a consequence that is audible in ordinary singing and that almost nobody notices until it is pointed at.

As a singer slides continuously upward, every partial slides with the pitch — but the formants do not move. So each partial sweeps through the resonance peaks in turn, being amplified as it passes one and falling away after. The loudness of the note therefore does not rise smoothly; it ripples.

The vowel in "heed", sung at 110 HzThe partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 270 Hz and 2290 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.F1270 HzF22290 Hz050010001500200025003000hertzamplitudethe filter — the shape of the mouththe partials — the pitch being sungand "hod", for comparisonPeterson & Barney, 1952 — mean adult male values7 partials carry most of the identity
Fig. 4 The partials of a 110 Hz note at the amplitudes two different tracts give them — /i/ against /ɑ/. The peaks are the formants, 270 and 2290 hertz for the first, and they stay where they are when the pitch changes because they are a property of the shape of the mouth rather than of the note.

That is the source–filter model drawn once rather than argued: the bars are the source, the curve is the filter, and a vowel is which curve. Nothing about the bars changed between the two readings.

The effect is largest for low voices, where the partials are close together and several are inside a formant at once, and it is largest of all for the fundamental itself in a high voice, where the whole loudness of the note depends on whether the fundamental happens to be near the first formant.

Two practical consequences follow. Singers learn to compensate, adjusting the tract shape continuously to keep the loudness even, which is part of what is meant by an even scale. And synthesised singing that keeps the formants fixed and slides the partials sounds correct in a way that synthesised singing which slides the whole spectrum together does not — the second is what happens if a recorded note is simply pitch-shifted, and it is why a pitch-shifted voice sounds like a different-sized person rather than the same person singing higher.

That last observation is worth stating as a rule, because it generalises. Shifting an entire spectrum multiplies the formant frequencies along with the fundamental, which is acoustically the signature of a shorter vocal tract — a smaller head. Shifting only the source keeps the formants and changes the pitch, which is the same speaker singing a different note. The two operations produce entirely different impressions, and the difference is precisely the separation this essay is about.

The problem that has no solution

The independence of source and filter has one hard limit, and it is the reason sopranos sing the way they do.

The filter boosts partials that fall near its peaks. If no partial falls near a peak, the peak does nothing — a resonance can only amplify energy that is present, and the energy is only present at multiples of the fundamental.

For a bass singing at 100 hertz, the partials are 100 hertz apart, and every formant has several partials to work with. For a soprano at 1000 hertz, the partials are 1000 hertz apart and the first of them is the fundamental — which is above the first formant of every vowel in the survey. The highest first formant among the nine is 730 hertz, for the vowel in “hod”.

So above about 700 hertz, the first formant has nothing to amplify. It is still there, the mouth is still the right shape, and no partial is anywhere near it.

The vowel in "hod", sung at 880 Hz. The partials of a 880 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.
Fig. 5 The vowel in “hod” at 880 Hz. The first formant, at 730 Hz, sits below the fundamental — there is no partial at or near it, and the peak is amplifying nothing. Below about this pitch the vowel is available; above it, it is not.

This is the acoustic content of the observation that high soprano singing is hard to understand. It is not a matter of diction or of effort. Above the first formant the information that distinguishes one vowel from another is largely absent from the signal, and no amount of care recovers it.

What singers do about it is well documented and is one of the more striking pieces of instrument-adaptation in music. Trained sopranos raise the first formant to meet the fundamental, principally by opening the jaw, and the amount they open it increases with pitch — so the mouth shape for a given vowel on a high note is not the mouth shape for the same vowel on a low one. Johan Sundberg’s measurements from the 1970s document the strategy, which is called formant tuning, and it means the vowel is being sacrificed to keep the note loud.

What a tube resonates at

The formants are not arbitrary numbers; they follow from the shape of a tube, and the simplest case is worth doing because it gets the right order of magnitude from nothing but a length.

A tube closed at one end and open at the other resonates at odd multiples of c/4Lc/4L, where cc is the speed of sound and LL is the length. An adult male vocal tract is about 17 centimetres from glottis to lips, so with c=343c = 343 metres a second the resonances are at

3434×0.17=504 Hz,1512 Hz,2520 Hz.\frac{343}{4 \times 0.17} = 504 \text{ Hz}, \quad 1512 \text{ Hz}, \quad 2520 \text{ Hz}.

Those are the formants of a neutral tract — one held as a uniform tube, which is roughly the shape for the vowel in “hud”. The measured values for that vowel are 640, 1190 and 2390, so the calculation is out by twenty per cent or so on the first two, which for a model consisting of one length and no anatomy is respectable.

Everything a speaker does with the tongue, jaw and lips is a departure from that uniform tube, and every vowel is a particular departure. Constricting the tube near a pressure antinode of a mode lowers that mode’s frequency; constricting near a node raises it. Which is why raising the tongue at the front — as in “heed” — lowers the first formant and raises the second, and why the vowel map has the shape it does.

The same calculation explains why the whole map scales with the speaker. A child’s tract is perhaps 11 centimetres, so its resonances are around 780, 2340 and 3900 hertz — every formant a factor of 1.5 higher, and the vowel space shifted bodily upward. That is the scaling a listener normalises for without effort, and it is why a child saying “heed” and an adult saying “heed” have almost no formant frequencies in common.

Where the model stops

Source and filter are not entirely independent. The model treats them as separable, and there is real acoustic coupling between the vocal tract and the vibrating folds, particularly when a formant sits close to the fundamental — which is precisely the soprano case. The model is at its weakest exactly where the interesting question is.

Formant frequencies vary with the speaker. Peterson and Barney’s own data shows enormous scatter: the vowel spaces of men, women and children are shifted and scaled relative to each other, because the vocal tract lengths differ. A listener normalises for this effortlessly and the mechanism by which that happens is not settled.

These are steady-state vowels. The measurements are of vowels held in a fixed context. Real speech is transitions, and a good deal of the information distinguishing vowels — and almost all of the information distinguishing consonants — is in the movement between targets rather than in the targets.

Two formants is a simplification with a purpose. It works well for the vowels of the languages it has been tested on and it is not a complete account. Nasalised vowels have additional resonances and anti-resonances that the model as drawn has no place for.

The bandwidths are assumed. The figures use 80, 110 and 160 hertz for the three formants, which are conventional values. Real bandwidths vary with the vowel, the speaker and the loudness, and they change the shape of the curve substantially even when the peaks stay put — much as a room’s modes depend on its absorption as well as its dimensions.

“Substantially” is the kind of word that ought to be a number, and it is one. Two quantities are worth watching as the three bandwidths are scaled together: the depth of the dip between the first formant and the second, which is what makes a formant look like a formant, and the mean decibel distance between two vowels’ partials at a male speaking pitch, which is what makes one vowel distinguishable from another.

bandwidths dip below F1 distance between two vowels
half, 40/55/80 Hz 14.6 dB 7.92 dB
as drawn, 80/110/160 8.9 5.52
double, 160/220/320 4.1 3.85
triple, 240/330/480 2.1 2.90

So the caveat is right about the picture: halving the bandwidths deepens the valley by two thirds and doubling them fills in more than half of it, on peaks that have not moved a hertz. A figure drawn at half these values would look like a much stronger claim than the same physics.

The two columns do not fall at the same rate, though, and the difference is the reassuring part. Doubling the bandwidths costs 54 per cent of the dip and only 30 per cent of the distance between two vowels. The identity is more robust than the drawing: what separates one vowel from another is which partials are lifted, and that survives a blurring that visibly ruins the curve. Which is also why a loud or breathy voice, whose first formant is genuinely broader, stays intelligible — it pays about a third of its vowel separation for a doubling, not the two thirds the picture would suggest.

Whose music, and when

The word formant is Ludwig Hermann’s, from 1890, and the observation is Helmholtz’s from the 1860s — he identified fixed resonances in vowels using tuned resonators and reported the first formant frequencies for several German vowels. The source–filter formulation in its modern form is Gunnar Fant’s, from 1960, and it was developed for speech synthesis rather than for singing.

Musically, the phenomenon shows up in several places at once and they are usually treated as unrelated.

The singer’s formant. Trained operatic male voices show a strong peak around 2800 to 3000 hertz that untrained voices do not, produced by clustering the third, fourth and fifth formants together. It is what allows a voice to carry over an orchestra, because it sits in a frequency region where the orchestral spectrum is comparatively weak. Sundberg’s account of it is the standard one, and it is a piece of acoustic engineering achieved entirely by training — a solution to the same crowding problem a low third runs into, reached from the other side.

The vowel in "hod", sung at 110 HzThe partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.F1730 HzF21090 Hz050010001500200025003000hertzamplitudethe filter — the shape of the mouththe partials — the pitch being sungand "hod, a 11.5 cm tract", for comparisonPeterson & Barney, 1952 — mean adult male values11 partials carry most of the identity
Fig. 6 The same vowel through two tracts of different lengths — a man’s seventeen and a half centimetres against a child’s eleven and a half. A tube’s resonances scale inversely with its length, so every formant moves up by the ratio of the two, and the vowel stays the same vowel.

A vowel is not a pair of frequencies; it is a pair of ratios. A child saying /ɑ/ puts both formants half again as high as a man does and is understood immediately, which is the strongest evidence that what a listener extracts is the shape of the curve rather than the numbers on it.

Instrument formants. A violin body has resonances that do not move when the note changes, and they function exactly as vocal formants do: they boost whatever partials fall near them, so the spectrum of a violin depends on the pitch in a complicated way. The same is true of a guitar body, of a bassoon’s bore and of any room the sound passes through, which is a filter of the same kind at a different scale.

Vowel imitation. The wah-wah pedal is a swept resonance in the region of the first and second formants, which is why it sounds like a vowel. The talk box routes an instrument’s sound through a player’s actual vocal tract, which is the source–filter model taken literally. Both date from the 1960s and both work by supplying a filter to a source that has none.

Text setting. Composers who set high notes to open vowels are working with the soprano problem rather than against it, and the convention is old. Italian operatic writing puts high climactic notes on “a” far more often than on “i” or “u”, which is exactly what the acoustics recommend, and which was arrived at by ear, in the way orchestration rules generally are.

The ladder from here

Later rungs on this anchor: the singer’s formant in detail, and what clustering three resonances requires. Consonants, where the information is in transitions and noise rather than in steady resonances. Formant tuning across the soprano range, measured. Instrument body resonances, and why a violin’s response curve is a better description of it than its spectrum at any one note. And the general question of what a filter does to a source, which is the same question a room asks and answers differently.

The vowel is the part that did not move, and above a certain pitch there is nothing left for it to do.

Part 2 of 13

One essay in the series on spectrum. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 19.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

FormantResonanceSource-filterSpectrumVowel