The sound a listener knows best
Assumes: The other instrument with a reed · A vowel is two resonances
The claim that opens this essay is one everybody has tested thousands of times without noticing.
A familiar voice is identified in about a syllable. It is identified whichever vowel that syllable contains, at whatever pitch it is spoken, over a telephone that has removed everything below 300 hertz and above 3,400, through a wall, over the noise of a crowded room where several other voices are being separated out at the same time, and in a whisper — where the vocal folds are not vibrating at all and there is no pitch to identify anything by — and across the seam where the folds change vibrating mechanism entirely, which alters the source spectrum by six decibels an octave and does not make anybody unrecognisable.
Whatever the ear is using has survived removal of the fundamental, removal of the vowel, and removal of the source. That is a short list of candidates.
What the vowel is, and what it is not
The source–filter model this ladder is built on separates two things completely. The folds chop the airflow into pulses, producing a spectrum that falls smoothly at about twelve decibels per octave and carries no information about vowels at all. The tract above them resonates, and the peaks of those resonances — the formants — are what a vowel is.
So a vowel is a shape of the envelope: the first two formants at 270 and 2,290 hertz make one vowel, and at 730 and 1,090 another. Changing vowel moves the formants relative to each other.
Which immediately rules the formants out as the carrier of identity, because a speaker says every vowel. If identity were a formant position it would change with every syllable.
The length that scales everything at once
The tract’s resonances are those of a tube closed at the larynx and open at the lips, so they sit at odd multiples of the quarter-wavelength: the speed of sound divided by four times the length, and three times that, and five times that. Every one of them is inversely proportional to the same length.
That is the whole mechanism. A shorter tract multiplies every formant by the same factor.
The sizes involved are not small. A neutral tract of 17.5 centimetres — an adult male — puts uniform-tube resonances at roughly 490, 1470, 2450 and 3430 hertz. At 14.5 centimetres they are 591, 1774, 2957 and 4140. At 11.5, a child’s, they are 746, 2237, 3728 and 5220.
| Tract | Resonances, Hz | Mean spacing |
|---|---|---|
| 11.5 cm | 746 · 2237 · 3728 | 1491 Hz |
| 14.5 cm | 591 · 1774 · 2957 | 1183 Hz |
| 17.5 cm | 490 · 1470 · 2450 | 980 Hz |
From an adult man to an adult woman the whole pattern moves up by 3.26 semitones; from an adult man to a child, by 7.27 — very nearly a perfect fifth. That is an enormous displacement of the spectral envelope, and it happens with the vowel completely unchanged.
Why the spacing is the invariant
The quantity that identifies the speaker cannot be a formant, because the vowel moves those. It cannot be the fundamental, because a whisper has none. What is left is the relation between formants — how far apart they are on a logarithmic axis — which is what the tract length sets. Whether the vowel leaves it alone is a separate question, it is the one the section after next puts to the data, and the answer for the first three formants is no.
Formant dispersion, the mean spacing between successive formants, is the simplest version of this. It is 980 hertz for the long tract and 1,491 for the short one, a ratio of 1.52 — which is the inverse of the length ratio, exactly, as the tube model requires.
That ratio is exact because it is computed from the tube, whose dispersion is the same for every vowel by construction: the uniform tube has no tongue in it and therefore has no vowels. The section below runs the same measure on nine real vowels of one real speaker, which is the test the model cannot perform on itself.
That orthogonality is the reason the system works. A listener hearing an unfamiliar speaker say an unfamiliar word has to solve for two unknowns from one spectrum, and it is soluble because the two unknowns act on the envelope in independent ways — one moves peaks relative to each other, the other scales all of them together.
The name for the second half of that in the speech literature is vocal-tract normalisation, and the reason it needed a name is that it has to happen before vowel identification can work at all. A child’s “heed” has a first formant at 746 hertz, which is higher than an adult man’s “hod” at 730. Without normalising for the speaker, the two would be confused constantly, and they are not.
The invariant, measured on a speaker
Peterson and Barney’s nine vowels are already in this collection, averaged over their adult male speakers, and computing a dispersion for each of them is subtraction.
| vowel | F1 | F2 | F3 | dispersion | tract that spacing implies |
|---|---|---|---|---|---|
| heed | 270 | 2,290 | 3,010 | 1,370 Hz | 12.5 cm |
| hid | 390 | 1,990 | 2,550 | 1,080 | 15.9 |
| head | 530 | 1,840 | 2,480 | 975 | 17.6 |
| who’d | 300 | 870 | 2,240 | 970 | 17.7 |
| hood | 440 | 1,020 | 2,240 | 900 | 19.1 |
| had | 660 | 1,720 | 2,410 | 875 | 19.6 |
| hod | 730 | 1,090 | 2,440 | 855 | 20.1 |
One speaker’s dispersion runs from 855 to 1,370 hertz, a ratio of 1.60. An adult man against a child, on the same vowel, is a ratio of 1.52. The quantity proposed as the speaker-invariant varies more across one person’s vowels than it does across the whole range of human tract lengths, and a listener reading tract length off it would place this one man anywhere from 12.5 centimetres to 20.1 — the first of those being a small child’s and the second longer than any adult’s.
So the orthogonality claimed two sections above is false as stated, and it is false in the direction that matters: the vowel does not merely move the formants relative to each other while leaving their mean spacing alone. It moves the mean spacing, hard. Taking only the gap from the second formant to the third makes it worse rather than better — 560 hertz on “hid” against 1,570 on “hawed”, a ratio of 2.8.
The repair is known and this collection’s data cannot carry it. The tongue’s effect on a tract falls away for the shorter wavelengths, so the higher formants — the fourth, the fifth and above — are far less vowel-dependent than the first three, and dispersion measured over those is the quantity the speech literature actually uses. Peterson and Barney published three formants; this site carries three; and three is exactly the range in which the tongue does most of its work.
What the argument of this rung therefore rests on is narrower than it looked. The scaling claim is exact and needs no data: a uniformly scaled resonator of any shape has uniformly scaled resonances, so a change of speaker really is one multiplication. What is not established here is that any measurable function of the first three formants recovers that multiplication, and the table above says one obvious candidate does not. The mechanism is secure and the read-out is owed.
The source is not silent about identity either
Setting the filter aside for a moment is worth doing, because the division of labour is not as clean as the model suggests.
The source is not silent about identity either, and setting the filter aside for a moment is worth doing because the division of labour is not as clean as the model suggests. The shape of the glottal pulse — how long the folds stay open, how abruptly they close — sets the slope of the source spectrum, and it differs between the two laryngeal mechanisms, between speakers, and between one speaker’s moods. So identity has two carriers and they behave differently. The filter carries a scale factor fixed for a given person that survives everything; the source carries pulse shape, breathiness, jitter and shimmer, which vary with effort, health and emotion within one speaker. A listener recognising somebody who has a cold is using the first; a listener noticing that somebody is upset is using the second.
So identity has two carriers and they behave differently. The filter carries a scale factor that is fixed for a given person and survives everything; the source carries pulse shape, breathiness, jitter and shimmer, which vary with effort, health and emotion within one speaker. A listener recognising somebody who has a cold is using the first; a listener noticing that somebody is upset is using the second.
That division explains an otherwise odd asymmetry. Impersonators can do a passable version of somebody’s prosody, pulse quality and articulatory habits and are limited by anatomy in exactly one respect: they cannot change the length of their own vocal tract by more than a small fraction. The imitable part is the source and the manner; the part that gives them away is the scale.
What a whisper proves
The whisper is the cleanest test, and it is available for nothing.
In a whisper the folds do not vibrate. The source is turbulent noise at the glottis — broadband, with no fundamental and no harmonics — and everything above it is unchanged. So a whisper is the tract’s filter with the source’s structure removed.
Vowels remain identifiable in a whisper, and so do speakers. That is the argument in its strongest form: identity survives the removal of the entire source, which means identity is in the filter. And within the filter it survives the change of vowel, which means it is in the part of the filter the vowel does not move.
The same argument, one instrument along
The structure of this claim is not peculiar to voices, and this collection has already made it about a different object.
The body is the filter for a string instrument in the way the tract is for a voice, and the consequence is identical. Two violins are told apart by their bodies, not their strings; a violin is told from a viola by a scale factor on the response; and neither identification depends on which note is being played.
Where the two part company is that a violin cannot change its filter and a speaker changes theirs several times a second. That is the entire difference between an instrument and a voice, and it is why the voice needed a two-part model when the string essays did not.
The part the model gets wrong
The uniform tube is a caricature and it should be said how bad a one.
A real vocal tract is not a tube of constant cross-section; it is a tube with a tongue in it, and the tongue is what makes a vowel. The uniform model’s resonances at 490, 1470 and 2450 hertz correspond to no vowel — they are the neutral schwa, roughly, and every actual vowel departs from them substantially.
What the model gets right is the scaling, and it gets it right for a reason that does not depend on the shape. Any resonator scaled uniformly has its resonances scaled by the inverse of the factor, whatever its shape is. So the claim “a shorter tract multiplies every formant by the same factor” is exact for a uniformly scaled tract of any shape at all, and the uniform tube is doing nothing but supplying illustrative numbers.
That matters because tracts are not uniformly scaled between people. The pharynx grows more than the oral cavity from childhood to adulthood, and men’s pharynges are proportionally longer than women’s — so the real relation between speakers is a scaling plus a distortion, and the distortion is why a synthesised voice made by simply shifting formants sounds like a small adult rather than a child.
What this predicts that is checkable
Two consequences follow directly and both are testable without any equipment beyond a recording.
Resampling a recording should change the speaker and not the words. Playing a voice back at 1.2 times speed multiplies every frequency by 1.2 — fundamental and formants alike — which is the tract-length transformation with a pitch change bundled in. The result sounds like a smaller person saying the same thing, which is what everybody who has sped up a recording has heard, and the striking part is that the vowels are unharmed. A transformation that moves every formant by a fifth ought to destroy vowel identity if vowels were formant positions, and it does not.
Shifting the pitch alone should not change the speaker. Moving the fundamental without moving the formants — which is what a pitch-shifter that preserves formants does, and what a singer does across their range — leaves the speaker recognisable. Both of these are routine in audio production, and the fact that the two operations have such different perceptual consequences is the strongest everyday evidence that the source and the filter are separately read.
The failure case is the one that used to be audible in cheap pitch-shifting, where the formants were dragged along with the pitch and every shifted voice sounded like a cartoon. That artefact was the model being violated, and it is why the fix has a name.
Which computation produced the numbers
The resonances are the closed-open tube formula: the nth is (2n−1) times the speed of sound divided by four times the length, with the speed taken as 34,300 centimetres a second. The three tract lengths are the standard ranges quoted for children, adult women and adult men, and are the one input here that is not computed.
The semitone shifts are 12 log₂ of the length ratio, which is exact and needs no acoustics at all: if every frequency scales by the same factor, the shift in semitones is the same for every one of them.
The vowel formants drawn in the figures are the Peterson and Barney survey values, which are adult male, and the scaled versions are those multiplied through by the length ratio.
Whose voices, and where the numbers came from
The vowel data is a 1952 survey of American English, and it is used across this site because it is the measurement everything else was calibrated against. It is not a description of vowels in general — vowel inventories differ enormously between languages, and a formant plot of a language with front rounded vowels or with pharyngeals looks nothing like this one.
The identity argument is not about a language. Recognising a familiar voice is not a linguistic skill and is not confined to speech; it works on a laugh, a cough and a hum. The mechanism proposed here — that a scale factor on the envelope is the carrier — is the standard account in the speech-perception literature and is quoted rather than demonstrated by anything drawn here.
The evolutionary version of the same argument is worth naming because it is where the strongest evidence sits. Formant dispersion tracks body size in a great many mammals, several of which use it in vocal displays, and a few of which have evolved elaborate machinery for lying about it — a descended larynx lengthens the tube and makes the animal sound bigger than it is.
And the long-term average spectra of an orchestra and of a trained solo voice show what the same filter can be made to do when the problem is audibility rather than identity: a lowered larynx clusters the third, fourth and fifth formants near three kilohertz and puts a peak where the orchestra has none. That is the same tract, reshaped for a different job.
What the picture cannot show
It cannot show learning. Everything here is about what information is present in the signal. Recognising a particular familiar voice from that information is a memory task with a large learned component, and nothing in an envelope says how many exposures it took.
It cannot show the source’s own contribution. Two speakers with identical tracts would still be distinguishable, by the shape of their glottal pulse, by jitter and shimmer, by breathiness, and by habits of articulation and prosody that are not acoustic properties at all. The filter carries a great deal of identity; it does not carry all of it.
It cannot show the room. Everything drawn here assumes the envelope reaching the ear is the envelope leaving the mouth, and a room imposes a filter of its own that is often stronger than the differences between two speakers. Voices are still recognisable in rooms, which means the ear is discounting the room’s contribution — a separate problem with its own machinery.
And it stops working at the top of a voice. A soprano above about 700 hertz has a fundamental higher than her own first formant, so the partials are too far apart to sample the envelope. Vowels become hard to identify there, and so do speakers — which is a prediction this model makes and which every listener to a high soprano line has confirmed without meaning to. It is also why the singer’s formant is a low-register phenomenon: a resonance can only be exploited if there are partials near it, and above the first formant there are not enough partials left to exploit anything.
The ladder from here
This rung and the next were both named as owed when this anchor opened, and they are opposite questions about the same thing. This one asks what makes one voice identifiable. The next asks what happens when there are sixteen of them on one note, and the answer is that the identifiability is exactly what goes — a choir is not a loud singer, and the reason is that everything this essay describes as a carrier of identity becomes, at sixteen sources, a spread.
What is still owed after that is the acquisition side: a listener’s ability to normalise for a tract they have never heard before is present in infancy and is not obviously learnable from the data available, and nothing here has anything to say about how it gets there.
Part 5 of 13
One essay in the series on the voice. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
FormantResonanceSource-filterSpectral envelopeTimbreVocal tractVoice identity
- An instrument is not one timbre resonance, source-filter, spectral envelope, timbre
- Which instrument is underneath formant, source-filter, spectral envelope
- A family resemblance in the heights resonance, timbre
- A note takes a number of periods to speak resonance, timbre
- Above a certain note the holes stop working source-filter, timbre
- An instrument points source-filter, timbre