The other instrument with a reed
Assumes: The ear hears the list, not the shape · A vowel is two resonances
This collection has fifteen essays about instruments. There are pipes in it, and strings, and a hammer, a bow, a membrane, a bell, a set of tone holes and a loudspeaker. There is nothing at all about the instrument that made most of the music every other essay is about, and which every reader owns.
Part of the reason is that the voice appeared to have been dealt with. A vowel is two resonances takes the source–filter model apart and shows that vowel identity lives entirely in the filter — that the peaks of the mouth’s response stay where they are while the partials of the note slide underneath them. That essay is about the filter, and it is a rung of the spectrum ladder for exactly that reason.
Which leaves the source undescribed, and the commonest thing said about it is wrong.
What the folds are not doing
The usual description is that the vocal folds vibrate, and that the vibration produces a tone which the mouth then colours. The word invites a picture of a string doing everything at once: a stretched object displaced from rest, returning through it, overshooting, and radiating a set of partials that are the modes of its own shape.
Almost nothing in that picture applies. The folds are not radiating anything — they are buried in the neck, they are small, and at a hundred hertz they are far smaller than a wavelength, which is the condition under which a source radiates essentially nothing at all. Their own modes are not what the partials are. And they are not being displaced from rest by a driving force and returning through it.
What they are doing is opening and shutting, once per period, across a stream of air that the lungs are pushing steadily upward. The sound is the interruption. Air passes when the gap is open and does not when it is shut, so the volume flow leaving the larynx is a train of pulses with a flat stretch of nothing between them — and the whole of the voice’s source spectrum is the transform of that pulse.
Two numbers describe the shape. The open quotient is the fraction of each period the folds are open at all, and the speed quotient is the ratio of the opening time to the closing time. That is the entire source model, and it is worth being precise about what follows from it: the source spectrum is not asserted anywhere. Two numbers go in, a shape comes out, and its harmonics are whatever the discrete transform of that shape says they are.
The twelve decibels nobody chose
The number every textbook quotes for the voice’s source is that it falls at about twelve decibels an octave. This site can produce it rather than repeat it.
The reason the slope exists is the closure. A waveform with a corner in it has partials falling as one over the harmonic number squared, which is twelve decibels an octave; a waveform with no corner — a pure sinusoid, say — has one partial and nothing else.
Two things about that are worth measuring rather than asserting, and the second corrects the sentence this section was about to end with. The computed slope at the default shape is −13.4 decibels an octave, which is a decibel and a half from the published twelve rather than within one of it — the model overshoots slightly, and saying so costs nothing.
And the slope barely moves with either parameter. Sweeping the open quotient from 0.3 to 0.9 takes it from −13.7 to −12.3; sweeping the speed quotient from 1 to 6 takes it from −13.4 to −12.3. Across the whole plausible range of both, the slope varies by about one decibel an octave, which is less than the model’s own offset from the published figure.
So the sharpness of the shutting is not what a bright voice and a breathy one differ in, at least not as a slope. Where they differ is in the level:
| open quotient | partial 4 | partial 8 |
|---|---|---|
| 0.3, firmly closed | −10 dB | −25 dB |
| 0.6, the default | −23 | −38 |
| 0.9, barely closing | −31 | −42 |
Twenty-one decibels at the fourth partial and seventeen at the eighth, between a firmly closing fold and one that barely shuts — an enormous difference in what reaches the mouth, produced by a shape whose slope is nearly constant. A breathy voice does not have a steeper spectrum than a bright one; it has the same spectrum shifted down, and the shift is in how much of the series exists rather than how fast it falls away.
The speed quotient runs the other way and by less: taking the opening-to-closing ratio from 1 to 6 raises the eighth partial from −43 decibels to −27, so a fold that shuts much faster than it opens is brighter by sixteen decibels up there. Both parameters therefore move the same quantity — the level of the upper series — and neither moves the slope. Which is a tidier account than the geometric one: the corner’s sharpness decides how much energy is in the discontinuity, and the one-over-n-squared falloff is a property of there being a corner at all rather than of how sharp it is.
This is the same reasoning that gives a bowed string its sawtooth — the Helmholtz corner running round the string is a discontinuity, and a discontinuity is a full harmonic series with a one-over-n envelope. The voice’s corner is the moment of closure, once a cycle, and it produces the same kind of spectrum for the same reason.
The valve
The physics that make the folds open and shut are not the physics of an oscillator being driven. They are the physics of a valve, and the site has already drawn one.
A larynx is that valve with a different geometry. Air arrives from below at some pressure; the gap between the folds is a constriction; the flow through the constriction depends on the pressure across it and on how wide the gap is; and the width of the gap depends, through the Bernoulli pressure drop inside it, on how much air is going through. That coupling is what makes the cycle self-sustaining rather than something a muscle has to perform six times a second, and it is the reason a singer cannot make the folds open and close by intending it.
Two consequences follow immediately, and they are both audible.
A note starts rather than fades in. Below a threshold pressure the coupling does not close the loop, the folds do not oscillate, and there is no sound at all — not a quiet sound, none. Above it there is a note. The pressure at which this happens is the phonation threshold, and it is the same kind of quantity as the pressure at which a clarinet speaks.
And loudness is not a control anybody has directly. Radiated level rises about nine decibels for every doubling of subglottal pressure — an exponent of one and a half on the pressure, not of one — so the useful dynamic range costs a great deal of pressure at the top and very little at the bottom.
Where the analogy with the reed stops, and it stops sharply
A clarinet reed does not choose the note. The tube does: the reed is a valve with a natural frequency far above anything the instrument plays, and what fixes the pitch is the standing wave in the bore, which the valve is merely feeding. Change the tube and the note changes; change the reed and the note does not.
The voice is the other way round, and the difference is a matter of numbers rather than of principle.
Those resonances are the ones the vowel essay draws as a filter curve. For an ordinary sung note at a hundred or two hundred hertz, the lowest of them is three to five times the fundamental — far away, in the sense that the tract’s resonance is not near enough to the fold oscillation to influence when the folds close. So the source runs at its own frequency, the filter shapes what comes out, and the two are independent. That independence is what makes the source–filter model work and it is an approximation with a stated condition: the resonances have to be well above the fundamental.
The condition fails in one place, which is where the next rung of this ladder starts. A soprano at the top of her range is singing above five hundred hertz, the first tract resonance is below the fundamental, and the two systems are coupled. What happens then is not a subtlety of tone; it is that the vowel becomes unrecoverable and the singer has to reshape the tract to get a note out at all.
What the closed phase is worth
The flat stretch in the flow pulse is the part of the model that does the most work and the part that is easiest to overlook, so it is worth putting a number on it.
At a hundred and ten hertz the period is 9.1 milliseconds. With the folds open half the time, they are shut for about four and a half milliseconds out of every nine — which is to say that for half of the duration of a sung note, no air is leaving the larynx at all.
That matters for three separate reasons.
It is why the spectrum is rich. A shape that spends part of every period at exactly zero cannot be a sinusoid, and the further it is from one the more partials it has. A pulse with no closed phase at all — the folds never quite meeting — is close to a sinusoid, and its second partial is more than twenty decibels below its first.
It is why the source spectrum has anything to do with effort. Louder singing closes the folds harder and faster, which sharpens the corner, which raises the high partials by more than it raises the low ones. So the same note played louder is not the same note amplified — a fact this site has already measured on the roughness side, arriving from the ear rather than from the larynx.
And it is why the voice is efficient at all. The lungs supply a steady flow and the larynx converts part of it into acoustic power. A valve that is shut half the time is passing half the air for the same sound, which is a large fraction of why a trained singer can hold a phrase.
The voice is also a strange object in one further respect. A struck or plucked note has an envelope that is most of what identifies it and no steady state worth speaking of; a sung note has a steady state that can be held for twenty seconds, and the source parameters can be moved continuously while it is being held. Nothing else here can do that except a bowed string, and the bow has two bounds it must stay between while the larynx has one.
Against the other instruments this collection has measured, the voice’s own spectrum is a full series falling steadily — nearer a bowed string’s than a clarinet’s, since the clarinet is missing its even partials because of where the ends of its tube are rather than because of anything its reed does. And the comparison worth making is not between waveshapes, which the ear has almost no access to, but between the presence and absence of a corner: a waveform with one carries a full series falling as one over n, and a waveform without one carries almost nothing above its fundamental. The glottal pulse has a corner because the folds slam shut, and that single fact fixes the shape of everything above the first partial.
The same argument, run backwards
There is a check on all of this that costs nothing, and it is the check the site’s habit demands: the model has to be able to fail.
Take the pulse’s two parameters to the values that describe a voice with no firm closure — open most of the period, opening and closing at nearly the same rate — and the model predicts a spectrum with almost everything in the fundamental and very little above it. That is a specific, falsifiable claim about breathy phonation, and it is right: the acoustic measure phoneticians use for exactly this is the level of the first partial above the second, and it separates pressed from breathy voice quality with no reference to any of the other things a voice does.
The failure mode is available too. If the source’s spectrum were a property of the folds’ mass and tension rather than of the pulse’s shape, then two voices at the same pitch would have the same source spectrum, since they would be running the same oscillator at the same rate. They do not, and the differences track exactly the things that change the pulse shape — how firmly the folds are adducted, how much air is flowing, how hard the closure is.
Whose voices these numbers are
Every value on this page is an adult modal voice, and mostly an adult male one.
The tract length of about seventeen centimetres is a male mean; adult female tracts are nearer fourteen and a half, which puts the neutral-vowel resonances about fifteen per cent higher, and a small child’s are shorter again. The formant values the vowel essay uses are Peterson and Barney’s 1952 means for adult male speakers, and the site says so there.
The pulse model is Rosenberg’s, which is a fit to inverse-filtered speech rather than a derivation. A more elaborate model — Fant, Liljencrants and Lin’s, which is the standard one — adds a return phase after closure, because real folds do not shut instantaneously, and that return phase is precisely what controls the highest partials. Nothing in this essay’s argument depends on which is used; the closed phase and the corner are in both.
And the whole of it is phonation, not singing. A sung note is a phonation plus a tract shape plus a vibrato plus an onset, and the last three are the next three rungs.
What the picture cannot show
There is no noise in this model, and a real voice has some. Air rushing through a narrow gap is turbulent, and turbulence is broadband. In a breathy voice the noise is a substantial part of the sound; in a pressed one it is a small part; in a whisper it is all of it, which is why a whisper has vowels and no pitch. The transform of a smooth pulse cannot produce any of that.
The source and the filter are not quite independent. The model treats the tract as a passive filter hanging off a source that does not know it is there. In fact the tract’s input impedance is part of what loads the folds, which is why a singer’s tone changes when the mouth opens, and the effect is largest exactly where the resonances come near the fundamental.
And the folds are not one object. The two-mass and body–cover models that actually reproduce self-sustained oscillation treat the fold as a stiff body with a loose mucosal cover that travels over it in a wave, and it is that wave — not any single stiffness — that supplies the phase difference between opening and closing. The pulse drawn here is the output of such a system, taken as given.
Whose instrument, and when
The larynx was described as a reed instrument long before any of this could be measured. Ferrein’s dissection experiments in 1741 blew air through excised larynxes and called the folds cordes vocales — vocal cords, which is the string metaphor, still in the language — and simultaneously demonstrated that they behaved like a reed, which is the valve metaphor. Both pictures have been available for nearly three centuries and the wrong one won the vocabulary.
Müller’s mid-nineteenth-century work on excised larynxes established the pressure dependence; van den Berg’s in the 1950s named the mechanism the myoelastic–aerodynamic theory, which is the valve description with the Bernoulli term made explicit; and the inverse filtering that produced the pulse shapes drawn here dates from the 1950s and 1960s.
What none of that history changes is the thing this rung exists to say. The voice is a valve on a stream of air, its spectrum is the transform of a pulse with a closed phase in it, and the two numbers that describe the pulse are enough to say why one voice is bright and another breathy — before any question about the mouth, the vowel or the note has been asked.
The ladder from here
Two mechanisms are available to the folds rather than one, and the seam between them is not a threshold. The next rung measures the overlap between them and the hysteresis at the crossing, which is the signature of a bifurcation and the thing no account of technique predicts.
Part 1 of 13
One essay in the series on the voice. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Boundary conditionGlottisOpen quotientPartialReedSource-filterSpectrumThreshold pressure
- An instrument points partial, source-filter, spectrum
- The body is the filter partial, source-filter, spectrum
- The hammer is not a point either boundary condition, partial, spectrum
- A bar's partials are the odd numbers, squared boundary condition, partial
- A beat has a depth, and six essays held it at one partial, spectrum
- A clarinet keeps what a string loses partial, spectrum