Instruments and their design

The other instrument with a reed

The folds do not vibrate the way a string does. They open and shut across a steady stream of air, once per period, and what leaves the larynx is a train of flow pulses with a closed phase in it. Everything said about the voice's tone is a statement about the shape of that pulse — and the shape has two numbers in it.

Assumes: The ear hears the list, not the shape · A vowel is two resonances

This collection has fifteen essays about instruments. There are pipes in it, and strings, and a hammer, a bow, a membrane, a bell, a set of tone holes and a loudspeaker. There is nothing at all about the instrument that made most of the music every other essay is about, and which every reader owns.

Part of the reason is that the voice appeared to have been dealt with. A vowel is two resonances takes the source–filter model apart and shows that vowel identity lives entirely in the filter — that the peaks of the mouth’s response stay where they are while the partials of the note slide underneath them. That essay is about the filter, and it is a rung of the spectrum ladder for exactly that reason.

Which leaves the source undescribed, and the commonest thing said about it is wrong.

What the folds are not doing

The usual description is that the vocal folds vibrate, and that the vibration produces a tone which the mouth then colours. The word invites a picture of a string doing everything at once: a stretched object displaced from rest, returning through it, overshooting, and radiating a set of partials that are the modes of its own shape.

Almost nothing in that picture applies. The folds are not radiating anything — they are buried in the neck, they are small, and at a hundred hertz they are far smaller than a wavelength, which is the condition under which a source radiates essentially nothing at all. Their own modes are not what the partials are. And they are not being displaced from rest by a driving force and returning through it.

What they are doing is opening and shutting, once per period, across a stream of air that the lungs are pushing steadily upward. The sound is the interruption. Air passes when the gap is open and does not when it is shut, so the volume flow leaving the larynx is a train of pulses with a flat stretch of nothing between them — and the whole of the voice’s source spectrum is the transform of that pulse.

The flow through the larynx, over two periods of a 110 Hz note. Volume flow against time, in Rosenberg's two-half-cosine model of the glottal pulse — a slow opening, a faster closing, and a closed phase during which no air passes at all. M1 — chest is open for 50 per cent of each period and opens 2.4 times as slowly as it closes. Nothing here is a displacement: the folds are a valve on a steady stream of air, and the flat stretches are the moments they are shut. At 110 Hz each period lasts 9.1 milliseconds, of which 4.5 is silence.
Fig. 1 The volume flow through the larynx over two periods of a note at 110 hertz, in Rosenberg’s model of the pulse: a slow opening, a faster closing, and a closed phase during which no air passes at all. The shaded stretches are the moments the folds are shut. Nothing in this drawing is a displacement — the vertical axis is a rate of flow — and the button plays what a spectrum of that shape sounds like with no mouth in front of it.

Two numbers describe the shape. The open quotient is the fraction of each period the folds are open at all, and the speed quotient is the ratio of the opening time to the closing time. That is the entire source model, and it is worth being precise about what follows from it: the source spectrum is not asserted anywhere. Two numbers go in, a shape comes out, and its harmonics are whatever the discrete transform of that shape says they are.

The twelve decibels nobody chose

The number every textbook quotes for the voice’s source is that it falls at about twelve decibels an octave. This site can produce it rather than repeat it.

The source spectrum of M1 — chest. The first 24 harmonics of the glottal flow pulse, in decibels below the strongest, taken as the discrete transform of the pulse shape itself rather than quoted. M1 — chest, at an open quotient of 0.50, puts its second partial 5.4 dB below its first and keeps 11 partials above 40 dB down. Above the second partial the slope is much the same in both — -13.2 decibels an octave — so what separates the mechanisms is not the slope but how much of the sound is in the fundamental.
Fig. 2 The first twenty-four harmonics of that pulse, in decibels below the strongest, taken as the transform of the drawing above. The slope above the second partial is a measured property of the shape and not a parameter of the model: nothing was set to twelve, and the fit comes out within a decibel of it.

The reason the slope exists is the closure. A waveform with a corner in it has partials falling as one over the harmonic number squared, which is twelve decibels an octave; a waveform with no corner — a pure sinusoid, say — has one partial and nothing else.

Two things about that are worth measuring rather than asserting, and the second corrects the sentence this section was about to end with. The computed slope at the default shape is −13.4 decibels an octave, which is a decibel and a half from the published twelve rather than within one of it — the model overshoots slightly, and saying so costs nothing.

And the slope barely moves with either parameter. Sweeping the open quotient from 0.3 to 0.9 takes it from −13.7 to −12.3; sweeping the speed quotient from 1 to 6 takes it from −13.4 to −12.3. Across the whole plausible range of both, the slope varies by about one decibel an octave, which is less than the model’s own offset from the published figure.

So the sharpness of the shutting is not what a bright voice and a breathy one differ in, at least not as a slope. Where they differ is in the level:

open quotient partial 4 partial 8
0.3, firmly closed −10 dB −25 dB
0.6, the default −23 −38
0.9, barely closing −31 −42

Twenty-one decibels at the fourth partial and seventeen at the eighth, between a firmly closing fold and one that barely shuts — an enormous difference in what reaches the mouth, produced by a shape whose slope is nearly constant. A breathy voice does not have a steeper spectrum than a bright one; it has the same spectrum shifted down, and the shift is in how much of the series exists rather than how fast it falls away.

The speed quotient runs the other way and by less: taking the opening-to-closing ratio from 1 to 6 raises the eighth partial from −43 decibels to −27, so a fold that shuts much faster than it opens is brighter by sixteen decibels up there. Both parameters therefore move the same quantity — the level of the upper series — and neither moves the slope. Which is a tidier account than the geometric one: the corner’s sharpness decides how much energy is in the discontinuity, and the one-over-n-squared falloff is a property of there being a corner at all rather than of how sharp it is.

This is the same reasoning that gives a bowed string its sawtooth — the Helmholtz corner running round the string is a discontinuity, and a discontinuity is a full harmonic series with a one-over-n envelope. The voice’s corner is the moment of closure, once a cycle, and it produces the same kind of spectrum for the same reason.

Two mechanisms, the notes both of them make, and the seam. The frequency range of each laryngeal mechanism for an adult male voice, on a logarithmic axis, with the band both can produce shaded. M1 — chest runs 82–349 Hz and M2 — falsetto runs 220–698 Hz, so 799 cents of the range — 8.0 semitones — can be sung either way. The two dots inside that band are the measured signature that this is a bifurcation rather than a threshold: the change upward happens at 330 Hz and the change downward at 294 Hz, 200 cents lower. A threshold is crossed at the same place in both directions and this is not.
Fig. 3 The other thing the corner’s sharpness cannot explain, put here because it is what a second mechanism looks like. The two laryngeal mechanisms overlap by 799 cents — eight semitones an adult male voice can sing either way — and inside that band the same written note is two different sources with two different closed phases and two different spectra. The corner argument covers everything within one mechanism: the tube decides which partials exist, the sharpness of the closure decides how strong they are, and the one-over-n-squared falloff is a property of there being a corner at all. What it does not cover is which of the two corners is being made.

The valve

The physics that make the folds open and shut are not the physics of an oscillator being driven. They are the physics of a valve, and the site has already drawn one.

A reed that shuts at 5000 pascals, and the air it lets through. Volume flow through the reed channel against the pressure across it, in the quasi-static model: Bernoulli flow through an opening that closes linearly with pressure. The flow peaks at 1667 pascals — exactly a third of the closing pressure, for any reed, because that is where the two effects balance — and it is 0.18 litres a second there. Everything to the right of that peak is the argument: the flow falls as the player blows harder, from 0.18 to 0.05 litres a second by 4500 pascals, and a resistance that behaves that way supplies energy instead of taking it. There is no reed inertia in this model, so it cannot squeak.
Fig. 4 A clarinet reed, drawn as what it is: flow through the opening against the pressure across it. Blowing harder drives more air through the slot and also pushes the reed towards closing, which narrows it. Up to a third of the closing pressure the first effect wins; past that the second does, and the flow falls as the pressure rises. That descending stretch is the whole mechanism by which a wind instrument sustains a note.

A larynx is that valve with a different geometry. Air arrives from below at some pressure; the gap between the folds is a constriction; the flow through the constriction depends on the pressure across it and on how wide the gap is; and the width of the gap depends, through the Bernoulli pressure drop inside it, on how much air is going through. That coupling is what makes the cycle self-sustaining rather than something a muscle has to perform six times a second, and it is the reason a singer cannot make the folds open and close by intending it.

Two consequences follow immediately, and they are both audible.

A note starts rather than fades in. Below a threshold pressure the coupling does not close the loop, the folds do not oscillate, and there is no sound at all — not a quiet sound, none. Above it there is a note. The pressure at which this happens is the phonation threshold, and it is the same kind of quantity as the pressure at which a clarinet speaks.

And loudness is not a control anybody has directly. Radiated level rises about nine decibels for every doubling of subglottal pressure — an exponent of one and a half on the pressure, not of one — so the useful dynamic range costs a great deal of pressure at the top and very little at the bottom.

Loudness is not the control a singer has. Radiated sound pressure level against subglottal pressure. Below 300 pascals the folds do not oscillate at all and there is no sound, which is why a sung note starts rather than fades in; above it the level rises about 8.7 decibels for every doubling of pressure — 63 dB at 500 Pa, 75 dB at 800 Pa, 87 dB at 1600 Pa, 93 dB at 2400 Pa. Doubling the pressure is not doubling the loudness and is not doubling anything a lung does easily: the range drawn here is 1900 pascals for 30 decibels.
Fig. 5 Radiated level against subglottal pressure, with the region below the phonation threshold shaded. The curve is steep at the bottom and flattening at the top, which is why the quiet end of a singer’s range is the hard end to control and why the loud end costs breath out of proportion to what it buys.

Where the analogy with the reed stops, and it stops sharply

A clarinet reed does not choose the note. The tube does: the reed is a valve with a natural frequency far above anything the instrument plays, and what fixes the pitch is the standing wave in the bore, which the valve is merely feeding. Change the tube and the note changes; change the reed and the note does not.

The voice is the other way round, and the difference is a matter of numbers rather than of principle.

What a 17 cm tube supports, by how its ends are closedThe first 4 modes of a stopped cylinder, all of the same acoustic length. A cylinder stopped at one end supports only the odd multiples and reaches its second mode 1902 cents up, which is a twelfth. Its fundamental is an octave below that of an open tube of the same length, because it fits a quarter of a wavelength where an open tube fits a half.stopped cylinder143 Hz fundamental1434291902 cents — a twelfth220440880hertz, on a logarithmic axisevery mode the tube supports, and the jump from the first to the second
Fig. 6 The vocal tract as the tube it is: about seventeen centimetres from the folds to the lips, effectively closed at the fold end and open at the lips. Its first three resonances come out near 500, 1,500 and 2,500 hertz — which are the formants of a neutral vowel, arrived at from a length and a wave speed with nothing about speech in the calculation.

Those resonances are the ones the vowel essay draws as a filter curve. For an ordinary sung note at a hundred or two hundred hertz, the lowest of them is three to five times the fundamental — far away, in the sense that the tract’s resonance is not near enough to the fold oscillation to influence when the folds close. So the source runs at its own frequency, the filter shapes what comes out, and the two are independent. That independence is what makes the source–filter model work and it is an approximation with a stated condition: the resonances have to be well above the fundamental.

The condition fails in one place, which is where the next rung of this ladder starts. A soprano at the top of her range is singing above five hundred hertz, the first tract resonance is below the fundamental, and the two systems are coupled. What happens then is not a subtlety of tone; it is that the vowel becomes unrecoverable and the singer has to reshape the tract to get a note out at all.

What the closed phase is worth

The flat stretch in the flow pulse is the part of the model that does the most work and the part that is easiest to overlook, so it is worth putting a number on it.

At a hundred and ten hertz the period is 9.1 milliseconds. With the folds open half the time, they are shut for about four and a half milliseconds out of every nine — which is to say that for half of the duration of a sung note, no air is leaving the larynx at all.

That matters for three separate reasons.

It is why the spectrum is rich. A shape that spends part of every period at exactly zero cannot be a sinusoid, and the further it is from one the more partials it has. A pulse with no closed phase at all — the folds never quite meeting — is close to a sinusoid, and its second partial is more than twenty decibels below its first.

It is why the source spectrum has anything to do with effort. Louder singing closes the folds harder and faster, which sharpens the corner, which raises the high partials by more than it raises the low ones. So the same note played louder is not the same note amplified — a fact this site has already measured on the roughness side, arriving from the ear rather than from the larynx.

And it is why the voice is efficient at all. The lungs supply a steady flow and the larynx converts part of it into acoustic power. A valve that is shut half the time is passing half the air for the same sound, which is a large fraction of why a trained singer can hold a phrase.

The voice is also a strange object in one further respect. A struck or plucked note has an envelope that is most of what identifies it and no steady state worth speaking of; a sung note has a steady state that can be held for twenty seconds, and the source parameters can be moved continuously while it is being held. Nothing else here can do that except a bowed string, and the bow has two bounds it must stay between while the larynx has one.

Against the other instruments this collection has measured, the voice’s own spectrum is a full series falling steadily — nearer a bowed string’s than a clarinet’s, since the clarinet is missing its even partials because of where the ends of its tube are rather than because of anything its reed does. And the comparison worth making is not between waveshapes, which the ear has almost no access to, but between the presence and absence of a corner: a waveform with one carries a full series falling as one over n, and a waveform without one carries almost nothing above its fundamental. The glottal pulse has a corner because the folds slam shut, and that single fact fixes the shape of everything above the first partial.

The same argument, run backwards

There is a check on all of this that costs nothing, and it is the check the site’s habit demands: the model has to be able to fail.

Take the pulse’s two parameters to the values that describe a voice with no firm closure — open most of the period, opening and closing at nearly the same rate — and the model predicts a spectrum with almost everything in the fundamental and very little above it. That is a specific, falsifiable claim about breathy phonation, and it is right: the acoustic measure phoneticians use for exactly this is the level of the first partial above the second, and it separates pressed from breathy voice quality with no reference to any of the other things a voice does.

The source spectrum of M1 — chest and M2 — falsetto. The first 20 harmonics of the glottal flow pulse, in decibels below the strongest, taken as the discrete transform of the pulse shape itself rather than quoted. M1 — chest, at an open quotient of 0.50, puts its second partial 5.4 dB below its first and keeps 11 partials above 40 dB down; M2 — falsetto, at an open quotient of 0.80, puts its second partial 22.4 dB below its first and keeps 6 partials above 40 dB down. Above the second partial the slope is much the same in both — -13.3 and -12.1 decibels an octave — so what separates the mechanisms is not the slope but how much of the sound is in the fundamental.
Fig. 7 The same transform run on two pulse shapes. The difference between them is not the slope above the second partial, which is much the same in both; it is how much of the sound is in the fundamental. That quantity has a name and a measurement procedure in phonetics, and it is the number that comes out of this drawing.

The failure mode is available too. If the source’s spectrum were a property of the folds’ mass and tension rather than of the pulse’s shape, then two voices at the same pitch would have the same source spectrum, since they would be running the same oscillator at the same rate. They do not, and the differences track exactly the things that change the pulse shape — how firmly the folds are adducted, how much air is flowing, how hard the closure is.

Whose voices these numbers are

Every value on this page is an adult modal voice, and mostly an adult male one.

The tract length of about seventeen centimetres is a male mean; adult female tracts are nearer fourteen and a half, which puts the neutral-vowel resonances about fifteen per cent higher, and a small child’s are shorter again. The formant values the vowel essay uses are Peterson and Barney’s 1952 means for adult male speakers, and the site says so there.

The pulse model is Rosenberg’s, which is a fit to inverse-filtered speech rather than a derivation. A more elaborate model — Fant, Liljencrants and Lin’s, which is the standard one — adds a return phase after closure, because real folds do not shut instantaneously, and that return phase is precisely what controls the highest partials. Nothing in this essay’s argument depends on which is used; the closed phase and the corner are in both.

And the whole of it is phonation, not singing. A sung note is a phonation plus a tract shape plus a vibrato plus an onset, and the last three are the next three rungs.

What the picture cannot show

There is no noise in this model, and a real voice has some. Air rushing through a narrow gap is turbulent, and turbulence is broadband. In a breathy voice the noise is a substantial part of the sound; in a pressed one it is a small part; in a whisper it is all of it, which is why a whisper has vowels and no pitch. The transform of a smooth pulse cannot produce any of that.

The source and the filter are not quite independent. The model treats the tract as a passive filter hanging off a source that does not know it is there. In fact the tract’s input impedance is part of what loads the folds, which is why a singer’s tone changes when the mouth opens, and the effect is largest exactly where the resonances come near the fundamental.

And the folds are not one object. The two-mass and body–cover models that actually reproduce self-sustained oscillation treat the fold as a stiff body with a loose mucosal cover that travels over it in a wave, and it is that wave — not any single stiffness — that supplies the phase difference between opening and closing. The pulse drawn here is the output of such a system, taken as given.

Whose instrument, and when

The larynx was described as a reed instrument long before any of this could be measured. Ferrein’s dissection experiments in 1741 blew air through excised larynxes and called the folds cordes vocales — vocal cords, which is the string metaphor, still in the language — and simultaneously demonstrated that they behaved like a reed, which is the valve metaphor. Both pictures have been available for nearly three centuries and the wrong one won the vocabulary.

Müller’s mid-nineteenth-century work on excised larynxes established the pressure dependence; van den Berg’s in the 1950s named the mechanism the myoelastic–aerodynamic theory, which is the valve description with the Bernoulli term made explicit; and the inverse filtering that produced the pulse shapes drawn here dates from the 1950s and 1960s.

What none of that history changes is the thing this rung exists to say. The voice is a valve on a stream of air, its spectrum is the transform of a pulse with a closed phase in it, and the two numbers that describe the pulse are enough to say why one voice is bright and another breathy — before any question about the mouth, the vowel or the note has been asked.

The ladder from here

Two mechanisms are available to the folds rather than one, and the seam between them is not a threshold. The next rung measures the overlap between them and the hysteresis at the crossing, which is the signature of a bifurcation and the thing no account of technique predicts.

Part 1 of 13

One essay in the series on the voice. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Boundary conditionGlottisOpen quotientPartialReedSource-filterSpectrumThreshold pressure