One voice over ninety players
Assumes: The other instrument with a reed · A vowel is two resonances
An opera house holds two thousand people. In the pit there are between sixty and a hundred players, and on the stage there is one person, unamplified, who is expected to be heard over all of them for four hours. The arrangement has worked for two hundred years and it is worth asking what it is that makes it work, because the obvious answer is wrong.
The obvious answer is that the singer is louder. A trained voice at full effort produces something like a hundred decibels at a metre; a large orchestra at full effort produces rather more than that, from a hundred sources, and the audience is much closer to the pit than to the stage. On any straightforward accounting of power the voice loses.
Level is the wrong quantity
The right quantity is not how much energy each party produces but where in the spectrum each puts it, because the ear does not hear a total. It hears in bands — each place on the cochlea responds to a range of frequencies rather than to one — and a sound is audible in the presence of another if it stands above it somewhere, not on average.
So the comparison to make is of long-term average spectra: each source’s mean spectrum over a long passage, drawn against its own peak so that shape is what is being compared rather than level.
The orchestra’s shape is what a room full of strings and low woodwind produces: a peak at a few hundred hertz and a steady decline above it, thirty decibels down by three kilohertz. That decline is not an accident of any particular orchestration. It is what the sources are — bowed strings with a body response that rolls off, and pipes and reeds whose partials fall as one over n or faster.
The voice’s shape does the same thing until about a kilohertz and then does something else.
Nineteen decibels, and where
Subtracting one curve from the other turns the description into a number.
Nineteen decibels is a large number. It is not the voice being nineteen decibels louder; it is the voice being nineteen decibels above the orchestra in that band, having been ten or fifteen decibels below it two octaves lower. The strategy is not amplification. It is relocation.
The mechanism is a resonance, and it is one the singer creates by a postural change rather than by effort. Lowering the larynx widens the pharynx just above it, and when the pharynx is several times wider than the tube at the top of the larynx, that short tube stops being part of the main resonator and starts behaving as a separate one with a resonance of its own.
The tube doing it is short: about three centimetres of laryngeal cavity, lowered clear of the pharynx above it, which behaves as a resonator in its own right and puts its lowest mode near three kilohertz. That is the whole mechanism — not a louder voice but a fifth formant clustered with the third and fourth, in a band the orchestra does not occupy.
That clustering is Sundberg’s account and it has a testable consequence which holds: the singer’s formant is present in trained male operatic voices, weak or absent in untrained ones, and weak in sopranos — who, as the next rung shows, have a different problem and a different solution.
The ear was already there
Here is the part that is not designed and is the more interesting for it.
So the frequency at which the voice stands furthest clear of the orchestra is within a hundred and fifty cents of the frequency at which the ear is most sensitive. Sweeping both curves rather than reading their labelled peaks, the two land on the same point — the widest gap and the lowest threshold are both at 3,151 hertz — though that is the resolution of two tabulated datasets meeting rather than a measurement of how close they really are, and the hundred and fifty cents is the honest bound. Neither party arranged that. The ear canal’s length is a fact about anatomy that long predates opera; the singer’s formant is a configuration of the vocal tract that a nineteenth-century training tradition arrived at by ear.
What connects them is not coincidence but selection. A technique that put its peak anywhere else would be less audible, would be less used, and would not have become the technique. The ear’s sensitivity is the fitness landscape and the singer’s formant is what grew on it.
Masking, which is the mechanism spelled out
“Audible over” is loose. The precise statement is about masking, and this collection has measured how masking spreads: a loud low tone raises the threshold for tones above it far more than for tones below it, and the spread upward grows with level.
That looks at first like bad news for a voice sitting above an orchestra whose energy is all low down. It is not, and the reason is not the obvious one about the skirt running out of reach.
The consequence is that the voice’s own low partials are thoroughly masked by the orchestra and its high ones are not. A listener at the back of an opera house is very largely hearing the top of the voice — which is the reason a recording made close to a singer sounds unlike the experience of the hall, and the reason recorded opera required a different balance from the beginning.
The skirt does reach, and it does not matter
Two things about that are worth computing rather than asserting, because the first is wrong and the second is the reason the first does not cost anything.
The skirt reaches. The upward spread of masking gets shallower as the masker gets louder — that is what the standard spreading function does — so the question is not whether a low masker can reach three kilohertz but at what level. Solving it: a 400-hertz masker’s skirt first clears the quiet threshold at 3,150 hertz when the masker is at 85 decibels. An orchestral tutti at a listener’s seat is louder than that, and in the pit considerably louder. So a singer’s formant is not competing with the room; at any level where the question arises, it is inside the skirt.
And the skirt is irrelevant anyway, because it is not the largest thing at that frequency. The orchestra’s own long-term spectrum is thirty decibels down at three kilohertz, which is a great deal more energy than its bass leaks up there by masking:
| orchestra at | its own 3.15 kHz content | the skirt from its 400 Hz energy |
|---|---|---|
| 80 dB | 50 dB | −23 |
| 90 dB | 60 | 11 |
| 100 dB | 70 | 45 |
The skirt runs forty to seventy decibels below the orchestra’s direct content in the same band, at every level. What a voice at three kilohertz has to stand above is the orchestra’s real energy up there, not the shadow of its bass — and standing above the orchestra’s real energy in that band is exactly what the nineteen-decibel figure measures.
So the masking section is answering a question the long-term spectra had already answered, and answering it with a mechanism that does not behave as described. That is a tidier result than it looks: the argument needs no masking model at all. Two spectra and a subtraction are the whole of it, and the reason they suffice is that the orchestra’s own roll-off is a much weaker filter than its masking skirt — it leaks a hundred times more energy into the singer’s band directly than it does by spreading.
Where the energy comes from
None of this creates power, so it has to come out of somewhere, and it does.
Two things follow. The first is that the effect depends on the source having partials that high at all, which is why it is a property of a firmly closing M1 production and largely absent from the breathy configuration the previous rung measured, whose second partial is already twenty decibels down.
The second is that a lowered larynx costs something elsewhere. It lengthens the tract, which lowers the first two formants, which is the reason a trained operatic vowel is darker than a spoken one — and, in the Italian pedagogical vocabulary, the reason the instruction is chiaroscuro: light and dark at once, meaning the singer’s formant and the lowered lower formants together.
And the mouth is beginning to point
One more consequence follows from the frequency alone, without any reference to voices.
A source two and a half centimetres across radiates in all directions at a few hundred hertz and begins to beam by three kilohertz, and the singer’s formant sits on the rising part of that curve — so the advantage is largest for a listener the singer is facing and falls off to the sides. That is a real and measurable part of stagecraft, and it is why a singer turning upstage loses more than a proportional share of audibility: the component being lost is the one that was doing the work.
That is a real and measurable part of stagecraft, and it is why a singer turning upstage loses more than a proportional share of audibility: the component being lost is the one that was doing the work.
The hall is not neutral about this either
The comparison so far is between two sources. A listener in the fourth circle hears both of them through a room, and the room treats three kilohertz differently from three hundred.
That works in the singer’s favour twice. The reverberant field takes over beyond a critical distance, and the critical distance is larger where the reverberation time is shorter — so the direct sound reaches further at three kilohertz than at three hundred. And the first eighty milliseconds are counted as part of the direct sound rather than against it, so the early reflections in that band add to the voice instead of blurring it.
It also works against the orchestra, whose low energy is exactly what a long bass reverberation time smears into a continuous bed. A hall that flatters a voice is a hall whose bass tail is long and whose treble is not, which is a description most nineteenth-century opera houses would recognise.
The test this account could fail
The claim is that carrying power lives in a band rather than in a level, and it is falsifiable in a way that costs a filter.
Remove the two-to-four-kilohertz band from a recording of a voice with an orchestra. The overall level barely moves — there is very little total energy up there, which is the whole reason the region was available — and the account predicts that the voice should become substantially harder to follow. That is the experiment, and it has been run: filtered that way, trained voices lose the quality listeners describe as carrying, while their measured level changes by a decibel or two.
And the converse should hold. A voice with no peak in that band should be hard to hear over an orchestra even when it is loud, which is the ordinary experience of an untrained loud singer in front of an ensemble, and is why the instruction to sing louder does not work.
There is a failure condition. If the peak were merely a by-product of loud singing rather than a configuration, it would appear in any voice pushed hard enough, and it does not: the measurements find it tracks larynx height rather than effort, and a trained singer can produce it quietly and can suppress it at full volume.
A tradition that does not have this resource
Operatic technique is one solution to one problem, and naming the problem exactly is what makes it possible to see the others.
A soprano above about seven hundred hertz cannot use it. Her partials are so widely spaced that there may be nothing at all near three kilohertz to amplify, and her first formant is a bigger problem than her third. What she does instead is the subject of a large literature on formant tuning, and it is a different manoeuvre with a different cost.
Belt technique does the opposite. Raising the first formant to sit on the second partial of a chest-register note puts a very large amount of energy around eight hundred to a thousand hertz — inside the orchestra’s own range rather than above it — and works because the amplified partial is a strong one and because musical-theatre orchestras are smaller and are usually amplified too.
Chinese opera places a great deal of energy above two kilohertz by a route that is not a lowered larynx at all: a high, narrow, often nasalised configuration, over an ensemble whose own spectrum is very different from a Western orchestra’s — percussion, plucked strings and the suona, which has plenty of energy up where a European orchestra has none. The problem is not the same problem, so the solution is not the same solution.
And the loudest thing in a European orchestra is the trumpet, whose spectrum runs well past three kilohertz because its bell radiates high frequencies efficiently and low ones badly. A voice does not out-carry a trumpet and never has. What the orchestration does about that is to not write both at once, which is a compositional fact rather than an acoustic one.
That exception is the one that shows the argument is about a shape rather than about a band. The singer’s formant works because the orchestra’s mean spectrum has a hole where the voice has a peak, and a trumpet fills the hole — a bell’s whole function is to radiate the high partials the tube would otherwise reflect, so a brass section is the one part of an orchestra whose spectrum does not roll off where the voice needs it to. Nineteen decibels of advantage over a mean that includes brass is not nineteen decibels of advantage over brass.
What the picture cannot show
A long-term average spectrum is an average. It is what a long passage does on the whole, and the moment a singer is actually fighting for is not the average moment — it is a tutti on a low chord under a high phrase. Both curves would move, and the site does not have the instantaneous data.
The orchestral spectrum depends entirely on the orchestration. Sundberg’s mean is over a particular repertoire. A Wagner orchestra with eight horns and a bass clarinet is a different spectrum from a Mozart one with two oboes, and one of the two composers is famous for the difficulty of being heard over him.
The nineteen decibels is not a signal-to-noise ratio. It compares two long-term means, each drawn against its own peak, in a band. What it says is that there is a region where the voice’s shape stands above the orchestra’s; what it does not say is what the absolute levels are at the seat.
And the masking calculation is a check rather than a step, which the section on it now says and the section did not originally. Nothing in the nineteen decibels depends on a masking model: it is a subtraction of two spectra. What the masking model was brought in to establish — that the orchestra’s bass does not bury the singer’s formant — turns out to be true for a reason the model makes clear only when run, which is that the orchestra’s direct content at three kilohertz swamps the skirt by forty to seventy decibels. A model that is invoked to defend a result and then found to be irrelevant to it is worth keeping in the essay and worth labelling as what it is.
And the ear canal figure is a plain quarter-wave tube. A real canal is curved, its termination is not rigid, and the head and pinna contribute their own gain above two kilohertz. The measured threshold minimum is at three thousand one hundred rather than at the tube’s three thousand four hundred, and the difference is all of the things the tube model leaves out.
Whose singing, and when
The singer’s formant is a nineteenth-century object. It is strong in the recorded operatic tradition from the beginning of recording; it is what the training produces; and it is absent from most other singing on earth, including a great deal of European singing before the orchestra got large.
That timing is not a coincidence either. The orchestra roughly doubled in size between 1800 and 1900, the pit got deeper, the houses got bigger, and the technique that produces a peak at three kilohertz became standard over the same century. Whether the technique was invented for the orchestra or merely selected by it is not something this site can settle. What it can say is that the resource exists, that it is worth nineteen decibels in the one band where nothing else is competing, and that the ear was already listening hardest there.
The ladder from here
The last thing this ladder has not measured is the note itself, which turns out not to be a note. The next rung draws the width of a sung pitch on the axis the tuning essays use, and finds it larger than every distinction those essays are about.
Part 3 of 13
One essay in the series on the voice. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
FormantHearing thresholdMaskingOrchestrationRegisterResonanceSource-filterSpectral balance
- An exit is worth nothing until the tutti is given up masking, orchestration, register
- An instrument is not one timbre register, resonance, source-filter
- Room is used up by whoever enters first masking, orchestration, register
- The blend table has a row for every note formant, register, source-filter
- The chord that has room for an entrance masking, orchestration, register
- The listener is given the top voice, and the bass as a sine masking, orchestration, spectral balance