Instruments and their design

The partial that gets louder as it goes sharp

Eleven earlier essays sweep a set of partials and hold their amplitudes still, and a real tract does not move with the fundamental. Put the formant account and the sweeping account together and every partial acquires an amplitude modulation at the vibrato rate: 0.07 decibels on the fundamental and 8.6 on the twelfth partial of the same note. They are in phase below a formant and anti-phase above one — not ninety degrees apart — and the whole note swings 0.82 decibels, because they cancel.

Assumes: The pitch that does not wobble · The sound a listener knows best

The pitch that does not wobble ended by naming the omission every rung above it shares:

Every rung above sweeps a set of partials and holds their amplitudes still, and a real singer’s tract does not move with the fundamental — so a partial sweeping past a formant is a partial whose level modulates at six hertz as well as its frequency.

Both halves of that are already in the collection and have been for a long time. The sound a listener knows best built the filter: a bank of resonances belonging to the shape of the mouth, fixed in absolute frequency, through which the partials of whatever note is being sung are passed. A note that is never at its pitch built the sweep: a fundamental moving 71 cents either side of the notated pitch at six cycles a second, taking every partial with it by the same number of cents.

Nothing new has to be measured to put them together. The tract’s response is a function of frequency; the frequency is a sinusoid; the level is therefore the response evaluated along the swing.

Partials 3, 4, 12 of "hod", over one vibrato cycle. The level of three partials of a 220 hertz note on the vowel in "hod", each about its own mean, over one cycle of a vibrato of ±71 cents at 6.0 hertz. The pale curve is the frequency deviation itself, for phase reference. Partial 3 at 660 hertz swings 4.22 decibels and peaks with the frequency; Partial 4 at 880 hertz swings 0.41 decibels and peaks twice a cycle; Partial 12 at 2640 hertz swings 8.62 decibels and peaks against it. The formants of this vowel are at 730, 1090, 2440 hertz and do not move; a partial below one rises as the frequency rises and one above it falls, so the modulations of a single note run in opposite directions at the same instant.
Fig. 1 Three partials of a 220 hertz note on the vowel in “hod”, each drawn about its own mean over one vibrato cycle, with the frequency deviation dashed behind them for phase reference. The third partial rises when the frequency rises, the twelfth falls, and the fourth peaks twice a cycle.

The filter has to be allowed to be still

The one modelling decision worth arguing about is whether the tract can be treated as responding instantly, and the arithmetic settles it comfortably.

A formant is a two-pole resonance of about eighty hertz bandwidth, and such a resonator’s envelope settles with a time constant of one over pi times the bandwidth — about four milliseconds. A vibrato period is 167. So the level at each instant is the level the filter would have if the frequency had been sitting there all along, and the lag between the frequency modulation and the amplitude modulation it produces is under nine degrees rather than ninety.

That matters because ninety degrees was the natural guess, and it is what the debt above proposed. Ninety degrees would mean the level tracked the rate of change of the frequency, which is what a filter with a long memory would do. This filter has a short one. The amplitude modulation is in phase with the frequency modulation where the spectral envelope rises, and exactly anti-phase where it falls, and there is no intermediate case except at a turning point, where the first-order term vanishes altogether.

The vowel in "hod", sung at 220 HzThe partials of a 220 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.F1730 HzF21090 Hz050010001500200025003000hertzamplitudethe filter — the shape of the mouththe partials — the pitch being sungPeterson & Barney, 1952 — mean adult male values7 partials carry most of the identity
Fig. 2 The vowel that produces those signs: the partials of a 220 hertz note, each drawn at the amplitude the tract’s resonances give it, with the dashed curve the resonances themselves. The peaks do not move when the note does. Everything on this page is that curve read along a swing.

Two orders of magnitude, inside one note

Reading the curve along the swing for every partial gives a map, and the map is startling in its range.

The vibrato read off "hod", sung at 220 hertz. How many decibels each partial of a 220 hertz note swings over one vibrato cycle of ±71 cents, on the vowel in "hod". The depth runs from 0.07 decibels at partial 1 to 8.62 at partial 12 — a factor of 115 within one note, because the depth is the slope of the vocal tract's response read along the swing and has nothing to do with effort. Bars drawn light are partials whose level rises as the frequency rises; bars drawn dark are above a formant and fall as it rises. The whole note's level swings 0.82 decibels, because those two sets cancel — and how completely they cancel is a fact about where the partials happen to fall rather than about the vibrato.
Fig. 3 How many decibels each partial of the same note swings over one vibrato cycle. Light bars rise as the frequency rises; dark bars are above a formant and fall as it rises. The horizontal line is what the whole note does.

The fundamental of that note swings 0.07 decibels and the twelfth partial swings 8.62 — a factor of a hundred and twenty, within one note, produced by one singer doing one thing. The depth is not a measure of effort or of the vibrato’s size in any local sense; it is the slope of the vocal tract’s response at that partial’s frequency, converted into decibels, and a partial that happens to sit on a steep flank modulates hugely while a partial on a plateau does not modulate at all.

The signs are as informative as the depths. Below the first formant every partial rises as the frequency rises; between the first and second every partial falls; between the second and third they rise again. One note carries partials modulating in opposite directions at the same instant, and the pattern of reversals is a fingerprint of the vowel, because the reversals sit at the formants.

Three of the sixteen partials are the interesting exception. The fourth partial at 880 hertz, the fifth at 1100 and the eleventh at 2420 sit at turning points of the envelope — near a formant peak or in the trough between two — and at a turning point the first-order term is zero. Their levels peak twice a cycle, at twelve hertz rather than six, with a depth that is second order and small: 0.41, 3.33 and 4.70 decibels against the 8.62 of the partial next door.

That is a real and testable prediction with no adjustable parameter in it. A partial sitting on a formant peak modulates at twice the vibrato rate. Nobody looking at a spectrogram would expect a component at 12 hertz in a signal whose only periodicity is 6, and it is there because a maximum is symmetric about itself.

It also gives the account a self-check that costs nothing. The doubling and the sign reversal have to happen at the same place, since both are consequences of the derivative passing through zero — so a partial whose level peaks twice a cycle must be a partial whose correlation with the frequency has collapsed, and one whose correlation is decisive must peak once. Every figure on this page tests exactly that, at every partial it draws, and a formant misplaced by a few hundred hertz would break it.

The note is not louder, and the reason is arithmetic

The debt above proposed that this term would explain why a vibrato is louder as well as wider. It does not, and the way it fails is the useful part.

Why the note is not louder, when its partials are. One vibrato cycle on the vowel in "hod" at 220 hertz, with the three deepest partials drawn about their own means and the total radiated level drawn on the same axis. The partials swing 5.7, 8.6, 5.1 decibels; the note swings 0.82. They cancel because a formant has two sides — 1 of the three rise as the frequency rises and 2 fall — so the energy moves about inside the spectrum without leaving it. A vibrato does not make a note louder. It makes the note's parts fluctuate while its total holds still, which is a different thing and is the one a listener has a mechanism for.
Fig. 4 The three deepest partials of the same note drawn about their own means, with the total radiated level on the same axis. The partials swing several decibels each; the note swings 0.82.

Sum the partials and the total level over the cycle swings 0.82 decibels. The deepest single partial swings 8.62. The whole note therefore fluctuates by a tenth of what its loudest-fluctuating part does, and the reason is the sign structure above: a formant has two sides, the partials below it rise while the partials above it fall, and the energy moves about inside the spectrum without leaving it.

Nought point eight of a decibel is at the edge of what a listener detects as a level change on a sustained tone. So the account this rung can give of what a vibrato buys is emphatically not loudness.

The cancellation is not a law, and the first version of the computation asserted it as one and was wrong. Sing the same vowel with the same vibrato at 330 hertz instead of 220 and the total swings 3.00 decibels rather than 0.82 — nearly four times as much — because at that fundamental the partials fall less evenly on the two sides of each formant and the two sets no longer match. So how nearly a vibrato cancels itself in the total is a fact about which partials the note happens to have, and it changes from note to note within a phrase. What is a law is only the inequality: the note cannot swing as far as its deepest partial does, because the deepest partial is being opposed by something. It is that the note’s parts fluctuate while its total holds still — which is a different thing, and, as it happens, the thing a listener has a mechanism for.

Both facts are worth holding together, because they are the sort of pair that gets collapsed. The energy in the singer’s spectrum is being stirred vigorously at six hertz. The energy leaving the singer is not.

The same shape has now appeared three times on this ladder and it is worth naming. The roughness a vibrato produces is large and fluctuating; the mean of it survives a listener’s window almost unchanged. The pitch swings 142 cents peak to peak and the extracted pitch barely moves. The partials swing several decibels each and the total holds within one. In every case a vibrato is enormous in the quantity being modulated and nearly invisible in the quantity a listener would report — which is not three coincidences but one property, that a vibrato is a fast excursion about a mean and every mechanism the ear has integrates.

Where the stirring happens is where the voice is audible

The partials that modulate deepest on this vowel are the tenth to the fifteenth, between 2,200 and 3,300 hertz, at three to eight and a half decibels each.

Partials 10, 11, 13 of "hod", over one vibrato cycle. The level of three partials of a 220 hertz note on the vowel in "hod", each about its own mean, over one cycle of a vibrato of ±71 cents at 6.0 hertz. The pale curve is the frequency deviation itself, for phase reference. Partial 10 at 2200 hertz swings 5.72 decibels and peaks with the frequency; Partial 11 at 2420 hertz swings 4.70 decibels and peaks twice a cycle; Partial 13 at 2860 hertz swings 5.11 decibels and peaks against it. The formants of this vowel are at 730, 1090, 2440 hertz and do not move; a partial below one rises as the frequency rises and one above it falls, so the modulations of a single note run in opposite directions at the same instant.
Fig. 5 The tenth, eleventh and thirteenth partials of the same note, in the region where a trained voice stands clear of an orchestra. The eleventh sits on the third formant and peaks twice a cycle; the two either side of it modulate in opposite directions.

That band is not an arbitrary place to find them. One voice over ninety players computed where a trained operatic voice stands furthest clear of the orchestra it has to be heard through, and the answer was 18.9 decibels of advantage at 3,171 hertz — the singer’s formant, a region where the voice has a peak and the orchestra’s long-term average spectrum has fallen away.

Long-term average spectra: an orchestra, playing forte against a trained operatic soloist. Each source's mean spectrum over a long passage, in decibels below its own strongest region, on a logarithmic frequency axis. An orchestra, playing forte peaks at 250 Hz and is 30 dB down by 3,150 Hz; a trained operatic soloist peaks at 250 Hz and is 11 dB down by 3,150 Hz. The shapes are the same until about 1 kHz and separate above it: at 3171 Hz the difference is 18.9 decibels, which is the largest anywhere in the range. Nothing here is about level. Both curves are drawn against their own peaks, so what is being compared is shape.
Fig. 6 The long-term average spectra of a trained soloist and of an orchestra playing forte. The voice’s advantage peaks near three kilohertz, which is where the deepest amplitude modulations on this page happen to fall.

The two formants that make a vowel are what identify it; this third region is what makes it audible at all. So the deepest modulations in a sung note land in the band where the voice already has the largest margin over the accompaniment — and the accompaniment there is, over a long passage, a nearly steady noise floor. That is a conjunction worth stating plainly and then hedging carefully.

Stated plainly: a fluctuating signal in a steady background is easier to detect than a steady signal of the same average level, which is the whole subject of comodulation and of every account of why a modulated target pops out of noise. The vibrato puts several decibels of six-hertz fluctuation exactly where the voice is trying to be heard.

Hedged: nothing on this page measures a listener, the orchestra’s spectrum is only steady in the long-term average and is violently unsteady instant by instant, and the mechanism just named is a real one that has never been tested on this configuration. The conjunction is a hypothesis with an arithmetic behind it rather than a result.

What the arithmetic does establish is that the band is not a coincidence of the vowel. The depths there are large because the spectral envelope is falling steeply above the highest formant, and it falls steeply above the highest formant of every vowel there is.

A different vowel is a different map

Changing the vowel changes the formants and therefore changes the whole picture, which is the strongest evidence that the depths are a property of the filter rather than of the vibrato.

The vibrato read off "heed", sung at 220 hertz. How many decibels each partial of a 220 hertz note swings over one vibrato cycle of ±71 cents, on the vowel in "heed". The depth runs from 0.07 decibels at partial 3 to 9.99 at partial 11 — a factor of 141 within one note, because the depth is the slope of the vocal tract's response read along the swing and has nothing to do with effort. Bars drawn light are partials whose level rises as the frequency rises; bars drawn dark are above a formant and fall as it rises. The whole note's level swings 1.37 decibels, because those two sets cancel — and how completely they cancel is a fact about where the partials happen to fall rather than about the vibrato.
Fig. 7 The same note, the same vibrato, on the vowel in “heed” — first formant at 270 hertz, second at 2,290, third at 3,010. The fundamental now sits below the first formant and the map is entirely rearranged.

The vowel in “heed” has its first formant at 270 hertz, which is below the second partial of this note. So the fundamental at 220 rises with the frequency and the second partial at 440 already falls — a reversal between the first two partials, which the open vowel does not have until the fourth. Its second and third formants are close together at 2,290 and 3,010, which puts partials ten to fifteen on alternating steep flanks: 9.91, 9.99, 1.65, 9.25, 8.81 and 8.01 decibels, with three sign changes among them.

The total for that vowel swings 1.37 decibels against the open vowel’s 0.82 — larger, because the reversals are not as evenly matched, and still an order of magnitude below the parts.

So a singer changing vowel under a held note changes the depth and the sign of the modulation on every partial while changing nothing about the note, the pitch or the vibrato. That is a prediction about a signal anybody can record, and it is the cleanest available test of everything above: hold a note, move the mouth, and the modulation pattern should rearrange itself in a way the formant positions predict exactly.

The depth is a straight line in the extent

One more quantity falls out and it is the one that connects this rung back to the measurements the ladder started from.

The depth is a slope times a swing, and to a good approximation over the measured range the swing is the only thing changing. The twelfth partial of the open vowel swings 4.06 decibels at an extent of ±34 cents, 8.62 at ±71 and 14.41 at ±123 — which are the bottom, the middle and the top of the range ten tenors were measured singing. The first two are a straight line to within two per cent and the third bends slightly, because at ±123 cents the partial is swinging far enough to reach a change of slope.

That makes the modulation depth a read-out of the vibrato’s extent in a unit a listener measures directly, which the extent in cents is not. A ratio of 71 cents to 34 is not something the ear reports; a difference of four and a half decibels on a partial is. Whether any listener uses it is not a question this collection can answer, and it is the first thing on this ladder for which the answer would have to come from a listener rather than from an arithmetic.

Which computation produced the numbers

The filter is the collection’s own formant bank: each formant a two-pole resonator of stated bandwidth — 80, 110 and 160 hertz — combined in power, with the centre frequencies from Peterson and Barney’s 1952 survey averaged over adult male speakers. The source falls at six decibels an octave, which is the usual simplification of a glottal pulse train at twelve plus lip radiation at six.

A partial’s level over the cycle is the filter’s gain at that partial’s instantaneous frequency, divided by the partial number for the source slope, evaluated at ninety-six equally spaced phases of one vibrato cycle. The depth is the maximum minus the minimum in decibels. The sign is the correlation between the level in decibels and the frequency deviation in cents over the same cycle, and it comes out within a hundredth of plus or minus one wherever the envelope has no turning point inside the swing.

The count of level maxima per cycle is counted on the drawn trace rather than inferred, which is what distinguishes a partial on a flank from one on a peak.

The total is the sum of the partial powers at each point in the cycle, over thirty partials, converted to decibels. Nothing is normalised between the two, so the partial curves and the total curve are on the same axis and the cancellation is visible rather than asserted.

The vibrato is ±71 cents at six hertz unless a figure says otherwise, which is the measured central value the ladder has used since its fourth rung.

Where the model stops

Three formants is not a tract. Real vocal tracts have four and five and higher resonances, and above about 3.5 kilohertz the model’s envelope is the tail of the third pole rather than anything measured. The depths in the singer’s-formant band are therefore the least trustworthy numbers on the page, and they are the ones the loudest claim rests on.

The singer’s formant is not in the filter. A trained operatic voice clusters its third, fourth and fifth formants to make the peak near three kilohertz, and the survey formants used here are speech values with no such cluster. So the model produces large depths in that band for the ordinary reason — a steep envelope — rather than for the trained one, and a real soloist’s peak would change the sign structure there.

The source is held constant. A real larynx does not produce an identical pulse at every instant of a vibrato: the fold tension is varying, which is what makes the pitch vary, and the open quotient varies with it. That would put a modulation on the source spectrum as well, in phase with the frequency, and this model has none.

And the vibrato is a sinusoid. Measured vibratos are close to sinusoidal and are not exactly, and a departure from a sinusoid puts harmonics into every modulation drawn here — which matters most for the doubling claim, since a twelve-hertz component would then have two possible origins.

What the picture cannot show

It cannot show what a listener does with any of it. Everything above is a property of a signal. Whether an eight-decibel fluctuation on the twelfth partial is heard as a fluctuation, as a timbre, or as nothing, is the question the ladder cannot answer without one.

Nor can it show two voices. The modulations of two singers on the same vowel have the same depths and independent phases, and what a section does with that is a different calculation — the one the choir rungs began with mistunings rather than with modulations.

It cannot show the onset. A vibrato takes a few hundred milliseconds to establish, so the beginning of every note is a straight tone with none of this in it — and the beginning of a note is where a great deal of identification happens.

And it cannot show the room. Every level here is at the singer. A hall’s own response has peaks and troughs of its own, at a spacing much finer than a formant’s, and a partial sweeping 142 cents in a real room is sweeping across those too.

Whose singing, and when

The vibrato is the Western operatic one of the twentieth century, measured on recordings of male soloists. The formants are American English speech values from 1952, averaged over seventy-six speakers, which is a claim about a language and a population rather than about vowels.

Neither is a claim about singing in any other tradition. A straight-tone practice has none of this at all, and a tradition whose ornament is a trill rather than a vibrato has a modulation at a rate and depth this model would have to be handed rather than deriving.

Where this ladder goes next

Twelve rungs. The larynx as a reed; two mechanisms and a seam; one voice over ninety players; a note that is never at its pitch; the sound a listener knows best; what a choir buys; two sections beating between partials; those partials moving; the roughness that motion produces; which of its three statistics survives a window; the note itself, which survives all of it; and now the amplitude modulation nobody put there, which is a read-out of the vowel rather than of the singer.

What is owed after this is the rate. Every figure on every rung above, this one included, gives its vibrato the same 6.0 hertz, and the measured spread is 5.5 to 7.5 — a parameter varied once in twelve rungs and never swept. It matters most exactly where this rung leaves off, because a section of singers each modulating its own partials at its own rate is a different object from a section all modulating at one, and because the collection already knows that the beating between mistuned voices rises with the partial while a modulation in cents does not. Two fluctuations that scale differently must cross somewhere, and where they cross is arithmetic this collection has every part of.

Part 12 of 13

One essay in the series on the voice. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

FormantMaskingModulationSource-filterSpectrumVibratoVocal tract