Instruments and their design

The rate that does not rise with the partial

Twelve earlier essays give every vibrato the same six hertz, and the measured spread is 5.5 to 7.5. Putting the two fluctuations a choir contains on one axis shows why the rate matters: the beating between mistuned voices rises with the partial and leaves the range a listener follows as fluctuation at 1,217 hertz, while the vibrato's own modulation is six hertz at every partial. Above that frequency a section fluctuates by vibrato alone — and if every singer had the same rate, it would barely fluctuate at all.

Assumes: The partial that gets louder as it goes sharp · Every partial beats at its own rate

Look back along this ladder and one number barely changes. Almost every vibrato in almost every figure on twelve rungs runs at 6.0 hertz — in the trace, in the beat rates, in the roughness, in the windowed roughness, in the autocorrelation, and in the amplitude modulation the formants produce. The extent has been swept repeatedly and drawn on a slider — across the whole range ten tenors were measured at, and against every comma the field argues about. The rate has been changed in exactly one figure of the twenty-nine that draw a vibrato, to 7.5, and never swept at all.

That is not carelessness; 6.0 is the measured central value and holding it constant is the right thing to do while the argument is about something else. But the same measurement that supplies the 6.0 supplies a spread with it — 5.5 to 7.5 hertz across singers — and a parameter with a documented spread that twelve rungs have held at a point is exactly the shape of an unasked question.

It turns out to matter in two ways, and only the second is about the spread. The first is about the rate being a rate at all.

Where a section's fluctuation stops being a beat, on 220 hertz. Two rates up the spectrum of a 220 hertz note sung by a section whose voices are spread by 15 cents. The rising line is the beat rate between a typical pair of them, which grows with the partial because a mistuning in cents is a difference in hertz that scales with frequency; it reaches the 15 hertz at which a beat stops being a beat by partial 5.5, at 1217 hertz. The flat line is the amplitude modulation the vibrato imposes through the formants, which is 6.0 hertz at every partial because the vibrato modulates every partial by the same number of cents at the same rate. The two are equal at 487 hertz. Above 1217 hertz the beating has become roughness and the only fluctuation left is the vibrato's — and that frequency is the same one an octave up, where it is partial 2.8 instead.
Fig. 1 The two fluctuations a section of singers produces, up the spectrum of one note. The rising line is the beat rate between a typical pair of them; the flat line is the vibrato’s own amplitude modulation. They cross where the beating leaves the range a listener follows as a beat.

Two fluctuations that scale differently

A choir on one note fluctuates for two reasons and the ladder has treated them as one.

The first is mistuning. Sixteen singers aiming at the same note are spread by something like fifteen cents, and every pair of them beats. Every partial beats at its own rate established the crucial property of that: a mistuning stated in cents is a difference in hertz that scales with frequency, so the k-th partials of a pair beat k times as fast as their fundamentals do, and somewhere up the spectrum the rate passes the point at which a beat stops being a beat and becomes roughness.

The second is vibrato. The rung below this one showed that a partial sweeping past a fixed formant has its level modulated as well as its frequency, by up to eight or nine decibels, at the vibrato rate.

And the vibrato’s modulation does not scale. The vibrato moves every partial by the same number of cents at the same rate, so every partial’s level fluctuates at six hertz whatever the partial is. One fluctuation rises linearly up the spectrum and the other is flat, which means they cross exactly once, and the crossing is worth locating.

Below the crossing, a choir is beats

At the bottom of the spectrum the mistuning wins easily. A typical pair of voices in a section spread by fifteen cents differs by about 21 cents, which on a 220-hertz fundamental is 2.7 beats a second — comfortably inside the range a listener follows as a fluctuation, and comfortably below the vibrato’s six.

The fluctuation stays; the rate goes. A unison of n voices with a spread of 15 cents, averaged over 5 draws. The depth of the amplitude fluctuation does not fall as voices are added — a choir is no steadier than a duet — but the fraction of that fluctuation in any single modulation component falls from 73 per cent at two voices to 21 at 32. Two voices make one beat and it can be counted; 16 make 120 and none of them is a rate. That is why a choir cannot be tuned by nulling anything.
Fig. 2 The fluctuation a unison produces from mistuning alone, against how many are singing. The depth does not fall as voices are added; what falls is the share of it in any single beat rate, which is why a choir cannot be tuned by nulling anything.

That is what a choir does that a soloist cannot, and it is entirely an account of the fundamental and the low partials. The two rungs after it — two sections against each other and then sixteen sweeps against sixteen — extended it to intervals and to moving partials, and every one of them is drawn at a coincidence low in the series, where the beating is a beating.

There is a reason none of them went higher, and it is a good one: a coincidence between two sections an interval apart is a coincidence between low partials by construction, since a major third brings the fifth partial of one against the fourth of the other and the ratios stop being small very quickly. So the whole choral account this ladder holds was built where the beating account is valid, and nothing in it was wrong. What it could not see is that the region it was built in has an edge.

Above it, a choir is vibrato

The beating passes fifteen hertz — the rate at which this collection stops calling a fluctuation a beat — at 1,217 hertz for a section spread by fifteen cents. Above that frequency the mistuning is producing roughness rather than fluctuation, and the only thing left fluctuating at a rate a listener can follow is the vibrato’s own modulation, flat at six hertz.

The number that matters about that crossing is not where it falls on the harmonic series. It is that it is a frequency.

Where a section's fluctuation stops being a beat, on 440 hertz. Two rates up the spectrum of a 440 hertz note sung by a section whose voices are spread by 15 cents. The rising line is the beat rate between a typical pair of them, which grows with the partial because a mistuning in cents is a difference in hertz that scales with frequency; it reaches the 15 hertz at which a beat stops being a beat by partial 2.8, at 1217 hertz. The flat line is the amplitude modulation the vibrato imposes through the formants, which is 6.0 hertz at every partial because the vibrato modulates every partial by the same number of cents at the same rate. The two are equal at 487 hertz. Above 1217 hertz the beating has become roughness and the only fluctuation left is the vibrato's — and that frequency is the same one an octave up, where it is partial 1.4 instead.
Fig. 3 The same two lines an octave higher. The crossing is at the same 1,217 hertz — because the beat rate depends on absolute frequency and nothing else — but it is now the third partial rather than the sixth.

On A3 the crossing is at the fifth or sixth partial. On A4 it is at the third. On a soprano’s high C it is barely above the fundamental. The partial number moves and the frequency does not, because the beat rate at a given frequency depends on the mistuning in cents and on that frequency, and not at all on which note is being sung.

So a choir’s sound is divided by a horizontal line rather than by a partial number, and every voice in the section crosses it at the same place. Below about 1,200 hertz a section fluctuates because its members are mistuned; above it, because they have vibrato.

The width of the section’s own tuning is what sets the line. A very well tuned section — five cents rather than fifteen — pushes it up to 3,665 hertz, so almost the whole audible spectrum is on the beating side. A ragged one at twenty-five cents pulls it down to 727, so almost nothing is. The crossing is therefore a measurement of a section’s tuning that has nothing to do with whether the tuning sounds good.

Which puts the singer’s formant on the vibrato side

That the crossing lands where it does is the part with a consequence.

The vibrato read off "hod", sung at 330 hertz. How many decibels each partial of a 330 hertz note swings over one vibrato cycle of ±71 cents, on the vowel in "hod". The depth runs from 0.20 decibels at partial 1 to 8.62 at partial 8 — a factor of 43 within one note, because the depth is the slope of the vocal tract's response read along the swing and has nothing to do with effort. Bars drawn light are partials whose level rises as the frequency rises; bars drawn dark are above a formant and fall as it rises. The whole note's level swings 3.00 decibels, because those two sets cancel — and how completely they cancel is a fact about where the partials happen to fall rather than about the vibrato.
Fig. 4 The depth of each partial’s amplitude modulation on an open vowel sung at 330 hertz. The deepest sit at the seventh and eighth partials, above two kilohertz — well past the crossing, where beating has already become roughness.

A trained voice stands furthest clear of an orchestra near 3,200 hertz, which is where the long-term spectra separate by nearly nineteen decibels. That is two and a half times the crossing frequency. The band a voice is actually heard through an accompaniment in is therefore entirely on the vibrato side of the line, at every pitch anybody sings.

The consequence for a section is direct. Whatever a choir’s mistuning is doing at 3 kilohertz, it is not producing a followable fluctuation there; it is producing roughness, which is a timbre. What produces fluctuation in that band is the vibrato, at six hertz, at a depth set by the slope of the vowel — several decibels, from the rung below.

That reverses the natural reading of a choral sound. The mistuning is what makes a choir sound like a choir low down, and it is the vibrato that does it high up, in the one band where the voices have to compete with anything.

What a spread of rates does

Now the spread, and here the arithmetic refuses the obvious answer twice.

Take sixteen voices, each modulating one partial by six decibels at its own rate, with independent phases. Give every one of them exactly 6.0 hertz, which is what twelve rungs have done, and the section’s level fluctuation collapses to 1.54 decibels — because sixteen modulations at one frequency are sixteen phasors at one frequency, and their resultant is a single draw whose size is then fixed for the whole note. Draw the rates from the measured spread instead and the fluctuation holds 2.65 decibels, seventy-two per cent more.

16 voices, and what a spread of rates keeps alive. The summed level of 16 voices whose partials each modulate by 6 decibels at their own vibrato rate, over 8 seconds. The pale trace gives every singer the rate used in eleven earlier essays — exactly 6 hertz — and the section's fluctuation collapses to 0.65 decibels, because 16 modulations at one frequency add to one modulation whose size is fixed for the whole note. The dark trace draws the rates from the measured spread of 5.5 to 7.5 hertz, and the fluctuation holds 2.21 decibels — 238 per cent more — because the modulations precess against each other and the sum revisits its large values within a second or two. What the spread does not do is add anything slow: 0 per cent of the fluctuation lies below three hertz, and the peak is still at 6.3.
Fig. 5 The summed level of sixteen voices over eight seconds, with one rate for everybody and with the measured spread. The spread does not add a new fluctuation; it stops the existing one from cancelling itself.

The mechanism is worth stating carefully because it is not the one a reader expects. With identical rates the phasors hold their relative phases forever, so whatever partial cancellation they happen to start with, they keep. With a spread of half a hertz they precess against each other on a timescale of a second or two, so the resultant wanders — and over any few seconds it revisits its large values. The depth measured over a phrase is therefore close to the largest the phasors can reach rather than to a typical draw.

A spread of rates does not create a fluctuation. It prevents an accidental cancellation from lasting.

That is the same lesson the mistuning rungs reached from the other side, and it is worth putting the two statements next to each other because they are not the same statement. A unison’s fluctuation does not fall with the number of voices because the beat rates are all different, so there is nothing for them to cancel into. A section’s vibrato modulation does fall, because the rates are nearly the same, and the measured spread is just wide enough to stop it falling as fast as it would. One mechanism is saved by its own disorder and the other is rescued by a little.

And what it does not do

The obvious guess about a spread of rates is that it produces slow beats between them: two singers at 5.8 and 6.4 hertz ought to make something at 0.6, and a section ought to acquire a slow swell at the difference frequencies. That would be a satisfying explanation of what a large choir sounds like, and it is the thing this rung was expected to find.

It is not there.

4 voices, and what a spread of rates keeps alive. The summed level of 4 voices whose partials each modulate by 6 decibels at their own vibrato rate, over 8 seconds. The pale trace gives every singer the rate used in eleven earlier essays — exactly 6 hertz — and the section's fluctuation collapses to 1.52 decibels, because 4 modulations at one frequency add to one modulation whose size is fixed for the whole note. The dark trace draws the rates from the measured spread of 5.5 to 7.5 hertz, and the fluctuation holds 5.23 decibels — 243 per cent more — because the modulations precess against each other and the sum revisits its large values within a second or two. What the spread does not do is add anything slow: 1 per cent of the fluctuation lies below three hertz, and the peak is still at 6.3.
Fig. 6 Four voices rather than sixteen, over the same eight seconds. The waxing and waning of the fluctuation is visible — that is the phasors precessing — but the fluctuation itself is still at the vibrato rate, and there is no slow component underneath it.

Across every section size and every spread computed, less than one per cent of the level fluctuation lies below three hertz, and the peak of the fluctuation spectrum stays within a fifth of a hertz of six. The difference frequencies do not appear.

The reason is that they are second order. The level of a sum is very nearly the sum of the levels, so a section’s fluctuation is very nearly the sum of its members’ modulations rather than any product of them, and only a product makes a difference frequency. The nonlinearity that would produce one is the logarithm in the decibel, and at these depths it is too gentle to matter.

That is a null and it is worth recording as one, because it removes an explanation that was available. The waxing and waning a listener hears in a large choir is not a slow beat between vibrato rates. It is the envelope of a six-hertz fluctuation whose amplitude wanders, which sounds different and has a different timescale — set by the spread of the rates rather than by their differences.

Against the size of the section

Sweeping the section size gives the whole of it in one picture.

What eleven earlier essays lost by giving everybody the same rate. The depth of a section's level fluctuation against how many are singing, under two assumptions about the vibrato rate. With one rate for everybody the fluctuation falls from 6.0 decibels at one voice to 1.13 at 32: the modulations are phasors at one frequency and their resultant is one draw, fixed for the note. With the measured spread of 5.5 to 7.5 hertz it holds 1.88 decibels at the same size, because the phasors precess and the sum keeps returning to its large values. The gap between the curves is what a spread of rates buys, and it is the one parameter of a vibrato nobody here has never varied.
Fig. 7 The depth of a section’s level fluctuation against how many are singing, under the two assumptions about the rate. The gap between the curves is what a spread of rates buys, and it is the parameter twelve essays held still.

Both curves fall, which is expected: adding voices averages a fluctuation down. What is not expected is that they fall at different rates and stay apart. At four voices the identical-rate assumption gives 3.07 decibels and the measured spread gives 5.10; at sixteen, 1.54 against 2.65; at thirty-two, 1.13 against 1.88. The ratio holds at about seventy per cent all the way along, which says the effect is not a small-number accident.

Seventy per cent more fluctuation in the band where the voices are trying to be heard is not a subtlety. And it is a gap that twelve rungs of this ladder have been on the wrong side of, silently, by using the mean of a measurement instead of the measurement.

Which computation produced the numbers

The beat rate at the k-th partial is k times the fundamental times the fractional difference between two voices, taken as the standard deviation of the section’s tuning times the square root of two, which is the standard deviation of the difference between two independent draws. Fifteen cents per voice therefore gives 21.2 cents per pair.

The rate at which a beat stops being a beat is this collection’s own stated fifteen hertz, which is a convention rather than a measurement and is used here as one; the crossing frequency scales inversely with it, so a limit of twenty hertz would put the crossing at 1,624 rather than 1,217.

The vibrato’s modulation rate is the vibrato rate, exactly, because the level of a partial is a function of its frequency and the frequency is periodic at that rate. The doubling at a formant peak, from the rung below, is the one exception and it is not drawn here.

The ensemble traces are the sum of the powers of n voices, each modulating as ten to the power of a sinusoid, with rates drawn from a normal distribution about 6.0 clipped to the measured 5.5 to 7.5 and phases drawn uniformly. Depth is the peak-to-trough range of the total in decibels over eight seconds; the spectrum is a Hann-windowed transform of it, and the slow share is the fraction of its power below three hertz. Every figure averages several draws with fixed seeds, so a picture of a choir is the same choir on every reading.

The six decibels of per-voice modulation is a representative value taken from the rung below rather than a measurement of anything; the ratio between the two curves is insensitive to it, and the absolute depths are not.

Where the model stops

Every voice modulates one partial. The traces above sum a single partial across a section, and a real section sums a whole spectrum in which every partial has its own depth and its own sign. Partials modulating in opposite directions within one voice would reduce that voice’s contribution before the section ever sums it.

The rates are drawn once and held. A real singer’s vibrato rate drifts within a phrase and speeds up towards the end of a note, which is documented and which would make the precession faster and the effect larger. The model has no time dependence in the rate at all.

The phases are independent and the rates are independent of them. Singers listening to each other may entrain, and if a section’s vibratos partially lock, this whole result collapses towards the identical-rate curve. Nothing here knows whether they do, and that is the single most important unmeasured quantity on the page.

And fifteen cents is one number for a whole practice. A trained chorus is tighter and an amateur one is not, and the crossing frequency moves inversely with it over a factor of five across the plausible range.

What the picture cannot show

It cannot show the roughness above the crossing. Saying that the beating has become roughness is not saying it has stopped mattering; roughness is a sensation and it has its own dependence on register. What the figure claims is only that it is no longer a fluctuation a listener follows.

Nor can it show a listener’s own filters. Two partials of a section land in the same auditory filter or in different ones depending on their spacing, and everything above sums them arithmetically, which is what a microphone does. Whether the sum is heard as one voice or many is a further question again.

It cannot show the room. A hall’s reverberation is a smoothing in time, and a six-hertz fluctuation arriving through a reverberant tail is a shallower fluctuation than one arriving direct. That is a real reduction and it applies to both curves equally.

And it cannot show what the fluctuation is worth. Whether seventy per cent more six-hertz modulation in the singer’s-formant band makes a section more audible is a question for a listener, and this ladder has now reached the point where nearly every remaining question is.

Whose choirs, and when

The vibrato rates and extents are twentieth-century Western operatic values from recordings of soloists, applied here to sections, which is an extrapolation the sources do not license. A chorus does not sing with a soloist’s vibrato and often deliberately reduces it.

The fifteen cents of tuning spread is the value this collection has used since the choir rungs, and it comes from measurements of choral unisons. It is not a claim about any particular ensemble.

And the whole argument assumes a practice that uses vibrato. Everything on this page is zero for a section singing straight tone, which is what a great deal of choral music asks for — and for such a section the crossing frequency still exists and there is simply nothing on the far side of it.

Where this ladder goes next

Thirteen rungs. The larynx as a reed; two mechanisms and a seam; one voice over ninety players; a note that is never at its pitch; the sound a listener knows best; what a choir buys; two sections beating between partials; those partials moving; the roughness that motion produces; which of its statistics survives a window; the note itself, which survives all of it; the amplitude modulation the formants produce; and now the two fluctuations a section contains, which change places at a frequency rather than at a partial.

What is owed after this is the onset, and it is the last thing on this ladder that can be computed without a listener. Every figure on every rung above draws a vibrato that has always been running. A real one takes two or three hundred milliseconds to establish, so the beginning of every sung note is a straight tone — with a definite pitch, no amplitude modulation, and none of the fluctuation this rung is about. That is between one and two vibrato cycles of a completely different signal at the front of every note, in exactly the window in which a listener decides what the note is and where it is. The collection has the establishment time from the same measurements the rate came from, and it has integration windows of the same order; putting them together would say what fraction of a note’s identification happens before any of this ladder’s last ten rungs applies. After that, the questions left are about listeners, and the ladder will have run out of arithmetic rather than out of subject.

Part 13 of 13

One essay in the series on the voice. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

BeatingCritical bandwidthFormantModulationRoughnessUnisonVibrato