Timbre and acoustics

What a choir does that a soloist cannot

Two singers on one note produce one beat and it can be counted. Sixteen produce a hundred and twenty at once, and the amplitude still fluctuates by as much as it did — a choir is no steadier than a duet. What has gone is not the fluctuation but its rate: the modulation energy that sat in a single line at two voices is spread across a band at sixteen, with no line in it. That is the choral sound, and it is also why the just-intonation drift this site measured describes only ensembles that hold their pitch still.

Assumes: The other instrument with a reed · Beats are arithmetic that anybody can hear

The beating ladder in this collection is about two tones. Two frequencies close together produce an amplitude fluctuation at their difference; a tuner counts it; a mistuned unison inside one instrument produces the other kind of wolf. Every rung assumes the rate can be followed, because a rate that can be followed is the whole use of the phenomenon.

A choir has sixteen people on a note. The arithmetic of two tones does not stop applying; it applies a hundred and twenty times at once.

What sixteen singers are, as a signal

Nobody sings a note exactly. Take a section of sixteen with a realistic spread — a standard deviation of fifteen cents, which is a good amateur choir and a slack professional one — and the object is sixteen sinusoidal components clustered within a few tens of cents.

Sixteen singers, and a hundred and twenty beat rates. The mistunings of 16 voices drawn from a normal distribution of 15 cents, and every pairwise beat rate they produce: 120 of them, from near zero to 6.4 hertz, with a median of 1.30. 108 of the 120 are inside the band a listener could follow as a beat if it were alone. None of them is alone.
Fig. 1 Sixteen singers’ fundamental frequencies drawn from a normal distribution of fifteen cents, and below them every beat rate the set produces: one for each pair, so a hundred and twenty of them. They run from very nearly zero to about nine hertz, with a median near one. The mistunings are drawn once with a fixed seed, so this is one choir on one night, and the sound buttons play exactly the frequencies drawn.

A hundred and twenty rates, all present simultaneously. The obvious question is what that sounds like, and the obvious guesses are both wrong. It is not sixteen audible beats — nobody can attend to a hundred and twenty rates. Nor is it a steadier note, which is what averaging usually does.

The fluctuation does not go away

This is the measurement that decides it, and it needs the right quantity. Counting beat rates is not enough, because what a listener hears is one envelope; the question is whether that envelope has a rate.

The fluctuation stays; the rate goes. A unison of n voices with a spread of 15 cents, averaged over 5 draws. The depth of the amplitude fluctuation does not fall as voices are added — a choir is no steadier than a duet — but the fraction of that fluctuation in any single modulation component falls from 77 per cent at two voices to 24 at 32. Two voices make one beat and it can be counted; 16 make 120 and none of them is a rate. That is why a choir cannot be tuned by nulling anything.
Fig. 2 Two measurements of the summed envelope against the number of voices, each averaged over five independent draws. The depth of the fluctuation — how much the level moves, relative to its mean — does not fall as voices are added; if anything it rises. What falls is the concentration: the fraction of the modulation energy sitting in any single component, which is 84 per cent at two voices and about a quarter at thirty-two.

Eighty-four per cent at two voices is a beat. One rate holds nearly all of the fluctuation, and a listener can count it, tune to it, and null it. At sixteen voices a quarter of the energy is in the largest component and the rest is spread over the band, and a quarter is not enough to be a rate — the envelope wanders rather than pulses.

So a choir is not steadier than a soloist. It is exactly as unsteady and the unsteadiness has stopped being periodic. That is worth stating in the negative because the intuition that many sources average out to something smooth is a good intuition about most things and is wrong here: the components are near-identical in frequency, so their sum does not converge, it wanders on a timescale of seconds.

That wandering is what the chorus effect is, and it is why every synthesiser has a knob for it. The electronic version detunes copies of one oscillator by a few cents, and it is doing precisely the arithmetic above with a smaller n.

220 Hz against 221.9 Hz. Two tones 1.9 hertz apart, added. The rapid oscillation is their average; the slow swelling is their difference, heard as 1.9 beats a second and used by every tuner who has ever worked by ear.
Fig. 3 Two tones fifteen cents apart at 220 hertz, and the envelope they make: a clean 1.9-hertz beat, countable by anybody. This is the two-voice case the whole account of beating is about, and it is the case a choir does not have. Adding fourteen more singers does not deepen this envelope or flatten it; it replaces it with something that has no period.

Which means a choir cannot tune by beats

The consequence for intonation is severe and it is a consequence rather than an observation.

Every method of tuning by ear in this collection works by nulling something. A tuner sets a fifth by slowing a beat to a stated rate; a bearing plan is a sequence of such counts; two players with no clock converge by hearing the discrepancy. All of it requires a rate to be present and countable.

A sixteen-voice section has no such rate. There is nothing to null, because the fluctuation is not a fluctuation at one frequency. A singer inside the section can hear their own contribution against the mass — which is a level and a roughness cue, not a beat — and can hear the section against another section, where the interval is large enough that the beating is between partials rather than fundamentals and is again a mass rather than a line.

This is the mechanism behind something choral directors say and rarely justify: that a section tunes by blend rather than by beats, and that the instruction to a section is different in kind from the instruction to a soloist. The physical difference is that one of them has a countable rate available and the other does not.

The correction it forces

There is a result in this collection that has to be re-read in the light of this, and it is one of the better ones.

The comma-pump essay shows that an ensemble tuning every chord in pure ratios cannot come back to where it started, and cites the published measurements of a cappella groups drifting downward over a piece by a comma or more. Those measurements are real and the arithmetic behind them is exact.

What has to be added is the condition under which they apply. A group that drifts by tuning each chord pure is a group that can tell when a chord is pure, which means a group with a countable beat rate available, which means a group whose pitch is steady enough for the beat to survive. Every measured a cappella drift result is from an ensemble of that kind — small, professional, and using little or no vibrato.

The other rung of this ladder makes the same point from the opposite side: an operatic vibrato of 142 cents peak to peak sweeps the beat rate through zero several times a second, so a vibrato singer has no stable beat either. Between them the two rungs bound the result: pure-intonation drift is a property of ensembles that suppress pitch variation, whether the variation comes from vibrato in one voice or from spread across many.

That is not a criticism of the drift result. It is the specification of its domain, which the result did not state and which this rung owed it.

Where the spread comes from, and why it will not go away

It is worth asking whether a choir could simply be more accurate, and the answer is that the spread has three sources and only one of them is a matter of skill.

Placement error. A singer aims at a pitch and misses by some amount. This is the part training reduces, and it can be reduced a long way — a good professional section is a few cents rather than fifteen.

Vibrato. A trained solo voice varies by ±50 to ±100 cents at five or six hertz. Sixteen such voices with unrelated vibrato phases are sixteen components each sweeping a band far wider than the placement spread, and no amount of accuracy removes it. Choral training that asks for straight tone is asking singers to disable the mechanism, which is why it is contentious.

Physical difference. No two tracts are the same length, so no two singers produce the same spectrum on the same fundamental even when the fundamentals agree exactly — which is the previous rung’s whole subject. The partials above the first therefore never coincide, and the beating between them is not removable by tuning at all.

The last of these is the interesting one, because it means a choir singing a perfect unison in fundamental would still produce a chorus effect from its upper partials. The spread is not a defect being tolerated; a substantial part of it is the difference between people.

A note sung at 220 Hz, drawn in centsDeviation from the notated pitch against time, for a vibrato of ±71 cents at 6.0 cycles a second — Seashore's and Prame's measured values, which agree. The note is 142 cents wide, and the shaded bands across the middle are the syntonic comma at 21.5 cents, the Pythagorean comma at 23.5 cents, the difference limen at 440 Hz at 5.6 cents. The pitch a listener reports is near the middle of the excursion rather than at either edge — and not exactly at the middle either: because cents are logarithmic and frequency is not, the mean frequency of this trace sits 0.73 cents above the centre line.the syntonic comma21.5 centsthe Pythagorean comma23.5 centsthe difference limen at 440 Hz5.6 centsthe note itself142 cents widemean 0.73 cents sharpof the centre line0.00.20.40.60.81.01.2-50050secondscents from the notated pitch
Fig. 4 One trained voice’s pitch against time, with three tuning intervals drawn as bands for scale. A single singer with vibrato is already sweeping a range far wider than a choir’s placement spread — so a section of sixteen with vibrato is sixteen components each moving, and the static picture above is a considerable understatement of how little periodicity is left.

Twenty singers are not twenty times one

The other half of what a choir does is level, and the arithmetic is the same arithmetic with a different consequence.

Twenty singers are not twenty times one. The level of n sources against n, on the two assumptions. Incoherent sources add in power and gain 3 dB per doubling, so 32 voices are 15 dB above one — about 6 times the pressure and rather less than that in loudness. Sources in phase would add in amplitude and reach 30 dB, which no choir does and no ensemble has ever needed to. The gap between the two curves is the whole reason a section of twenty exists rather than a soloist told to sing louder.
Fig. 5 The level of n sources against n, on the two assumptions. Sources in phase would add in amplitude — six decibels per doubling — and reach thirty decibels at thirty-two voices. Sources with unrelated phases add in power, three decibels per doubling, and reach fifteen. No ensemble achieves the upper curve and none has ever needed to.

Three decibels a doubling is a modest return: sixteen singers are twelve decibels above one, which is loud but is not sixteen times anything. If the point of a chorus were volume it would be an inefficient way to get it.

That the return is so modest is the argument that volume is not the point. A section of sixteen exists because sixteen slightly different sources produce a sound one source cannot produce at any level — the aperiodic envelope of the figures above — and the twelve decibels are incidental.

The same arithmetic runs through the rest of this collection wherever sources are multiplied. A piano’s unison strings are two or three sources deliberately mistuned, for exactly this reason at exactly this scale, and the piano tuner’s decision about how much to mistune them is a decision about where on the curve between “one beat” and “a chorus” the note should sit.

What a section sounds like against another section

Everything above concerns one note. A choir spends nearly all its time on chords, and the arithmetic changes in a way worth stating even though this rung does not develop it.

A fifteen-cent spread is a fiftieth of a semitone, and the critical bandwidth is between one and three semitones across the whole register — so a choir’s placement spread sits far inside a single band at every pitch anybody sings. Whatever a section is doing to the sound, it is not being done by separating the voices into different channels.

So a choir has two regimes at once and they are not the same phenomenon. Within a section, sixteen components inside one critical band, producing a slow aperiodic envelope. Between sections, partials from thirty-two sources meeting across bands, producing roughness whose amount depends on the interval and on how spread each section is.

The obvious thing to expect is that a spread section is less rough than a precise one, because a spread of pitches spreads the coincidences too and no single pair of partials sits at the worst separation. Measuring it says that expectation is backwards for every interval a choir sings in tune.

Two sections of sixteen, each voice a complex tone, every cross pair scored with this collection’s own roughness model and averaged over forty draws:

interval between the sections spread 0 15 cents 30 cents change
unison 0.0045 0.087 0.139 thirty times rougher
octave 0.0045 0.056 0.083 eighteen times
fifth 0.053 0.068 0.076 +45%
fourth 0.101 0.109 0.113 +12%
major third 0.145 0.147 0.148 +2%
minor third 0.193 0.192 0.193 0%
tritone 0.101 0.101 0.099 −2%
semitone 0.260 0.256 0.239 −8%

The sign of the effect is decided by which side of the dissonance curve the interval sits on, and the reasoning that produced the wrong expectation is the right reasoning applied at the wrong place. Averaging a quantity over a spread of inputs moves it towards the average of its neighbourhood. A consonance is a minimum — the well the whole consonance ladder is about — so its neighbourhood is higher than it is and spreading can only climb. A dissonance sits near a maximum, so spreading falls. The scattering argument is sound on a peak and reversed in a well, and a choir spends its time in wells.

The sizes make the point rather than the signs. On the dissonant intervals the gain is a few per cent, which is nothing; on the unison and the octave the loss is more than an order of magnitude, because two identical complex tones are not rough at all and a spread one is. A section’s spread costs it precisely in the intervals it is most exposed on, and that is a prediction the model makes rather than an observation borrowed from anywhere: the unison and the octave are the two intervals a choral director is heard complaining about, and they are the top two rows.

It also puts a boundary on the previous section. Between sections, roughness is what the interaction is; and the interaction gets worse with spread nearly everywhere, while within a section the same spread is the entire point. The spread is not one thing being traded against another — it is one thing helping in one regime and hurting in the other.

Sixteen singers, and a hundred and twenty beat rates. The mistunings of 4 voices drawn from a normal distribution of 15 cents, and every pairwise beat rate they produce: 6 of them, from near zero to 3.4 hertz, with a median of 2.54. 5 of the 6 are inside the band a listener could follow as a beat if it were alone. None of them is alone.
Fig. 6 The same drawing at four voices rather than sixteen: six beat rates instead of a hundred and twenty. Four is roughly where the transition happens — few enough that individual rates are still sometimes followable, many enough that the envelope has begun to wander. A quartet and a chorus are different objects, and the difference is not one of degree in any quantity a listener attends to.

That is a useful place to mark, because it corresponds to something musicians already treat as a boundary. One to a part is chamber music and is tuned as chamber music: by ear, by beats, chord by chord. Several to a part is choral and is tuned by blend. The change of method is usually explained by the difficulty of coordinating more people, and the figures here suggest a simpler reason — past about four, the signal a small ensemble tunes by is no longer there.

The fluctuation stays; the rate goes. A unison of n voices with a spread of 5 cents, averaged over 5 draws. The depth of the amplitude fluctuation does not fall as voices are added — a choir is no steadier than a duet — but the fraction of that fluctuation in any single modulation component falls from 70 per cent at two voices to 43 at 16. Two voices make one beat and it can be counted; 16 make 120 and none of them is a rate. That is why a choir cannot be tuned by nulling anything.
Fig. 7 The same measurement with a five-cent spread rather than fifteen — a professional chamber choir rather than a good amateur one. The concentration falls with voice count in the same way and the beat rates are three times slower, so the fluctuation is slower and no more periodic. Tightening the tuning does not restore a countable beat; it moves the whole distribution down.

Which computation produced the numbers

The mistunings are drawn from a normal distribution by Box–Muller from a fixed seed, so every figure here is the same choir every time it is drawn. The frequencies are the nominal fundamental multiplied by two to the power of the drawn cents over 1200.

The beat rates are all pairwise differences: sixteen voices give 120 pairs, thirty-two give 496.

The envelope is the magnitude of the sum of the components as unit phasors, sampled at 400 hertz over two seconds. Its modulation spectrum is a direct transform up to thirty hertz — past which a fluctuation is heard as roughness rather than as a beat, which is the consonance ladder’s own boundary and not a new one. Concentration is the largest bin’s share of the total modulation power.

An earlier version of this measurement scored periodicity by autocorrelation and returned a number that did not move with the number of voices at all. A slowly varying envelope correlates well with itself at short lag whatever it is doing, so the peak sat on the search floor at every size and the figure was measuring the floor. The modulation spectrum asks the question directly, and the fact that the first attempt produced a plausible flat line is the reason it is recorded here.

The level curves are 10 log₁₀ n and 20 log₁₀ n, which are the incoherent and coherent cases and involve no acoustics beyond the choice of which applies.

Whose choirs, and when

The fifteen-cent spread is a stated assumption, not a measurement, and it is the parameter everything else scales with. Published measurements of choral unison spread vary from a few cents in a professional chamber choir to well over thirty in a congregation, and the shape of the argument is unchanged across that range — a wider spread moves the beat rates up and the concentration down, which makes the effect stronger rather than different.

The claim about drift and vibrato applies to the specific published a cappella studies, which are mostly European art-music ensembles of four to eight singers recorded in the last forty years. Traditions with large unison singing and deliberate heterophony — Georgian polyphony, Balkan village singing, congregational psalmody, Sardinian cantu a tenore — are a different object, and the wide unison is often the aesthetic point rather than an imprecision to be minimised.

What the picture cannot show

Every voice here is a single sinusoid. A real singer produces a whole spectrum, and the beating between two singers happens at every partial at once, faster at each higher one because the same cents difference is more hertz. That makes the real modulation spectrum broader than the one drawn and strengthens the argument; it also means the depth figures here are a floor.

The mistunings do not move. A real choir’s pitches wander continuously — that is most of what makes the sound — and the model draws them once and holds them. A model with drifting components would produce an envelope with even less periodicity than this one shows.

It has no onsets. Sixteen singers do not begin a note together — the spread in attack time is tens of milliseconds — and the first fifty milliseconds of a note are most of what identifies it. A model that starts every component at once has removed one of the largest differences between a section and a soloist before it begins.

And there is no room. A room adds its own decorrelation between sources arriving from different positions, and the reverberant field of a large church does a substantial part of what this essay attributes to the spread. Separating the two would need a measurement in an anechoic chamber, and choirs are not usually found in them.

The ladder from here

The two rungs owed when this anchor opened are now written, and between them they have changed what an earlier result means rather than only adding to it.

What the ladder still lacks is the acquisition question — a listener’s ability to normalise for an unfamiliar tract, which appears very early and is not obviously learnable from what an infant hears — and the ensemble question one level up: what a section does against another section, where the interval is large and the beating is between partials.

The second is now opened rather than owed. The table above gives the sign and the size of what spread does to each interval, and it does so with two sections held at a fixed distance; what it does not do is let the sections move. A real choir’s two sections negotiate their interval while each is spreading, and the quantity that decides where they settle is the gradient of that table rather than its values. A section sitting in a well has a restoring force towards the pure interval and a section sitting on a slope has a push along it, so the same measurement run as a derivative would say which intervals a choir is driven towards and which it merely occupies — which is the shape of the rung after this one, and it needs no new machinery.

Part 6 of 13

One essay in the series on the voice. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 19.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

BeatingChorus effectIncoherent additionJust intonationModulationUnisonVibrato