What a choir does that a soloist cannot
Assumes: The other instrument with a reed · Beats are arithmetic that anybody can hear
The beating ladder in this collection is about two tones. Two frequencies close together produce an amplitude fluctuation at their difference; a tuner counts it; a mistuned unison inside one instrument produces the other kind of wolf. Every rung assumes the rate can be followed, because a rate that can be followed is the whole use of the phenomenon.
A choir has sixteen people on a note. The arithmetic of two tones does not stop applying; it applies a hundred and twenty times at once.
What sixteen singers are, as a signal
Nobody sings a note exactly. Take a section of sixteen with a realistic spread — a standard deviation of fifteen cents, which is a good amateur choir and a slack professional one — and the object is sixteen sinusoidal components clustered within a few tens of cents.
A hundred and twenty rates, all present simultaneously. The obvious question is what that sounds like, and the obvious guesses are both wrong. It is not sixteen audible beats — nobody can attend to a hundred and twenty rates. Nor is it a steadier note, which is what averaging usually does.
The fluctuation does not go away
This is the measurement that decides it, and it needs the right quantity. Counting beat rates is not enough, because what a listener hears is one envelope; the question is whether that envelope has a rate.
Eighty-four per cent at two voices is a beat. One rate holds nearly all of the fluctuation, and a listener can count it, tune to it, and null it. At sixteen voices a quarter of the energy is in the largest component and the rest is spread over the band, and a quarter is not enough to be a rate — the envelope wanders rather than pulses.
So a choir is not steadier than a soloist. It is exactly as unsteady and the unsteadiness has stopped being periodic. That is worth stating in the negative because the intuition that many sources average out to something smooth is a good intuition about most things and is wrong here: the components are near-identical in frequency, so their sum does not converge, it wanders on a timescale of seconds.
That wandering is what the chorus effect is, and it is why every synthesiser has a knob for it. The electronic version detunes copies of one oscillator by a few cents, and it is doing precisely the arithmetic above with a smaller n.
Which means a choir cannot tune by beats
The consequence for intonation is severe and it is a consequence rather than an observation.
Every method of tuning by ear in this collection works by nulling something. A tuner sets a fifth by slowing a beat to a stated rate; a bearing plan is a sequence of such counts; two players with no clock converge by hearing the discrepancy. All of it requires a rate to be present and countable.
A sixteen-voice section has no such rate. There is nothing to null, because the fluctuation is not a fluctuation at one frequency. A singer inside the section can hear their own contribution against the mass — which is a level and a roughness cue, not a beat — and can hear the section against another section, where the interval is large enough that the beating is between partials rather than fundamentals and is again a mass rather than a line.
This is the mechanism behind something choral directors say and rarely justify: that a section tunes by blend rather than by beats, and that the instruction to a section is different in kind from the instruction to a soloist. The physical difference is that one of them has a countable rate available and the other does not.
The correction it forces
There is a result in this collection that has to be re-read in the light of this, and it is one of the better ones.
The comma-pump essay shows that an ensemble tuning every chord in pure ratios cannot come back to where it started, and cites the published measurements of a cappella groups drifting downward over a piece by a comma or more. Those measurements are real and the arithmetic behind them is exact.
What has to be added is the condition under which they apply. A group that drifts by tuning each chord pure is a group that can tell when a chord is pure, which means a group with a countable beat rate available, which means a group whose pitch is steady enough for the beat to survive. Every measured a cappella drift result is from an ensemble of that kind — small, professional, and using little or no vibrato.
The other rung of this ladder makes the same point from the opposite side: an operatic vibrato of 142 cents peak to peak sweeps the beat rate through zero several times a second, so a vibrato singer has no stable beat either. Between them the two rungs bound the result: pure-intonation drift is a property of ensembles that suppress pitch variation, whether the variation comes from vibrato in one voice or from spread across many.
That is not a criticism of the drift result. It is the specification of its domain, which the result did not state and which this rung owed it.
Where the spread comes from, and why it will not go away
It is worth asking whether a choir could simply be more accurate, and the answer is that the spread has three sources and only one of them is a matter of skill.
Placement error. A singer aims at a pitch and misses by some amount. This is the part training reduces, and it can be reduced a long way — a good professional section is a few cents rather than fifteen.
Vibrato. A trained solo voice varies by ±50 to ±100 cents at five or six hertz. Sixteen such voices with unrelated vibrato phases are sixteen components each sweeping a band far wider than the placement spread, and no amount of accuracy removes it. Choral training that asks for straight tone is asking singers to disable the mechanism, which is why it is contentious.
Physical difference. No two tracts are the same length, so no two singers produce the same spectrum on the same fundamental even when the fundamentals agree exactly — which is the previous rung’s whole subject. The partials above the first therefore never coincide, and the beating between them is not removable by tuning at all.
The last of these is the interesting one, because it means a choir singing a perfect unison in fundamental would still produce a chorus effect from its upper partials. The spread is not a defect being tolerated; a substantial part of it is the difference between people.
Twenty singers are not twenty times one
The other half of what a choir does is level, and the arithmetic is the same arithmetic with a different consequence.
Three decibels a doubling is a modest return: sixteen singers are twelve decibels above one, which is loud but is not sixteen times anything. If the point of a chorus were volume it would be an inefficient way to get it.
That the return is so modest is the argument that volume is not the point. A section of sixteen exists because sixteen slightly different sources produce a sound one source cannot produce at any level — the aperiodic envelope of the figures above — and the twelve decibels are incidental.
The same arithmetic runs through the rest of this collection wherever sources are multiplied. A piano’s unison strings are two or three sources deliberately mistuned, for exactly this reason at exactly this scale, and the piano tuner’s decision about how much to mistune them is a decision about where on the curve between “one beat” and “a chorus” the note should sit.
What a section sounds like against another section
Everything above concerns one note. A choir spends nearly all its time on chords, and the arithmetic changes in a way worth stating even though this rung does not develop it.
A fifteen-cent spread is a fiftieth of a semitone, and the critical bandwidth is between one and three semitones across the whole register — so a choir’s placement spread sits far inside a single band at every pitch anybody sings. Whatever a section is doing to the sound, it is not being done by separating the voices into different channels.
So a choir has two regimes at once and they are not the same phenomenon. Within a section, sixteen components inside one critical band, producing a slow aperiodic envelope. Between sections, partials from thirty-two sources meeting across bands, producing roughness whose amount depends on the interval and on how spread each section is.
The obvious thing to expect is that a spread section is less rough than a precise one, because a spread of pitches spreads the coincidences too and no single pair of partials sits at the worst separation. Measuring it says that expectation is backwards for every interval a choir sings in tune.
Two sections of sixteen, each voice a complex tone, every cross pair scored with this collection’s own roughness model and averaged over forty draws:
| interval between the sections | spread 0 | 15 cents | 30 cents | change |
|---|---|---|---|---|
| unison | 0.0045 | 0.087 | 0.139 | thirty times rougher |
| octave | 0.0045 | 0.056 | 0.083 | eighteen times |
| fifth | 0.053 | 0.068 | 0.076 | +45% |
| fourth | 0.101 | 0.109 | 0.113 | +12% |
| major third | 0.145 | 0.147 | 0.148 | +2% |
| minor third | 0.193 | 0.192 | 0.193 | 0% |
| tritone | 0.101 | 0.101 | 0.099 | −2% |
| semitone | 0.260 | 0.256 | 0.239 | −8% |
The sign of the effect is decided by which side of the dissonance curve the interval sits on, and the reasoning that produced the wrong expectation is the right reasoning applied at the wrong place. Averaging a quantity over a spread of inputs moves it towards the average of its neighbourhood. A consonance is a minimum — the well the whole consonance ladder is about — so its neighbourhood is higher than it is and spreading can only climb. A dissonance sits near a maximum, so spreading falls. The scattering argument is sound on a peak and reversed in a well, and a choir spends its time in wells.
The sizes make the point rather than the signs. On the dissonant intervals the gain is a few per cent, which is nothing; on the unison and the octave the loss is more than an order of magnitude, because two identical complex tones are not rough at all and a spread one is. A section’s spread costs it precisely in the intervals it is most exposed on, and that is a prediction the model makes rather than an observation borrowed from anywhere: the unison and the octave are the two intervals a choral director is heard complaining about, and they are the top two rows.
It also puts a boundary on the previous section. Between sections, roughness is what the interaction is; and the interaction gets worse with spread nearly everywhere, while within a section the same spread is the entire point. The spread is not one thing being traded against another — it is one thing helping in one regime and hurting in the other.
That is a useful place to mark, because it corresponds to something musicians already treat as a boundary. One to a part is chamber music and is tuned as chamber music: by ear, by beats, chord by chord. Several to a part is choral and is tuned by blend. The change of method is usually explained by the difficulty of coordinating more people, and the figures here suggest a simpler reason — past about four, the signal a small ensemble tunes by is no longer there.
Which computation produced the numbers
The mistunings are drawn from a normal distribution by Box–Muller from a fixed seed, so every figure here is the same choir every time it is drawn. The frequencies are the nominal fundamental multiplied by two to the power of the drawn cents over 1200.
The beat rates are all pairwise differences: sixteen voices give 120 pairs, thirty-two give 496.
The envelope is the magnitude of the sum of the components as unit phasors, sampled at 400 hertz over two seconds. Its modulation spectrum is a direct transform up to thirty hertz — past which a fluctuation is heard as roughness rather than as a beat, which is the consonance ladder’s own boundary and not a new one. Concentration is the largest bin’s share of the total modulation power.
An earlier version of this measurement scored periodicity by autocorrelation and returned a number that did not move with the number of voices at all. A slowly varying envelope correlates well with itself at short lag whatever it is doing, so the peak sat on the search floor at every size and the figure was measuring the floor. The modulation spectrum asks the question directly, and the fact that the first attempt produced a plausible flat line is the reason it is recorded here.
The level curves are 10 log₁₀ n and 20 log₁₀ n, which are the incoherent and coherent cases and involve no acoustics beyond the choice of which applies.
Whose choirs, and when
The fifteen-cent spread is a stated assumption, not a measurement, and it is the parameter everything else scales with. Published measurements of choral unison spread vary from a few cents in a professional chamber choir to well over thirty in a congregation, and the shape of the argument is unchanged across that range — a wider spread moves the beat rates up and the concentration down, which makes the effect stronger rather than different.
The claim about drift and vibrato applies to the specific published a cappella studies, which are mostly European art-music ensembles of four to eight singers recorded in the last forty years. Traditions with large unison singing and deliberate heterophony — Georgian polyphony, Balkan village singing, congregational psalmody, Sardinian cantu a tenore — are a different object, and the wide unison is often the aesthetic point rather than an imprecision to be minimised.
What the picture cannot show
Every voice here is a single sinusoid. A real singer produces a whole spectrum, and the beating between two singers happens at every partial at once, faster at each higher one because the same cents difference is more hertz. That makes the real modulation spectrum broader than the one drawn and strengthens the argument; it also means the depth figures here are a floor.
The mistunings do not move. A real choir’s pitches wander continuously — that is most of what makes the sound — and the model draws them once and holds them. A model with drifting components would produce an envelope with even less periodicity than this one shows.
It has no onsets. Sixteen singers do not begin a note together — the spread in attack time is tens of milliseconds — and the first fifty milliseconds of a note are most of what identifies it. A model that starts every component at once has removed one of the largest differences between a section and a soloist before it begins.
And there is no room. A room adds its own decorrelation between sources arriving from different positions, and the reverberant field of a large church does a substantial part of what this essay attributes to the spread. Separating the two would need a measurement in an anechoic chamber, and choirs are not usually found in them.
The ladder from here
The two rungs owed when this anchor opened are now written, and between them they have changed what an earlier result means rather than only adding to it.
What the ladder still lacks is the acquisition question — a listener’s ability to normalise for an unfamiliar tract, which appears very early and is not obviously learnable from what an infant hears — and the ensemble question one level up: what a section does against another section, where the interval is large and the beating is between partials.
The second is now opened rather than owed. The table above gives the sign and the size of what spread does to each interval, and it does so with two sections held at a fixed distance; what it does not do is let the sections move. A real choir’s two sections negotiate their interval while each is spreading, and the quantity that decides where they settle is the gradient of that table rather than its values. A section sitting in a well has a restoring force towards the pure interval and a section sitting on a slope has a push along it, so the same measurement run as a derivative would say which intervals a choir is driven towards and which it merely occupies — which is the shape of the rung after this one, and it needs no new machinery.
Part 6 of 13
One essay in the series on the voice. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 19.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
BeatingChorus effectIncoherent additionJust intonationModulationUnisonVibrato
- A beat has a depth, and six essays held it at one beating, unison
- A consensus with nothing to hold it beating, just intonation
- A roughness with a rate of its own beating, vibrato
- An interval is two errors beating, vibrato
- An orchestra is given a note beating, just intonation
- Counting beats moves the price of a chord, not the tuning beating, just intonation