Pitch and tuning

A note that is never at its pitch

Eleven essays in this field argue about differences of one to twenty-four cents. An ordinary operatic vibrato is a hundred and forty cents wide and completes six excursions a second, so every one of those distinctions fits inside a single sung note several times over — and the beat rate a tuner would null passes through zero eleven times a second.

Assumes: The other instrument with a reed · How small a difference is audible

This field has eleven essays in it and every one of them is an argument about a small number of cents. Twelve fifths miss seven octaves by 23.5. The syntonic comma is 21.5. A quarter-comma meantone fifth is 3.4 cents narrow of equal. The schisma is under two, which was the point of the essay it appears in.

Not one of those essays has drawn the width of the note the difference is supposed to be heard on.

A hundred and forty cents

A note sung at 440 Hz, drawn in centsDeviation from the notated pitch against time, for a vibrato of ±71 cents at 6.0 cycles a second — Seashore's and Prame's measured values, which agree. The note is 142 cents wide, and the shaded bands across the middle are the syntonic comma at 21.5 cents, the Pythagorean comma at 23.5 cents, the difference limen at 440 Hz at 4.0 cents. The pitch a listener reports is near the middle of the excursion rather than at either edge — and not exactly at the middle either: because cents are logarithmic and frequency is not, the mean frequency of this trace sits 0.73 cents above the centre line.the syntonic comma21.5 centsthe Pythagorean comma23.5 centsthe difference limen at 440 Hz4.0 centsthe note itself142 cents widemean 0.73 cents sharpof the centre line0.00.20.40.60.81.0-50050secondscents from the notated pitch
Fig. 1 A sung note, drawn on the axis this field uses. The trace is the pitch against time for an ordinary vibrato — plus and minus seventy-one cents at six cycles a second, which are the measured central values — and the shaded bands across the middle are the syntonic comma, the Pythagorean comma and the difference limen at this pitch. Every one of them is inside the trace. The buttons play the note steady and with the vibrato on it, and then a note a full Pythagorean comma sharp, steady, for comparison.

The numbers are Seashore’s, from the 1930s, and Prame’s, from 1994. Seashore measured commercial recordings with a photographic pitch-tracking apparatus; Prame re-measured a thousand and eighteen tones from ten tenors singing the same Schubert song. The two agree closely, which is worth noting because almost nothing else about singing has been measured twice sixty years apart with the same answer.

The rate is between five and a half and seven and a half cycles a second, and it is remarkably fixed: it does not follow the tempo, it does not follow the pitch, and singers cannot vary it much on request. The section on what the pitch of such a note is finds where that band sits in the range the ear allows, and it is at the very bottom of it — which is a fact about what a vibrato is for rather than about what the larynx can do. The extent is between about thirty-four and a hundred and twenty-three cents either side of centre, with a mean near seventy. So the note is between seventy and two hundred and fifty cents wide, and the ordinary case is about a hundred and forty — considerably more than a semitone.

Every quantity in the field, on one axis

Everything this field argues about, and the width of one sung note. Eight quantities in cents, on one linear axis. Seven of them are distinctions the tuning essays are about, from a beat once in two seconds, on a fifth at 1.31 cents to the Pythagorean comma at 23.5. The eighth is the peak-to-peak width of a single note sung with an ordinary vibrato of ±71 cents, which is 142 cents — 6.1 times the Pythagorean comma and 35 times the difference limen at 440 Hz. Every bar above the axis is smaller than the note it would be measured on.
Fig. 2 Eight quantities in cents. Seven are distinctions the tuning essays are about; the eighth is the peak-to-peak width of one sung note. The largest of the seven is the Pythagorean comma, and the note is six times wider than it.

That is the whole observation, and the rest of this essay is about what does and does not follow from it. Two things follow immediately, and both are stronger than they look.

A vibrato tone cannot be tuned by beats

The tuner’s method is not to listen for pitch at all. It is to listen to two notes together, hear the beat between a partial of one and a partial of the other, and adjust until the beat is at the rate the temperament asks for — usually until it stops.

That method is exact. It is what makes tuning by ear possible at all, and it is why a fifth can be set to a third of a cent while two successive notes cannot be distinguished at four. But it needs both notes to hold still.

What a tuner would have to null, on a note with vibrato. The instantaneous beat rate between the third partial of a steady 440 Hz note and the second partial of a fifth above it, where the upper note carries a vibrato of ±71 cents at 6.0 cycles a second. The rate runs from −55.1 to 55.1 beats a second and passes through zero 11 times in 1.0 second. A tuner's method is to adjust until the beat stops; here it stops 11 times a second whatever the tuning is, and the excursion either side is 37 times the 1.5 beats a second an equal-tempered fifth produces on this note. At the extremes the rate is past 55 a second, which is not a beat any more — it is a rate the ear reads as roughness. Beat-counting is a method for held tones, and a sung tone is not one.
Fig. 3 The beat rate between the third partial of a steady note and the second partial of a fifth above it, with a vibrato on the upper note. It sweeps from more than fifty beats a second in one direction to more than fifty in the other, crossing zero eleven times in a second. There is nothing here to null; the beat is at every rate, including the right one, in every cycle.

The reason the excursion is so violent is that beats live in hertz and vibrato lives in cents. Seventy cents on a note at 660 hertz is about twenty-seven hertz of deviation on that partial, and the beat rate is a difference of frequencies, so the whole seventy cents shows up at full size in the beat. At the extremes the rate is past fifty a second, which is no longer a beat at all but a roughness.

So the entire apparatus by which keyboards, harps and organs are put in tune is unavailable to a singer with a vibrato, and the question of what a vibrato singer’s intonation even means has to be asked before it can be measured.

What the pitch of such a note is

It has an answer and the answer is not obvious a priori. A note that spends its time everywhere between minus seventy and plus seventy cents could plausibly be heard at its highest point, at its lowest, at its mean, or as no definite pitch at all.

What listeners report is a single, definite pitch, near the middle of the excursion. Sundberg’s matching experiments put it within a few cents of the mean, and the effect is robust: a vibrato tone is not heard as a smear or a trill but as one note, with the vibrato heard as a quality of the note rather than as a movement of it.

That is a fact about the ear rather than about singing, and the usual explanation for it — that six hertz is too fast for the pitch mechanism to follow and too slow to turn the tone into a chord of sidebands — has two boundaries in it, neither of which is where the explanation needs them.

The upper boundary is a long way off. Frequency modulation puts sidebands at plus and minus the modulation rate, and they are heard as separate components only when they fall outside one critical band. At 440 hertz that band is 72 hertz wide, so the vibrato rate would have to be twelve times the measured one before a sung note became a chord. That end of the argument is safe by an order of magnitude.

The lower boundary is not there at all. A half cycle of a six-hertz vibrato lasts 83 milliseconds, and a tone of that length can have its frequency specified to about 23 cents — against an excursion of 142. So a listener has, in principle, six times the resolution needed to follow the excursion, and would not lose that until the rate reached about 38 hertz.

So the ear is not averaging because it cannot resolve. It is averaging while able to resolve, which makes the single-pitch percept a grouping decision rather than a resolution limit — the same kind of decision that makes a set of partials one note rather than several, taken over time instead of over frequency. A listener hears one note with a quality because a coherent six-hertz modulation is exactly the evidence that one thing is making the sound, and hearing it as a pitch going up and down would be hearing two facts where there is one.

That reading also explains why the measured rate sits where it does. The averaging régime runs from about six hertz to about seventy, and singers use the very bottom of it — as slow as a vibrato can be while still fusing, which is where it is most nearly audible as movement without becoming movement. A vibrato at twenty hertz would fuse just as completely and would be inaudible as an expressive gesture, which is presumably why nobody produces one.

There is a small arithmetical curiosity inside the averaging, and it is worth computing because it is the kind of thing that gets asserted without one.

Cents are logarithmic and frequency is not. A trace that is symmetric in cents is not symmetric in hertz: the sharp half of each cycle spends its time at frequencies further above the centre than the flat half spends below it. So the mean frequency of a vibrato tone is above its centre, and the amount is computable. For a sinusoidal vibrato of ±71 cents it is 0.73 cents; at ±123 cents it is 2.2.

That is smaller than the schisma and comfortably inside the noise of any measurement of where a singer’s vibrato is centred — and it is worth noticing that it is also smaller than the offset the same arithmetic gives for the widest measured vibrato, which is 2.2 cents and is still under the limen. So the curiosity is negligible across the whole measured range and not merely at its middle, which is the form the check had to take: a quantity that is small at one point in a range is not established as small. It is drawn on the first figure as a line above the centre line, and its whole significance is that it is negligible — which is worth establishing rather than assuming, because the reasoning that produces it would produce a large number for a wider modulation.

A note sung at 440 Hz, drawn in cents. Deviation from the notated pitch against time, for a vibrato of ±34 cents at 7.5 cycles a second — Seashore's and Prame's measured values, which agree. The note is 68 cents wide, and the shaded band across the middle is the Pythagorean comma at 23.5 cents. The pitch a listener reports is near the middle of the excursion rather than at either edge — and not exactly at the middle either: because cents are logarithmic and frequency is not, the mean frequency of this trace sits 0.17 cents above the centre line.
Fig. 4 The narrowest vibrato in Prame’s measured set — thirty-four cents either side, at seven and a half cycles a second — against the Pythagorean comma alone. Even at the bottom of the measured range the note is nearly three times the width of the discrepancy that this whole field is named after.

What this settles, and what it does not

The temptation is to conclude that intonation is meaningless for a vibrato voice. That is wrong, and the reason it is wrong is the same averaging.

The smallest audible difference, and what has to clear it. The difference limen for frequency, converted from Wier, Jesteadt and Green's 1977 fit into cents, against the intervals and commas the rest of these essays argue about. Anything drawn below the curve is a quantity nobody can hear as a change of pitch; anything well above it is a quantity a listener can be asked about. The limen is for pure tones, successive, with trained listeners — the most favourable case there is, and therefore the right one to test a claim against.
Fig. 5 The difference limen for successive tones across the range, with the largest and smallest of this field’s own quantities drawn against it. The vibrato is not on this axis: at 142 cents it is six times the top line here and would flatten everything below it. The limen is a few cents, the note is a hundred and forty cents wide, and yet the centre of that note can be placed by a singer and detected by a listener to within about ten cents, because both are working with the average rather than with the excursion.

So a singer can be flat by a comma, and it can be heard, and it means what it usually means. The vibrato does not destroy the pitch — it destroys the method. What is lost is the ability to tune by nulling a beat, and what is left is the ability to place a mean by ear against a remembered or sounded reference, which is the coarser of the three resolutions this site has measured rather than the finest.

That distinction turns out to be the whole of the disagreement in the literature on how ensembles tune.

Which is why the choir results are about a particular kind of choir

Somebody has to pay the comma works out what happens to an ensemble with no frets and no keys: it can tune every chord exactly, and a common progression sung that way arrives a comma flat every circuit. The freedom to tune each chord pure is what makes the drift.

That freedom rests on the beat method. An ensemble tunes a chord pure by adjusting until the beating between shared partials stops, which is the same procedure a tuner uses with more people. And the figure above says that a singer with an ordinary operatic vibrato cannot do it, because there is no beat to stop.

The period is still there, and it is wider. The autocorrelation of a 12-partial complex on 220 hertz, drawn twice: steady, and averaged over one cycle of a 71-cent vibrato. A vibrato moves every partial by the same number of cents, so the complex is exactly harmonic at every instant and nothing is mistuned — what moves is the period the extractor is looking for. The peak survives. It loses 6 per cent of its height above the surrounding lags and gains 11 per cent in width, because the vibrato swings the period by 0.37 milliseconds against a peak 0.90 wide. Its maximum also moves, to 2.4 cents sharp of the still tone's, which is a prediction with a sign in it.
Fig. 6 The autocorrelation of a twelve-partial complex on 220 hertz, drawn twice: steady, and averaged over one cycle of a seventy-one-cent vibrato. A vibrato moves every partial by the same number of cents, so the complex stays a complex.

The method a tuner uses needs two steady tones and an envelope that rises and falls at a countable rate. A vibrato tone has no steady partial to beat against, and this is why: the peak the pitch is read from is broadened rather than displaced, so the note keeps a pitch and loses the thing a beat is made of.

The conclusion is not that the comma-drift results are wrong. It is that they describe choirs that do not vibrate — which is exactly the repertoire the measurements were made on. The published studies of a cappella drift are overwhelmingly of close-harmony, barbershop and early-music ensembles, and all three traditions suppress vibrato explicitly, for reasons their own practitioners describe in terms of the chords locking.

An operatic chorus is a different object. Its members are each a hundred and forty cents wide, they cannot lock, and nobody claims they do.

The same argument, on instruments

Vibrato is not confined to voices and the numbers differ, which is a useful control.

A violinist’s vibrato is narrower — typically ±20 to ±40 cents in orchestral playing, wider in solo playing — and it is produced by rocking the finger, which shortens and lengthens the string. Narrower means the argument above is weaker but not absent: forty cents is still nearly twice the Pythagorean comma.

A wind player’s vibrato is usually an amplitude and airflow modulation with a smaller frequency component, because the pitch is held by a standing wave in a tube that does not care how hard it is being driven. A flute’s vibrato moves the pitch by a few cents and the loudness by several decibels.

And a keyboard has none at all, which is the reason this entire field exists. The whole of the temperament literature is about instruments whose pitches are fixed before the music starts, and the two facts — a fixed pitch and an audible comma — are the same fact.

The two things a wider vibrato does to octave. Against how wide the vibrato is: the depth of the roughness fluctuation, peak to peak over its own mean, and the ratio of the moving average to the still value. The depth rises steadily and reaches 1.18 at 200 cents, so an operatic vibrato makes the roughness of a sustained interval swing by most of its own size. The mean goes from 1.00 at no vibrato, which is one by definition, to 24.30 — so on an interval sitting in a deep minimum the shift is the larger of the two effects and the depth is the more obvious one. Which of the two a listener uses is the question this figure cannot answer.
Fig. 7 Against how wide the vibrato is: the depth of the roughness fluctuation, peak to peak over its own mean, and the ratio of the moving average to the still value.

The depth rises steadily and the mean barely moves, which is this ladder’s recurring shape. Set beside the thirds of four temperaments — spread over a few cents — the vibrato is an order of magnitude wider than every tuning distinction it is supposed to be obscuring, and yet the mean survives it.

The listener’s tolerance, which is the same width

There is a coincidence here worth putting a number on, because this collection has already measured the other half of it.

The ear sorts intervals into boxes, and the boxes are wide. Identification of a melodic interval swings from one name to the next over about thirty cents for trained listeners, and the centre of each box is where the name is, so an interval can be forty or fifty cents off its nominal size and still be identified without hesitation as what it is.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 8 Identification curves for three adjacent intervals. What matters here is not where the boundaries are but how much room there is between them: an interval sitting anywhere in the flat part of a curve gets the same name with the same confidence, and the flat parts are most of the axis.

That is the reason temperament works at all: a major third fourteen cents wide of just is still, categorically, a major third, and the whole of equal temperament trades on it.

So the production side and the perception side have arrived at the same width from opposite directions. A voice produces a note a hundred and forty cents wide; an ear assigns names over boxes about a hundred cents apart with transitions of thirty. Neither is a consequence of the other and there is no reason to expect them to match. What follows is a workable practice: singers can be wide, listeners can be tolerant, and the pitch that is agreed on is an average on one side and a category on the other.

The simple ratios against the twelve equal steps put the largest disagreement at about sixteen cents, on the thirds. Every quantity in this essay is larger than that: an ordinary vibrato is ±70 cents, so the note sweeps past all four of the historical temperaments’ thirds several times a second and arrives, somehow, at one pitch.

Whose singing, and when

Continuous vibrato is not a fact about the human voice. It is a practice, it has a date, and the date is later than most of the repertoire.

The evidence for what earlier singing sounded like is documentary and it is contested, but it is consistent on one point: the writers who mention a tremolo or trembling of the voice before about 1800 describe it as an ornament applied to particular notes rather than as a permanent condition of the tone. Mersenne, in 1636, treats it as a device. Geminiani, in 1751, recommends it on long notes. Mozart complained in a letter that a particular singer’s tone “trembles”, and he complained about it as a fault — which requires that the alternative was available.

What changed over the nineteenth century is the same thing that changed in the previous rung: the orchestra grew, the halls grew, and the technique that carried won. A vibrato helps a voice carry — it is a form of movement in a texture that is otherwise steady, and the ear attends to movement — and the same century that produced the singer’s formant produced continuous vibrato.

The historical-performance movement of the twentieth century reversed it deliberately, and the reversal is why measurements of a cappella just intonation exist at all: it takes an ensemble that has decided not to vibrate before there is anything to measure.

What the picture cannot show

A real vibrato is not a sine wave. The traces here are sinusoidal because the arithmetic is then exact and reproducible. Measured vibratos are approximately sinusoidal in the middle of a note and are not at all sinusoidal at its beginning and end — the onset typically has no vibrato for the first tenth of a second or more, and the extent grows.

Frequency is not the only thing modulating. A sung vibrato moves the loudness as well, by several decibels, and it moves the spectrum, because the partials are sliding under formants that do not move and each partial’s amplitude therefore rises and falls as it crosses a resonance. Part of what is heard as vibrato is that amplitude pattern, and it is not in any of these drawings.

And the measured extent depends on the analysis window. A pitch tracker with a window of a hundred milliseconds is averaging over most of a vibrato cycle and will report a smaller extent than one with a window of twenty. The published values are not always comparable for that reason, and the range from thirty-four to a hundred and twenty-three cents is partly a range of methods.

The beat calculation assumes both notes are complex tones with real partials at the ratios drawn. For voices they are, and for the specific partials named — a third and a second — they are strong. For a pair of near-sinusoidal M2 productions there is much less to beat with in the first place.

Where this ladder has got to

Four rungs, and none of them is about tone in the sense the word usually carries. The source is a valve with a closed phase; the seam in the middle of the range is a bifurcation with hysteresis; the carrying power is a peak in a band the orchestra has vacated; and the note is a hundred and forty cents wide.

What is left undone is the thing every other field here has: a measurement of what a listener does with all of it. The voice is the sound humans are best at identifying, most attentive to and most tolerant of, and none of those three has been given a number on this site. That is where the ladder goes.

Part 4 of 13

One essay in the series on the voice. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 15.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

BeatingCentsIntonationJust-noticeable differencePitch perceptionTemperamentTuning by earVibrato