Instruments and their design

The pitch that does not wobble

Three earlier essays have treated a vibrato as a modulation of roughness. The reason singers use one is what it does to the note, and there is an extractor here that turns a set of partials into a pitch and has never been asked what it does with partials that will not hold still. The period survives, at a cost that rises with the extent — and the practice stops within a hair of where the cost becomes total.

Assumes: The mean survives the window · A note that is never at its pitch

The mean survives the window ended the ninth and tenth rungs’ business with roughness and said what all three of them had left out:

Everything in the last three rungs treats a vibrato as a modulation of a sensory quantity, and the reason singers use one is that it does something to the note — a listener hears a steady pitch, more presence, and a voice that carries through a texture.

The first of those three is the one this collection can compute. The pitch that moves the wrong distance and the rungs around it built an account of how a pitch is extracted from a set of partials — a period, found by asking at what lag the waveform most resembles itself — and every one of them handed it a signal that was standing still.

The period is still there, and it is wider. The autocorrelation of a 12-partial complex on 220 hertz, drawn twice: steady, and averaged over one cycle of a 71-cent vibrato. A vibrato moves every partial by the same number of cents, so the complex is exactly harmonic at every instant and nothing is mistuned — what moves is the period the extractor is looking for. The peak survives. It loses 6 per cent of its height above the surrounding lags and gains 11 per cent in width, because the vibrato swings the period by 0.37 milliseconds against a peak 0.90 wide. Its maximum also moves, to 2.4 cents sharp of the still tone's, which is a prediction with a sign in it.
Fig. 1 The autocorrelation of a twelve-partial complex on A3, steady and with a vibrato of seventy-one cents either side, averaged over one cycle. The peak is still there; it is six per cent shorter and eleven per cent wider.

Nothing is mistuned

The first thing to be clear about is what a vibrato is not.

A vibrato modulates the fundamental, so every partial moves by the same number of cents. At every instant the complex is exactly harmonic: the ratios between the partials never change, no partial is ever mistuned against any other, and every cue this collection has for whether a set of components belongs to one sound says yes at every moment.

So the difficulty is not identification. It is that the answer will not stay put. A pitch mechanism that looks for a lag at which the waveform repeats has a target that is moving, at six hertz, over a hundred and forty cents peak to peak — which is a swing wider than any interval this collection calls a comma.

Two widths, and the whole argument is their ratio

The autocorrelation of a steady harmonic complex has a peak at the period, and that peak has a width. It is set by the highest partial carrying appreciable energy: a complex with many strong partials has a sharp waveform and a sharp peak, and one with three has a broad one. For a string-like spectrum on A3 the peak is 0.90 milliseconds wide at half height, against a period of 4.55.

A vibrato moves the period. A swing of ±71 cents is a fractional frequency swing of ±4.2 per cent, so the period swings by 0.37 milliseconds.

Those two numbers are the whole of it. 0.37 against 0.90: the vibrato smears the peak by less than half its own width, so averaging over a cycle leaves a peak that is slightly lower and slightly broader and unmistakably in the same place.

A vibrato is a large interval and a small smear, and the reason it is both is that a pitch mechanism is not an interval mechanism.

What it costs, and where the practice stops

Sweeping the extent turns that into a curve.

What a vibrato costs the pitch, and where singers stop. How much of the autocorrelation peak survives, against how wide the vibrato is. The shaded band is the measured extent of operatic vibrato — 34 to 123 cents either side, across ten tenors and a thousand tones — and inside it the peak keeps between 85 and 98 per cent of its salience. The faint curve is the swing the vibrato puts on the period, divided by the peak's own width: it passes one at about 200 cents, which is a little above the widest vibrato anybody was measured singing. So the practice occupies the region in which a period-based pitch mechanism is degraded and not defeated, and stops close to where it would be.
Fig. 2 How much of the autocorrelation peak survives, against how wide the vibrato is. The shaded band is what ten tenors were measured singing.

The peak keeps 98.5 per cent of its salience at 34 cents, 93.9 at 71, and 84.5 at 123. Those three numbers are the bottom, the middle and the top of the range Prame measured across a thousand tones, and the fourth rung of this ladder is where they came into the collection.

Above that band it falls away fast: 70.5 per cent at 200 cents, 46.3 at 340. And the period swing passes the peak’s own width at about 200 — a little above the widest vibrato anybody in that study was recorded singing.

So the measured practice occupies exactly the region in which a period mechanism is degraded and not defeated, and stops close to where it would be. That is the kind of coincidence worth stating carefully. It is not evidence that singers are optimising a pitch extractor; nobody is doing that. It is that a vibrato which destroyed the pitch would not be a vibrato, it would be a bad note, and the boundary between the two is computable and turns out to be where the practice’s own edge is.

The vocabulary agrees. A vibrato past that boundary has a name — a wobble — and what everybody says about it is that the note stops having a definite pitch. The model says the peak has stopped being narrower than the swing.

And the register drops out entirely

The obvious next question is whether it is harder low or high, and the answer is neither, exactly.

The pitch account has no register in it, and the roughness account does. The same 71-cent vibrato at seven pitches, and how much of the autocorrelation peak survives at each. The line is flat to 2e-16 across three octaves, and it has to be: the period, the peak's width and the swing the vibrato puts on it are all proportional to one over the frequency, so the ratio between them has no pitch in it. That is the opposite of what the same model gave for roughness, where the critical band is a width in hertz and therefore a scale, and the cost of a vibrato changes by a large factor up the compass. Two accounts of one signal: the sensory one has a register and the pitch one does not.
Fig. 3 The same vibrato at seven pitches over three octaves. The line is flat, and it is flat for a reason rather than by accident.

The period goes as one over the frequency. The peak’s width goes as one over the frequency, because it is set by the highest partial and the partials scale with the fundamental. The vibrato’s swing on the period goes as one over the frequency, because the extent is in cents.

Three quantities all proportional to the same thing, and the answer is a ratio of two of them. There is no pitch in it. A seventy-one-cent vibrato costs the same six per cent of the pitch’s salience on a bass’s low A as on a soprano’s high one.

That is worth setting against what the tenth rung found for the same signal. Roughness has a scale in it — the critical band is a width in hertz, not in cents — so the cost of a vibrato in roughness changes by a large factor up the compass, and that rung drew the change.

Two accounts of one physical signal, one of which has a register and one of which cannot. A singer moving up an octave is changing what their vibrato does to the sensory dissonance of the texture and changing nothing at all about what it does to their own pitch.

The window is a function of the register, so the survival is too. The same interval taken up the compass. Roughness fluctuates faster in a higher register — 42 hertz at the bottom of this range and 227 at the top — and the window a listener integrates it over is a fixed number of its cycles, so the window shortens from 96 milliseconds to 18. A short window follows a vibrato and a long one flattens it, so the share of the fluctuation that survives rises from 47 per cent to 58. That is a prediction about where the moving roughness of a vibrato is audible as movement, and it says high rather than low.
Fig. 4 The earlier register figure, for comparison: the roughness account of the same signal, which does move up the compass because the critical band is a bandwidth rather than an interval.

A prediction with a sign in it

The averaged peak is not only lower and wider. Its maximum has moved, and the direction is not the obvious one.

The natural guess is flat. A vibrato is symmetric in cents, so it spends equal time above and below; a period is one over a frequency, and the average of one over a symmetric swing is longer than one over the middle. That is the same Jensen argument the ninth rung used on roughness, and it points to a pitch slightly below the mean.

The computation gives sharp: 2.4 cents at typical extent, 9.1 at the top of the measured range, 36 at 200 cents.

The mechanism is that the smear is proportional to the lag. Each instant’s autocorrelation contributes a peak at its own period, and the contributions at longer lags are individually broader — so averaging erodes the long-lag side of the peak more than the short-lag side, and the maximum slides toward the short end. The asymmetry is in the instrument, not in the signal.

Two and a half cents is a small quantity and it is not an unmeasurable one: it is inside what a listener can do comparing two tones directly, and the effect is predicted to grow with the extent, which makes it a slope rather than an offset. The perceived pitch of a vibrato tone should be a few cents sharp of its mean, and sharper the wider the vibrato.

That is the one claim on this page that could be shown to be wrong by a listener rather than by a better model.

What a tuner would have to null, on a note with vibrato. The instantaneous beat rate between the third partial of a steady 440 Hz note and the second partial of a fifth above it, where the upper note carries a vibrato of ±71 cents at 6.0 cycles a second. The rate runs from −55.1 to 55.1 beats a second and passes through zero 11 times in 1.0 second. A tuner's method is to adjust until the beat stops; here it stops 11 times a second whatever the tuning is, and the excursion either side is 37 times the 1.5 beats a second an equal-tempered fifth produces on this note. At the extremes the rate is past 55 a second, which is not a beat any more — it is a rate the ear reads as roughness. Beat-counting is a method for held tones, and a sung tone is not one.
Fig. 5 The instantaneous beat rate between the third partial of a steady note and the second partial of a fifth above it, where the upper note carries a vibrato of ±71 cents at six cycles a second.

The rate runs through zero and out the other side, which is the sharpest form of what a vibrato does to a coincidence: the two partials are not slightly mistuned, they are alternately sharp and flat of each other. Anything that averages a beat rate over a vibrato cycle is averaging a signed quantity, and the mean is not what a listener hears.

The extractor’s own choice matters, and it is visible here

There is a second reading of the sharp bias and it belongs in the open rather than in the caveats.

Two mechanisms are on offer in the literature for how a set of partials becomes a pitch, and this collection has drawn both. One is periodicity: find the lag at which the waveform repeats, which is what is computed above. The other is a harmonic template: slide a comb of whole-number multiples along the frequency axis and find where it fits.

On a steady complex the two agree, which is why the argument between them has needed carefully constructed stimuli — a shifted complex, whose partials are equally spaced but not multiples of anything, on which the two accounts give different answers and listeners side with periodicity.

A vibrato is a stimulus of the same kind and nobody has used it that way. A template does not care that the whole comb is sliding: it fits perfectly at every instant, and averaging its output over a cycle gives the mean with no loss of salience and no bias. Periodicity loses six per cent of its salience and gains two and a half cents of sharpness.

So the two accounts differ by a measurable amount on the commonest signal in Western singing, and the difference is in the direction of a quantity somebody could put a slider on.

Why a singer would want this

None of the above says a vibrato is good for anything. It says the pitch survives one, which is a licence rather than a reason.

The reason the ladder can now assemble is a negative one and it is worth having. A voice with a vibrato has, at every instant, a spectrum spread over a hundred and forty cents; two voices with independent vibratos do not beat in any steady way; a section of them fills a band rather than a line. All of that is what makes a choral sound and it is bought with the frequency axis.

What it is not bought with is pitch definiteness. The one thing a singer cannot afford to lose is the note, and this figure says the note is the cheapest thing in the transaction — six per cent of a salience, at an extent that changes the roughness of an octave by a factor of nineteen.

That asymmetry is the answer to why the practice exists in the form it does. Every other quantity a vibrato moves, it moves by a lot. The pitch it barely touches.

Everything this field argues about, and the width of one sung note. Eight quantities in cents, on one linear axis. Seven of them are distinctions the tuning essays are about, from a beat once in two seconds, on a fifth at 1.31 cents to the Pythagorean comma at 23.5. The eighth is the peak-to-peak width of a single note sung with an ordinary vibrato of ±71 cents, which is 142 cents — 6.1 times the Pythagorean comma and 35 times the difference limen at 440 Hz. Every bar above the axis is smaller than the note it would be measured on.
Fig. 6 The size of a vibrato against everything else this site measures in cents. It is larger than every comma and every tempering decision, and the pitch mechanism does not notice.
The roughness of major third, over one second of vibrato. Every roughness figure until now is one number for a steady spectrum. This is the same computation evaluated at every instant of a 50-cent vibrato at 6 hertz on both voices. It has three quantities where there was one: a mean of 1.48e-1, a peak-to-peak depth of 57 per cent of that mean, and a fluctuation rate of 12 hertz. The still value — the roughness at the two nominal frequencies — is 1.45e-1, so the moving mean is 1.02 times it. And 12 hertz is slow enough to be counted, so the roughness itself is a thing with a beat rate — which is the recursion this figure ran into.
Fig. 7 The roughness of one interval evaluated at every instant of a fifty-cent vibrato at six hertz, on both voices at once. Every roughness figure until now is a single number for a steady spectrum; this one has three quantities where there was one.

A mean, a depth and a rate — and only the first of them is what the steady calculation reports. The pitch does not wobble and the roughness does, which is the asymmetry this essay is named for: the extractor smooths one quantity into stability and leaves the other fluctuating at the vibrato rate.

Which computation produced the numbers

The complex is twelve partials with the amplitudes this collection calls a string spectrum, on a fundamental of 220 hertz unless a figure says otherwise. Twelve rather than six because the peak’s width stops changing above about twelve for this envelope, which falls as one over the partial number squared.

The autocorrelation of a sum of steady partials is evaluated in closed form as the amplitude-weighted sum of cosines, which is the function this collection has used for organ pipes and loudspeakers since the missing-fundamental ladder. There is no sampling and no window in it.

The moving case is that function evaluated at forty-eight equally spaced phases of one vibrato cycle and averaged. That is the assumption doing the most work here and it is stated rather than derived: it says a listener’s pitch mechanism integrates over at least one vibrato period, which at six hertz is 167 milliseconds, and the integration windows this collection uses elsewhere are of that order.

The peak’s position is found by parabolic interpolation on the lag grid, because the grid is 0.01 milliseconds and the quantity being reported is a shift of 0.005.

The salience is the peak’s height above the mean of the lags more than a quarter of a period away from it, which is the contrast measure the jitter rung argued for: a bare peak height settles at the density of the signal rather than at zero, and what a mechanism could use is the difference.

The vibrato extents are the fourth rung’s, from ten tenors and 1,018 tones.

Where the model stops

One vibrato cycle is asserted as the window. Everything here averages over exactly one, and a listener’s window is neither exactly that nor rectangular. A shorter window tracks the wobble and a longer one does what this does; the interesting region is between them and the model has no shape for it.

Autocorrelation is one of several accounts. A harmonic-template model would give a different answer to the same question — templates match in frequency and would follow the sweep with no smearing at all, so the salience cost here would be zero and the sharp bias would not exist. The two accounts differ in a way this rung makes visible, which is more useful than either being right.

The spectrum is fixed. A real singer’s partials do not merely slide: the vocal tract’s formants are fixed, so a sweeping fundamental moves its partials through stationary resonances and their amplitudes modulate as well as their frequencies. The formant account is on this ladder and joining it to this is a real calculation nobody here has done.

And there is no noise. A real signal reaches the extractor through auditory filters with their own ringing and their own jitter, and the salience computed here is an upper bound in the way every noiseless calculation is.

What the picture cannot show

It cannot show a listener. Everything above is a property of a computation, and whether the human mechanism is this computation is the question the missing-fundamental ladder spends nine rungs on.

Nor can it show two voices. A choir is many independent vibratos, and the autocorrelation of their sum is not the sum of their autocorrelations. What a choir buys is the rung about the sum and it is about loudness rather than pitch.

It cannot show the onset. A singer’s vibrato takes a few hundred milliseconds to establish, so the beginning of a note is a straight tone and the pitch is available before the vibrato starts. That may be the whole answer to how a listener gets a definite pitch out of a wobbling note, and it is a fact about time that no averaged autocorrelation contains.

And it cannot show intonation. Whether a singer with a vibrato is in tune is a question about the relation between two moving pitches, and this rung has computed one.

Whose singing, and when

The vibrato is the Western operatic one of the twentieth century: about six hertz, thirty to a hundred and twenty cents either side, measured on recordings of male soloists.

Everything on this page collapses on a straight tone, which is what Renaissance and early baroque treatises describe and what a great deal of recorded early music does. That is not a limitation of the model; it is that the model’s whole subject is absent.

The one historical observation available is the one about the boundary. The vibrato that would defeat a period mechanism is around two hundred cents, and singing described as wobbling is described that way at extents somewhere near there — but nobody has measured a wobble, because it is a criticism rather than a technique, and the studies that measure vibrato measure singers who are thought to be singing well.

Where this ladder goes next

Eleven rungs. The larynx as a reed; two mechanisms and a seam; one voice over ninety players; a note that is never at its pitch; the sound a listener knows best; what a choir buys; two sections beating between partials; those partials moving; the roughness that motion produces; which of its three statistics survives a window; and now the note itself, which survives all of it.

What is owed after this is the formant. Every rung above sweeps a set of partials and holds their amplitudes still, and a real singer’s tract does not move with the fundamental — so a partial sweeping past a formant is a partial whose level modulates at six hertz as well as its frequency, and the two modulations are ninety degrees out of phase with each other on one side of the resonance and in phase on the other. This ladder has the formant account from its first rung and the sweeping account from its eighth, and the thing they make together is an amplitude modulation nobody put there, at the vibrato rate, whose depth is a function of where in the vowel the partial sits. That is the term that would say why a vibrato is louder as well as wider, and it is the last thing on this ladder that can be computed without a listener.

Part 11 of 13

One essay in the series on the voice. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AutocorrelationCritical bandwidthIntegration windowPitch perceptionResidue pitchThe voiceVibrato