The pitch that does not wobble
Assumes: The mean survives the window · A note that is never at its pitch
The mean survives the window ended the ninth and tenth rungs’ business with roughness and said what all three of them had left out:
Everything in the last three rungs treats a vibrato as a modulation of a sensory quantity, and the reason singers use one is that it does something to the note — a listener hears a steady pitch, more presence, and a voice that carries through a texture.
The first of those three is the one this collection can compute. The pitch that moves the wrong distance and the rungs around it built an account of how a pitch is extracted from a set of partials — a period, found by asking at what lag the waveform most resembles itself — and every one of them handed it a signal that was standing still.
Nothing is mistuned
The first thing to be clear about is what a vibrato is not.
A vibrato modulates the fundamental, so every partial moves by the same number of cents. At every instant the complex is exactly harmonic: the ratios between the partials never change, no partial is ever mistuned against any other, and every cue this collection has for whether a set of components belongs to one sound says yes at every moment.
So the difficulty is not identification. It is that the answer will not stay put. A pitch mechanism that looks for a lag at which the waveform repeats has a target that is moving, at six hertz, over a hundred and forty cents peak to peak — which is a swing wider than any interval this collection calls a comma.
Two widths, and the whole argument is their ratio
The autocorrelation of a steady harmonic complex has a peak at the period, and that peak has a width. It is set by the highest partial carrying appreciable energy: a complex with many strong partials has a sharp waveform and a sharp peak, and one with three has a broad one. For a string-like spectrum on A3 the peak is 0.90 milliseconds wide at half height, against a period of 4.55.
A vibrato moves the period. A swing of ±71 cents is a fractional frequency swing of ±4.2 per cent, so the period swings by 0.37 milliseconds.
Those two numbers are the whole of it. 0.37 against 0.90: the vibrato smears the peak by less than half its own width, so averaging over a cycle leaves a peak that is slightly lower and slightly broader and unmistakably in the same place.
A vibrato is a large interval and a small smear, and the reason it is both is that a pitch mechanism is not an interval mechanism.
What it costs, and where the practice stops
Sweeping the extent turns that into a curve.
The peak keeps 98.5 per cent of its salience at 34 cents, 93.9 at 71, and 84.5 at 123. Those three numbers are the bottom, the middle and the top of the range Prame measured across a thousand tones, and the fourth rung of this ladder is where they came into the collection.
Above that band it falls away fast: 70.5 per cent at 200 cents, 46.3 at 340. And the period swing passes the peak’s own width at about 200 — a little above the widest vibrato anybody in that study was recorded singing.
So the measured practice occupies exactly the region in which a period mechanism is degraded and not defeated, and stops close to where it would be. That is the kind of coincidence worth stating carefully. It is not evidence that singers are optimising a pitch extractor; nobody is doing that. It is that a vibrato which destroyed the pitch would not be a vibrato, it would be a bad note, and the boundary between the two is computable and turns out to be where the practice’s own edge is.
The vocabulary agrees. A vibrato past that boundary has a name — a wobble — and what everybody says about it is that the note stops having a definite pitch. The model says the peak has stopped being narrower than the swing.
And the register drops out entirely
The obvious next question is whether it is harder low or high, and the answer is neither, exactly.
The period goes as one over the frequency. The peak’s width goes as one over the frequency, because it is set by the highest partial and the partials scale with the fundamental. The vibrato’s swing on the period goes as one over the frequency, because the extent is in cents.
Three quantities all proportional to the same thing, and the answer is a ratio of two of them. There is no pitch in it. A seventy-one-cent vibrato costs the same six per cent of the pitch’s salience on a bass’s low A as on a soprano’s high one.
That is worth setting against what the tenth rung found for the same signal. Roughness has a scale in it — the critical band is a width in hertz, not in cents — so the cost of a vibrato in roughness changes by a large factor up the compass, and that rung drew the change.
Two accounts of one physical signal, one of which has a register and one of which cannot. A singer moving up an octave is changing what their vibrato does to the sensory dissonance of the texture and changing nothing at all about what it does to their own pitch.
A prediction with a sign in it
The averaged peak is not only lower and wider. Its maximum has moved, and the direction is not the obvious one.
The natural guess is flat. A vibrato is symmetric in cents, so it spends equal time above and below; a period is one over a frequency, and the average of one over a symmetric swing is longer than one over the middle. That is the same Jensen argument the ninth rung used on roughness, and it points to a pitch slightly below the mean.
The computation gives sharp: 2.4 cents at typical extent, 9.1 at the top of the measured range, 36 at 200 cents.
The mechanism is that the smear is proportional to the lag. Each instant’s autocorrelation contributes a peak at its own period, and the contributions at longer lags are individually broader — so averaging erodes the long-lag side of the peak more than the short-lag side, and the maximum slides toward the short end. The asymmetry is in the instrument, not in the signal.
Two and a half cents is a small quantity and it is not an unmeasurable one: it is inside what a listener can do comparing two tones directly, and the effect is predicted to grow with the extent, which makes it a slope rather than an offset. The perceived pitch of a vibrato tone should be a few cents sharp of its mean, and sharper the wider the vibrato.
That is the one claim on this page that could be shown to be wrong by a listener rather than by a better model.
The rate runs through zero and out the other side, which is the sharpest form of what a vibrato does to a coincidence: the two partials are not slightly mistuned, they are alternately sharp and flat of each other. Anything that averages a beat rate over a vibrato cycle is averaging a signed quantity, and the mean is not what a listener hears.
The extractor’s own choice matters, and it is visible here
There is a second reading of the sharp bias and it belongs in the open rather than in the caveats.
Two mechanisms are on offer in the literature for how a set of partials becomes a pitch, and this collection has drawn both. One is periodicity: find the lag at which the waveform repeats, which is what is computed above. The other is a harmonic template: slide a comb of whole-number multiples along the frequency axis and find where it fits.
On a steady complex the two agree, which is why the argument between them has needed carefully constructed stimuli — a shifted complex, whose partials are equally spaced but not multiples of anything, on which the two accounts give different answers and listeners side with periodicity.
A vibrato is a stimulus of the same kind and nobody has used it that way. A template does not care that the whole comb is sliding: it fits perfectly at every instant, and averaging its output over a cycle gives the mean with no loss of salience and no bias. Periodicity loses six per cent of its salience and gains two and a half cents of sharpness.
So the two accounts differ by a measurable amount on the commonest signal in Western singing, and the difference is in the direction of a quantity somebody could put a slider on.
Why a singer would want this
None of the above says a vibrato is good for anything. It says the pitch survives one, which is a licence rather than a reason.
The reason the ladder can now assemble is a negative one and it is worth having. A voice with a vibrato has, at every instant, a spectrum spread over a hundred and forty cents; two voices with independent vibratos do not beat in any steady way; a section of them fills a band rather than a line. All of that is what makes a choral sound and it is bought with the frequency axis.
What it is not bought with is pitch definiteness. The one thing a singer cannot afford to lose is the note, and this figure says the note is the cheapest thing in the transaction — six per cent of a salience, at an extent that changes the roughness of an octave by a factor of nineteen.
That asymmetry is the answer to why the practice exists in the form it does. Every other quantity a vibrato moves, it moves by a lot. The pitch it barely touches.
A mean, a depth and a rate — and only the first of them is what the steady calculation reports. The pitch does not wobble and the roughness does, which is the asymmetry this essay is named for: the extractor smooths one quantity into stability and leaves the other fluctuating at the vibrato rate.
Which computation produced the numbers
The complex is twelve partials with the amplitudes this collection calls a string spectrum, on a fundamental of 220 hertz unless a figure says otherwise. Twelve rather than six because the peak’s width stops changing above about twelve for this envelope, which falls as one over the partial number squared.
The autocorrelation of a sum of steady partials is evaluated in closed form as the amplitude-weighted sum of cosines, which is the function this collection has used for organ pipes and loudspeakers since the missing-fundamental ladder. There is no sampling and no window in it.
The moving case is that function evaluated at forty-eight equally spaced phases of one vibrato cycle and averaged. That is the assumption doing the most work here and it is stated rather than derived: it says a listener’s pitch mechanism integrates over at least one vibrato period, which at six hertz is 167 milliseconds, and the integration windows this collection uses elsewhere are of that order.
The peak’s position is found by parabolic interpolation on the lag grid, because the grid is 0.01 milliseconds and the quantity being reported is a shift of 0.005.
The salience is the peak’s height above the mean of the lags more than a quarter of a period away from it, which is the contrast measure the jitter rung argued for: a bare peak height settles at the density of the signal rather than at zero, and what a mechanism could use is the difference.
The vibrato extents are the fourth rung’s, from ten tenors and 1,018 tones.
Where the model stops
One vibrato cycle is asserted as the window. Everything here averages over exactly one, and a listener’s window is neither exactly that nor rectangular. A shorter window tracks the wobble and a longer one does what this does; the interesting region is between them and the model has no shape for it.
Autocorrelation is one of several accounts. A harmonic-template model would give a different answer to the same question — templates match in frequency and would follow the sweep with no smearing at all, so the salience cost here would be zero and the sharp bias would not exist. The two accounts differ in a way this rung makes visible, which is more useful than either being right.
The spectrum is fixed. A real singer’s partials do not merely slide: the vocal tract’s formants are fixed, so a sweeping fundamental moves its partials through stationary resonances and their amplitudes modulate as well as their frequencies. The formant account is on this ladder and joining it to this is a real calculation nobody here has done.
And there is no noise. A real signal reaches the extractor through auditory filters with their own ringing and their own jitter, and the salience computed here is an upper bound in the way every noiseless calculation is.
What the picture cannot show
It cannot show a listener. Everything above is a property of a computation, and whether the human mechanism is this computation is the question the missing-fundamental ladder spends nine rungs on.
Nor can it show two voices. A choir is many independent vibratos, and the autocorrelation of their sum is not the sum of their autocorrelations. What a choir buys is the rung about the sum and it is about loudness rather than pitch.
It cannot show the onset. A singer’s vibrato takes a few hundred milliseconds to establish, so the beginning of a note is a straight tone and the pitch is available before the vibrato starts. That may be the whole answer to how a listener gets a definite pitch out of a wobbling note, and it is a fact about time that no averaged autocorrelation contains.
And it cannot show intonation. Whether a singer with a vibrato is in tune is a question about the relation between two moving pitches, and this rung has computed one.
Whose singing, and when
The vibrato is the Western operatic one of the twentieth century: about six hertz, thirty to a hundred and twenty cents either side, measured on recordings of male soloists.
Everything on this page collapses on a straight tone, which is what Renaissance and early baroque treatises describe and what a great deal of recorded early music does. That is not a limitation of the model; it is that the model’s whole subject is absent.
The one historical observation available is the one about the boundary. The vibrato that would defeat a period mechanism is around two hundred cents, and singing described as wobbling is described that way at extents somewhere near there — but nobody has measured a wobble, because it is a criticism rather than a technique, and the studies that measure vibrato measure singers who are thought to be singing well.
Where this ladder goes next
Eleven rungs. The larynx as a reed; two mechanisms and a seam; one voice over ninety players; a note that is never at its pitch; the sound a listener knows best; what a choir buys; two sections beating between partials; those partials moving; the roughness that motion produces; which of its three statistics survives a window; and now the note itself, which survives all of it.
What is owed after this is the formant. Every rung above sweeps a set of partials and holds their amplitudes still, and a real singer’s tract does not move with the fundamental — so a partial sweeping past a formant is a partial whose level modulates at six hertz as well as its frequency, and the two modulations are ninety degrees out of phase with each other on one side of the resonance and in phase on the other. This ladder has the formant account from its first rung and the sweeping account from its eighth, and the thing they make together is an amplitude modulation nobody put there, at the vibrato rate, whose depth is a function of where in the vowel the partial sits. That is the term that would say why a vibrato is louder as well as wider, and it is the last thing on this ladder that can be computed without a listener.
Part 11 of 13
One essay in the series on the voice. The essays either side of this one:
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
AutocorrelationCritical bandwidthIntegration windowPitch perceptionResidue pitchThe voiceVibrato
- A dissonance has to last critical bandwidth, integration window
- A part entering is not a change of level critical bandwidth, integration window
- A roughness with a rate of its own critical bandwidth, vibrato
- Sixteen sweeps against sixteen critical bandwidth, vibrato
- The dissonance arrives and the dynamic does not critical bandwidth, integration window
- The rate that does not rise with the partial critical bandwidth, vibrato