The pitch that moves the wrong distance
Assumes: The note that is not there · A pitch with nothing to match
Eight essays ago this collection took the oldest result in psychoacoustics — that a complex tone missing its fundamental still has the pitch of that fundamental — and started asking where the pitch comes from.
The second rung removed one answer: the ear does generate difference tones, they are measurable, and they are not what produces the residue, because the residue survives at levels and in noise conditions where the distortion products cannot be there.
That left two accounts standing and they have been standing since. Either the ear is finding how often the waveform repeats, or it is finding which harmonic series best fits the partials it has resolved. Both predict the right answer for an ordinary complex tone, which is why the question survived so long.
There is a stimulus on which they predict different numbers, it was built in the nineteen-fifties, and it is the last thing this ladder has to do.
The stimulus
Take three partials an equal distance apart — 1800, 2000 and 2200 hertz — which are the ninth, tenth and eleventh harmonics of 200. The pitch is 200 and there is nothing at 200 in the sound.
Now move all three up by the same number of hertz. Not by the same ratio — by the same number, so the spacing between them is untouched.
That is the whole experiment. The spacing is a property of the differences between the partials and the shift does not change any difference. Anything that computes the pitch from how often something repeats is looking at a number the manipulation cannot move.
What happens
The pitch moves, and it moves by the shift divided by the harmonic number.
At a shift of forty hertz the best template is 204.0, which is 200 plus 40/10 exactly. At eighty it is 208. Reported pitch matches: this is the first effect of pitch shift, it was measured by de Boer and by Schouten, Ritsma and Cardozo, and the linear relation with slope one over the mean harmonic number is what they found.
The envelope rate is 200 throughout. A mechanism that returns the repetition rate of the envelope returns the wrong answer at every shift except zero, and by up to twenty hertz — which at these frequencies is a hundred and sixty-five cents, well over a semitone and not a subtle discrepancy.
And there is more than one answer
The second thing the experiment produces is not a discrepancy but a multiplicity, and it is at least as awkward for a rate account.
Listeners in these experiments report more than one pitch, they can be led to hear one or another by context, and the set of pitches they report is the set of good template fits. A repetition rate is a single number and cannot be ambiguous. A template fit is a search over readings and is ambiguous exactly when two readings fit comparably well, which is what the figure shows.
The sweep above has a sawtooth in it for the same reason: as the shift grows, the reading one harmonic higher becomes the better fit, the solid line jumps down, and at a full spacing the stimulus is once again an exact harmonic series — of the tenth, eleventh and twelfth harmonics of 200 — so the pitch is back where it started. The function is periodic in the shift with period equal to the spacing, and that periodicity is a signature of a template search and of nothing else.
The same search drawn as this collection’s residue machinery normally draws it gives the goodness of each reading, and it is the same function the seventh rung used on an inharmonic set — pointed here at a set that is inharmonic in a controlled way rather than because a bar or a bell happens to be shaped as it is.
And a check on the other side of the same mechanism: mistune one partial of an otherwise harmonic complex and it stops belonging — at a few per cent it is heard as a separate tone and stops contributing to the pitch. That is a template search reporting a poor fit for one element, and it is the same operation the shift experiment runs on all the elements at once.
What the experiment does not settle
It would be neat to stop there, and it would be wrong, because the two surviving accounts are not “envelope rate” and “template”. They are three, and the third survives.
A waveform’s envelope repeats at the spacing. Its fine structure does not: shifting the partials changes the phase relationship inside each envelope period, so the waveform as a whole is not periodic at 1/200 second at all. A mechanism operating on the fine structure — an autocorrelation of the neural firing pattern, which is the standard temporal account — sees that.
Run this collection’s own autocorrelation on the shifted complex and it does not stay at 200. At a shift of forty its highest peak in the relevant range is at 4.903 milliseconds, which is 203.97 hertz — the same number the template gives, to four significant figures.
That is not a coincidence and it is not a bug. A harmonic template fit and an autocorrelation peak are, for a resolved complex, two ways of asking the same question; they agree wherever the partials are far enough apart to be resolved, and the shift experiment does not separate them. Where they do differ is which branch of the ambiguity wins: at a shift of a hundred hertz the template’s best fit is 191.0 and the autocorrelation’s highest peak is 209.9 — the same two candidates, ranked the other way.
So the honest statement of what nine rungs have established is:
The residue is not a distortion product. That is rung two, and it is settled.
The residue is not the envelope repetition rate. That is this rung, and it is settled by the shift.
Whether it is a spectral template or a temporal autocorrelation is not settled, and the shift experiment cannot settle it, because on a resolved complex the two compute the same thing.
Where the two do come apart, and why it does not help
“Cannot settle it” is a strong claim and the sentence four paragraphs above already contains a counterexample to it, so it is worth running the comparison properly rather than at one shift.
Sweeping the shift from nought to a full spacing in one-hertz steps and taking both predictions at every step: the template fit and the highest autocorrelation peak agree to within a cent almost everywhere. They differ by more than fifty cents at six of the two hundred and one shifts — a window three hertz wide either side of ninety-eight — and inside it they differ by up to 164 cents, which is well over a semitone and not a subtlety.
The window is the branch flip, and the two mechanisms flip at different places. The template abandons the reading centred on the tenth harmonic at a shift of 95 hertz; the autocorrelation holds it until 101. Between those two shifts one mechanism has moved to the lower branch and the other has not, and they name pitches a minor second and a half apart. Everywhere else in the sweep they are on the same branch and therefore on the same number.
So the experiment is not blind to the distinction. It discriminates in three per cent of its own parameter range, and the three per cent is the worst possible place to try to measure anything.
That is the point, and it is worth stating carefully because it is the reason the last three decades of this argument happened somewhere else. At the branch flip the two candidate readings fit equally well by construction — that is what a flip is — so the pitch is at its most ambiguous, its weakest and its most susceptible to whatever context the listener was given. It is the region where listeners report two pitches rather than one, where the report depends on what preceded the stimulus, and where a mean over trials is a mean over a bimodal distribution and describes neither mode. A discriminating measurement needs a stimulus on which the two mechanisms differ and on which the percept is stable, and the shift experiment offers the first only where it destroys the second.
A window that narrow is also a window that any correction closes. The six hertz is a property of this model — three partials, no filtering, an autocorrelation of the exact sum of sinusoids. A real temporal mechanism autocorrelates the output of each cochlear filter and sums across channels, which smooths the peak structure and moves the flip; a real spectral one has a tolerance and a prior over harmonic numbers, which moves the other. Neither correction is small compared to six hertz. So the honest reading is not that the shift experiment discriminates a little, but that whether it discriminates at all is decided by modelling choices downstream of the experiment itself — which is a stronger reason to go elsewhere than “the two compute the same thing”.
Above about the eighth partial the ear cannot separate neighbours, so a spectral mechanism has nothing to fit and a temporal one still has a waveform — which is why the argument moved decades ago to unresolved complexes, where pitch is weaker and still present. That is the strongest evidence for a temporal component, and it sits at the same boundary the dominance region does, measured for a different purpose.
It is worth saying why that is not a disappointing outcome. The two surviving accounts are not rival theories of different things; they are two descriptions of one computation, and a great deal of the literature since the nineteen-seventies has been about whether the distinction is even well posed for a resolved complex. What the shift experiment eliminates is the mechanism that was actually simple — count the repeats — and the reason it was worth eliminating is that it is the one everybody’s intuition supplies.
The second effect, which this model does not reproduce
There is a further detail in the measured data and it is worth recording as a failure rather than omitting.
The observed pitch shift is slightly larger than the shift divided by the mean harmonic number — by something like ten per cent, in the direction of the lower partials mattering more. The usual explanation is that the ear weights its readings toward the low, well-resolved partials rather than treating all three equally.
That explanation is testable in one line, and here it does not work. A least-squares template fit that weights each partial by one over its harmonic number squared — a very strong bias toward the low end — gives a slope of 1.0135 times the unweighted prediction. The measured excess is about 1.1. The mechanism produces one and a bit per cent where ten is wanted, so unequal weighting of a three-partial complex is not the explanation, or is not the whole of it.
There is a sharper way to put that, and it is exact rather than a single evaluation. A partial at harmonic n taken alone implies a fundamental of (n·spacing + shift)/n, which is the spacing plus shift/n; against the nominal prediction of shift/10 that is a slope of exactly 10/n. So the ninth partial alone gives 10/9, the tenth gives 1, and the eleventh gives 10/11. Any weighted least-squares fit over the three is a weighted compromise between those, so every possible weighting of this stimulus produces a slope between 0.909 and 1.1111, and the ceiling is 10/9.
The measured excess of about 1.1 sits at ninety-nine per cent of that ceiling. That reframes the objection rather than removing it: unequal weighting is not arithmetically incapable of producing the measured shift, it is capable of producing it only by discarding the other two partials almost entirely. Solving for the exponent that reaches 1.10 gives about 25 — which weights the ninth partial some two hundred times the eleventh, and is a description of a mechanism that listens to one partial rather than of a mechanism that prefers low ones.
So the explanation fails by saturating rather than by falling short. It has one free parameter, the parameter has a hard ceiling barely above the number it has to reach, and reaching it costs the assumption that made the account plausible — that the ear is combining its resolved partials with a bias, rather than picking one. Either the excess comes from somewhere else, or the three-partial stimulus is too impoverished to distinguish “weights the low partials” from “uses the lowest partial”, which is itself a reason to want more partials.
The ordinary case for comparison is a harmonic complex with its fundamental deleted: the spacing and the fit agree, every reading points at the same fundamental, and there is nothing to be ambiguous about. And the three kinds of partial set this ladder has dealt with — harmonic, membrane, bell — differ in how much a template has to forgive, which is the axis the shifted complex moves along under control.
Closing the ladder
Nine rungs, and the model is a set of partials with a pitch assigned to it. What bounds a model is having said something about every variable it has, and this one has five.
Is the fundamental present? Rung one: it does not have to be.
What produces the pitch? Rung two removed the ear’s own distortion; this rung removes the envelope rate; what is left is a template fit or its temporal equivalent, and the two are not separated.
Are the partials harmonic? Rungs three and four: a bell and a drum are the two ways of not being, one converging on a harmonic subset and one not converging at all.
Which partials? Rung five: the dominance region, and it is low.
Where do they come from? Rung six: two pipes can supply them, because the mechanism does not care which source a partial arrived from, and rung eight is the same fact exploited in a loudspeaker.
And what if nothing fits? Rung seven: the fit degrades gracefully and the pitch becomes weak and multiple rather than absent.
Every variable the model has now has a rung, and this one is the last because it is the only one that discriminates between mechanisms rather than describing behaviour. The anchor closes at nine.
What is not on the list belongs elsewhere. How a residue pitch behaves in a chord is consonance; how it interacts with a second source is auditory scene analysis; how finely it can be compared is pitch acuity. Those are different models rather than further rungs.
Which computation produced the numbers
The shifted complex is built as (n₀ − 1, n₀, n₀ + 1) times the spacing, plus a constant shift added to each. Nothing else is done to it: no envelope, no level differences, no phase manipulation.
The template fit is the same least-squares search this collection has used for the residue since the fifth rung. For every assignment of harmonic numbers up to twenty, the fundamental minimising the squared error is computed in closed form, the worst deviation in cents is checked against a tolerance, and the survivors are ranked by fit and then by how few gaps the reading leaves. Sixty cents is the tolerance in the figures here, against thirty in the earlier essays, because a shifted complex is deliberately inharmonic and a tight tolerance rejects every reading.
The autocorrelation is computed from the partial frequencies directly — the sum over partials of cos(2πfτ), weighted by amplitude — which is the exact autocorrelation of a sum of sinusoids and needs no sampled waveform. The lag search runs from 3.5 to 7 milliseconds, a window either side of the nominal period.
The weighted fit is the same closed form with each partial’s residual scaled by n to a power; the reported slope is the fitted pitch at a shift of a hundred hertz minus its value at zero, over the prediction. The 10/9 ceiling is not a result of that sweep but of the algebra above it, so it holds for any weighting scheme whatever and not only for this family of exponents.
The comparison between the two mechanisms is the same two computations run at every shift from nought to two hundred hertz in steps of one, with the autocorrelation peak taken as the largest local maximum strictly inside the lag window rather than the largest value in it. That distinction is load-bearing and is the kind of error this collection has made before: the autocorrelation of three high partials rises toward short lags, so a search for the largest value returns the edge of the window at every shift and reports a constant, plausible-looking pitch that is an artefact of where the window was cut.
What the picture cannot show
There is no listener anywhere in this essay. The reported pitches, the ambiguity and the ten-per-cent excess are all quoted from the experimental literature. What is computed here is what each mechanism predicts, which is the half that can be computed.
Three partials is the classic stimulus and it is a hard case for both accounts. With more partials the template’s ambiguity shrinks fast and the autocorrelation’s peaks sharpen; the experiment is designed to be marginal, which is what makes it discriminating and also what makes both mechanisms unstable on it.
The autocorrelation here has no auditory filtering in it. A real temporal model runs the stimulus through a bank of cochlear filters and autocorrelates each channel, and the summary across channels is what produces the pitch. That would change the ranking of the branches, which is exactly the quantity the last section says is unresolved.
And nothing here is at a stated level. Both the dominance region and the resolvability limit move with level, and every number in this essay assumes the partials are resolved, which for the ninth to eleventh harmonics of 200 hertz is marginal at best.
The ladder from here
There is no further rung. What the closure leaves open is the one thing the nine rungs could not decide, and it is not a gap in the ladder but a gap in the field: the shift experiment separates the envelope from everything else and leaves a spectral fit and a temporal one indistinguishable, and every experiment since has been an attempt to build a stimulus on which they differ. That work belongs to whoever writes about unresolved harmonics, and this collection’s contribution is to have shown exactly why the obvious experiment does not do it.
Part 9 of 9
One essay in the series on missing fundamental. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
AutocorrelationHarmonicityOctave ambiguityPeriodicityResidue pitchResolvabilityTemporal codingVirtual-pitch
- Eleven partials is one too many harmonicity, resolvability
- The beat that is not in the air periodicity, temporal coding
- The root an ear supplies residue pitch, virtual-pitch