A pitch with nothing to match
Assumes: The note that is not there · Which harmonics carry the pitch
Two hundred clicks a second, band-passed so that nothing survives below three kilohertz, has a pitch. It is a poor, thin, buzzy pitch, and it is at two hundred hertz, and a listener can match it to a tone and can hear it move when the rate moves.
Nothing in that sound is at two hundred hertz. The partials that are present run from the fifteenth harmonic upward, and — this is the part that matters — not one of them is separable from its neighbours by the ear’s own frequency analysis. There is no pattern to match, because there is no pattern the ear can read.
The argument, and what it is an argument against
The first rung of this anchor chose between two accounts of the missing fundamental. Helmholtz’s said the ear’s own nonlinearity generates a real tone at the fundamental; Schouten’s said the auditory system fits a harmonic series to the partials present and reports its fundamental. The shifted-residue experiment settled it against Helmholtz, and this site has been running on the pattern account ever since.
A pattern account needs partials it can see. The previous two rungs counted how many of them a note actually delivers and found that the answer collapses in the bass, and that the case the anchor opened with — a bass line through a small speaker — delivers none at all.
This rung supplies the stimulus that isolates the alternative. Not a variation on a musical sound, which always has resolved partials somewhere, but a sound engineered so that a template match is impossible and a pitch is nevertheless heard.
Harmonics fifteen to twenty-five of a 200 Hz train fall inside the three-to-five-kilohertz band, where the ear’s analysis bandwidth is between 350 and 560 hertz against a spacing of 200. Every filter up there therefore contains several harmonics at once, which is the definition of an unresolved region and is the condition this whole essay is about.
The mechanism, stated as a computation
If the harmonics are not separable, what is left in the filter’s output is a series of events in time: the filter rings once per click and the rings do not overlap. The quantity a mechanism could read is then the repeat interval, and the standard way of extracting a repeat interval from a signal is to correlate it with a delayed copy of itself.
That is the autocorrelation account of pitch, and it is old — Licklider proposed a neural version of it in 1951, and it has been reinvented in a dozen forms since. Its prediction for a click train is exact: the correlation has a maximum at every multiple of the repetition period, and the pitch reported is the first of them.
Everything in this essay’s figures is that correlation, computed rather than sketched, with one modelling choice stated in every caption. The ear does not receive impulses. A band-pass filter of equivalent rectangular bandwidth rings for about , and two clicks closer together than that are one event in its output. At four kilohertz is 456 hertz, so the ring is 2.2 milliseconds, and each click is smoothed by half that before the correlation is taken.
Which computation produced the number
A bare peak height is the wrong measure of what a pitch mechanism could use, and using it understates the collapse badly. A heavily jittered train still has a highest point somewhere in its correlation, and that point’s height settles at the average density of the train rather than at zero. What matters is the contrast between the peak and the lags around it, and that is what the figures plot: the peak minus the mean of the correlation at lags more than a quarter of a period away from it, averaged over eight seeded trains.
The threshold is stated as a fraction of the period rather than as a number of milliseconds, and that is a claim which one rate cannot support.
The model’s own answer is that the cue halves at a jitter of roughly a seventh of the period. The listener’s answer is a published measurement and it is quoted as a range rather than a number, because it depends on the rate, the level, the band and the task, and because reports differ between listeners by more than they differ between experiments. What the literature broadly says is that the pitch of a jittered train weakens progressively and becomes unusable somewhere between a tenth and a quarter of the period.
It is worth being explicit about what the argument survives that being wrong by. The claim here is not that the threshold is fourteen per cent. It is that a jitter small enough to leave the average rate and the long-term spectrum alone destroys the pitch, and that no spectral account has anything to say about why. A published threshold anywhere from five per cent to fifty per cent leaves that intact. A finding that the pitch survives jitter of a whole period would refute it, and nobody reports that.
There is one measurement the model makes that a published threshold could contradict rather than merely fail to confirm, and it is the scaling. The collapse happens at a fixed fraction of the period at both rates tested, which is what a correlation predicts; a mechanism with a fixed timing resolution — a clock rather than a correlator — would collapse at a fixed number of microseconds, and that is a different curve on the same axes. Two published thresholds at two rates would separate them, and the difference between the predictions at a hundred and two hundred clicks a second is a factor of two, which is well outside the scatter of anything here.
The window the experiment lives in is narrower than it looks
The stimulus needs three things to be true at once, and only the first is usually stated.
The partials in the band have to be unresolved, which needs a bandwidth wider than the repetition rate. The clicks have to be separate events in the filter’s output, which needs a bandwidth wider than about twice the rate: a filter that rings for longer than the gap between clicks turns them into a continuous tone and there is nothing left to correlate. And the nerve has to be locking to the waveform at all, which fails above about five kilohertz.
Between 3.5 and 5 kilohertz, and nowhere else, a two-hundred-a-second train has a pitch that no spectral mechanism can be reading and a temporal mechanism can. That is where the classic high-harmonic experiments were run, and the arithmetic says it was not a free choice.
It also says why two hundred a second is close to the highest rate anybody used. Half an octave of usable band is not much room for a filter with skirts on it, and the next section finds where the room runs out entirely — which is close enough to two hundred that the choice of rate in the published work looks less like a convention than like the edge of what the stimulus allows.
And it closes altogether at 283 a second
Two boundaries rising and one fixed is a window that must shut, and solving for where it does is one line.
| rate | separate events above | window to 5 kHz |
|---|---|---|
| 50 | 698 Hz | 2.84 octaves |
| 100 | 1,624 | 1.62 |
| 200 | 3,477 | 0.52 |
| 250 | 4,403 | 0.18 |
| 283 | 5,000 | 0 |
| 300 | 5,330 | closed |
Above about 283 clicks a second there is no band at all in which the experiment can be run. Below that frequency the filter rings for longer than the gap between clicks and the correlation has nothing to find; above it the nerve is not locking. The two conditions cross at a repetition rate of 283 hertz, which is a semitone below D4.
That is a bound on the evidence rather than on the phenomenon, and the distinction matters. Nothing here says a temporal mechanism stops working above 283 hertz; the phase-locking limit is at five kilohertz and a residue pitch is reported up to about 1.4. What it says is that the stimulus that isolates the temporal account from the spectral one does not exist above 283 hertz, so every direct demonstration of timing-based pitch is a demonstration in the bottom two octaves of the musical range, and the account is extended upward by assumption.
Which puts a shape on how far this rung’s conclusion travels. The bass, where the previous rung found that a note’s resolved harmonics collapse, is exactly the register in which the temporal account can be shown to work — and it is also, by that rung’s arithmetic, the register in which the spectral account has least to work with. The two results are the same register seen from opposite sides, and the coincidence is not one: both are consequences of the critical band being wide compared with a low fundamental.
Above middle C the position reverses, and neither result applies. There the partials are resolved and the spectral account is comfortable, and the experiment that would test the alternative cannot be built.
The window also closes with rate. At a hundred clicks a second it runs from 1.6 kilohertz to 5 and is comfortable; at three hundred it has almost shut; above about two hundred and fifty a second there is no band in which both conditions hold, because the bandwidth needed to keep the clicks separate is above the phase-locking limit. The experiment is possible only for repetition rates in the bottom two octaves of the musical range — which is, not by accident, the register where the residue is used and where resolvability is worst.
Where the model stops, and it stops on this ladder’s own preferred account
Two things have to be said here, and the second is uncomfortable for the position this anchor has been holding.
The first is that autocorrelation is a model and not a mechanism, exactly as the least-squares template fit is. Nothing in the auditory system computes . Neural versions of the idea are built out of coincidence detectors and delay lines, the delay lines have never been found anatomically, and the whole family is argued about. The figures here state what a correlation predicts and not what happens.
The second is that the experiment which made the pattern account necessary does not separate the pattern account from this one.
The correlation gets the shifted residue right, and it gets the ambiguity right too: its local maxima in that window sit at 226, 204, 186 and 170 hertz, which are exactly the four fundamentals a template search over the same partials returns. Two models that were introduced as rivals turn out to agree on the case that was supposed to arbitrate them, and to agree on the errors as well.
Run the template search on the click train’s own partials — six of them, from the fifteenth harmonic up — and on paper it works: the series fits 200 hertz exactly and nothing else comes close. On paper is the whole difficulty, because the search is being handed partials the ear never separated.
So the jittered train is doing work the shifted residue cannot. It is the case where the two accounts come apart, because a template match over the long-term spectrum is completely indifferent to click timing and a correlation is not.
What the picture cannot show
The correlation drawn is of the pressure in the air. A real mechanism would correlate the output of each auditory filter separately and sum the results, which is a stronger and messier object — it weights each channel by how well that channel resolves anything, and it is what the published summary-autocorrelation models compute.
The octave problem is drawn and not solved. A correlation has maxima at every multiple of the period, so nothing in it prefers 200 hertz to 100. Every model in this family needs a further rule to pick the first peak, and the rules are ad hoc. That the same octave ambiguity shows up in the template account, and in an organ’s two pipes, is a point in favour of the ambiguity being real rather than of any one model.
Level is absent from all of it, and level moves both boundaries. Auditory filters broaden as the input grows, so a loud train is unresolved over a wider range and its filters ring for a shorter time — which widens the usable window at both ends. A figure drawn at one level is a figure about one level, and the same caution applies here as to every masking measurement on this site.
And the jitter is Gaussian and independent, which no real timing noise is. A drifting rate, a rate modulated by something else, and a rate with occasional missing clicks all leave different marks on a correlation, and all three occur in sounds people actually hear.
Above three kilohertz the partials stop being separable, and in this site’s usual units the reason is stark: adjacent harmonics of a 200 Hz train up at four kilohertz are eighty-four cents apart — less than a semitone — while the ear’s band up there is several times that wide. Nothing in the periphery is resolving them.
The same ceiling, met from a completely different direction
The five-kilohertz limit in these figures is not a fact about pitch. It is a fact about the auditory nerve: above roughly four to five kilohertz a nerve fibre’s firings stop being locked to the phase of the waveform and carry only its envelope and its average rate.
That property has a second consequence somewhere else entirely on this site. Locating a sound by the difference in arrival time at the two ears requires comparing the phase of two waveforms, and it fails at the top of the spectrum for the same reason — which is why localisation switches from a timing cue to a level cue partway up, and why the duplex theory has two halves rather than one.
So the ceiling on residue pitch and the crossover in binaural localisation are the same measurement seen twice, in two fields that do not otherwise touch. Neither is a property of pitch or of space; both are the point at which a nerve fibre stops being a clock.
Whose sound, and where it occurs
A band-passed click train is a laboratory object, and its musical relatives are real enough.
Vocal fry, the creaky register at the bottom of a speaking voice, is a train of glottal pulses at rates between about thirty and eighty a second, and its pulses are irregular. That irregularity is what makes fry sound rough rather than pitched, and it is the same manipulation as the figures above, applied by a larynx rather than by a random-number generator. A singer descending into fry crosses the boundary this essay computes.
A rattle, a snare buzz or a badly seated reed is the same object again at a rate too low or too irregular to have a pitch at all — as is an untuned drum, from the other end of the argument: a membrane’s partials support no common fundamental, so its waveform has no repeat for a correlation to find and no series for a template to fit, and both accounts return nothing together. And going the other way, a rhythm accelerated past about twenty events a second becomes a pitch, which is this essay’s stimulus approached from the bottom of the same continuum rather than from the top.
The engineering case is the one the anchor keeps returning to. A bass note through a small loudspeaker arrives with every resolved partial removed and a set of unresolved upper harmonics left, and what those harmonics have in common is that their sum repeats at the rate of the note. The next rung computes that case — and it is the one place where the argument of this essay is worth money.
Where the ladder goes next
This rung was written to supply evidence for the temporal account, and it does, and the honest summary of the phase is that it bounds the ladder’s own preferred model from the other side. The pattern account is not wrong. It is under-determined: it is one of at least two models that predict the shifted residue, the octave errors and the dominance of the low harmonics equally well, and the case that separates them is a stimulus nobody would call music.
What is left is a division of labour rather than a winner. Where partials are resolved, a template can be matched and the pitch is strong and definite. Where they are not — at the bottom of the bass, through a small speaker, in an organ’s resultant, in a click train above three kilohertz — a template has nothing to work on and the timing is all there is, and the pitch that results is exactly the weak, ambiguous, octave-prone thing the literature describes.
The last rung of this anchor takes that division into an engineering decision: what a three-inch driver leaves of a forty-hertz note, why the note is heard anyway, and what a manufacturer is really buying when a device adds harmonics of a bass it cannot reproduce.
Part 7 of 9
One essay in the series on missing fundamental. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
AutocorrelationCritical bandwidthJitterMissing fundamentalPeriodicityResidue pitchResolvabilityTemporal coding
- A bell has no fundamental periodicity, residue pitch
- Eleven partials is one too many critical bandwidth, resolvability
- Every member of a beat family is the same depth critical bandwidth, resolvability
- The beat that is not in the air periodicity, temporal coding
- The played notes already name the ghost bass missing fundamental, residue pitch
- The series has three tops critical bandwidth, resolvability