Perception and the listener

A pitch with nothing to match

Filter a click train into a band where no partial is separable from its neighbours and it still has a pitch at its repetition rate. Displace each click by a fraction of a millisecond, leaving the average rate and the long-term spectrum exactly where they were, and the pitch goes. The mechanism is reading the timing — which bounds the account endorsed here from the start.

Assumes: The note that is not there · Which harmonics carry the pitch

Two hundred clicks a second, band-passed so that nothing survives below three kilohertz, has a pitch. It is a poor, thin, buzzy pitch, and it is at two hundred hertz, and a listener can match it to a tone and can hear it move when the rate moves.

Nothing in that sound is at two hundred hertz. The partials that are present run from the fifteenth harmonic upward, and — this is the part that matters — not one of them is separable from its neighbours by the ear’s own frequency analysis. There is no pattern to match, because there is no pattern the ear can read.

A 200-a-second click train, correlated with itself. The autocorrelation of a click train at 200 a second, smoothed by the ring of an auditory filter centred at 4000 Hz — an equivalent rectangular bandwidth of 456 Hz, so a ring of 2.2 ms. The regular train peaks at 5.0 ms, one period. With each click displaced by a standard deviation of 20 per cent of the period — 1.00 ms — the peak's contrast against the surrounding lags falls from 0.41 to 0.12. The average rate and the long-term spectrum are unchanged by the jitter; only the timing is.
Fig. 1 The autocorrelation of that train against itself, at every lag out to two and a half periods. The regular train peaks hard at five milliseconds, one period, and the peak stands well clear of everything around it. The second curve is the identical train with each click displaced by a Gaussian of standard deviation a fifth of the period — one millisecond. The average rate is unchanged, the long-term spectrum is very nearly unchanged, and the peak has lost most of its contrast against its surroundings.

The argument, and what it is an argument against

The first rung of this anchor chose between two accounts of the missing fundamental. Helmholtz’s said the ear’s own nonlinearity generates a real tone at the fundamental; Schouten’s said the auditory system fits a harmonic series to the partials present and reports its fundamental. The shifted-residue experiment settled it against Helmholtz, and this site has been running on the pattern account ever since.

A pattern account needs partials it can see. The previous two rungs counted how many of them a note actually delivers and found that the answer collapses in the bass, and that the case the anchor opened with — a bass line through a small speaker — delivers none at all.

This rung supplies the stimulus that isolates the alternative. Not a variation on a musical sound, which always has resolved partials somewhere, but a sound engineered so that a template match is impossible and a pitch is nevertheless heard.

Harmonics fifteen to twenty-five of a 200 Hz train fall inside the three-to-five-kilohertz band, where the ear’s analysis bandwidth is between 350 and 560 hertz against a spacing of 200. Every filter up there therefore contains several harmonics at once, which is the definition of an unresolved region and is the condition this whole essay is about.

The mechanism, stated as a computation

If the harmonics are not separable, what is left in the filter’s output is a series of events in time: the filter rings once per click and the rings do not overlap. The quantity a mechanism could read is then the repeat interval, and the standard way of extracting a repeat interval from a signal is to correlate it with a delayed copy of itself.

That is the autocorrelation account of pitch, and it is old — Licklider proposed a neural version of it in 1951, and it has been reinvented in a dozen forms since. Its prediction for a click train is exact: the correlation has a maximum at every multiple of the repetition period, and the pitch reported is the first of them.

Everything in this essay’s figures is that correlation, computed rather than sketched, with one modelling choice stated in every caption. The ear does not receive impulses. A band-pass filter of equivalent rectangular bandwidth bb rings for about 1/b1/b, and two clicks closer together than that are one event in its output. At four kilohertz bb is 456 hertz, so the ring is 2.2 milliseconds, and each click is smoothed by half that before the correlation is taken.

A note with its first 8 partials removed. The spectrum of a 200 Hz tone with the lowest 8 partials deleted, and the wave that remains. The wave still repeats 200 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 2 The boundary drawn on an ordinary complex rather than on a click train. Eight harmonics of a 200 Hz note are resolved at this pitch, so deleting exactly those eight leaves a stimulus made entirely of partials the ear cannot separate — the ninth to the twelfth, all inside shared filters. The wave underneath still repeats 200 times a second. This is the click-train experiment with four partials instead of eleven, and it shows that the boundary this essay works either side of is a property of the note and not of the stimulus being exotic.
A 200-a-second click train, and the same train jittered. The onset times of a click train at 200 a second over 60 milliseconds. The faint rules are the exact grid. The lower row has each click displaced by a Gaussian of standard deviation 20 per cent of the period, which is 1.00 ms; the largest displacement drawn is 1.88 ms. The average rate is identical and the long-term spectrum is very nearly so. The pitch is not: this is the manipulation that removes it, and the picture is deliberately unconvincing about how small it is.
Fig. 3 The stimulus itself, and the reason this essay needs an autocorrelation rather than a picture. The faint rules are the exact grid and the coloured lines are the clicks. The jitter that removes the pitch is a fifth of a period — one millisecond at this rate — and at a scale where sixty milliseconds fits on a page it is very nearly invisible. The manipulation is a small one and the argument it settles is not.

Which computation produced the number

A bare peak height is the wrong measure of what a pitch mechanism could use, and using it understates the collapse badly. A heavily jittered train still has a highest point somewhere in its correlation, and that point’s height settles at the average density of the train rather than at zero. What matters is the contrast between the peak and the lags around it, and that is what the figures plot: the peak minus the mean of the correlation at lags more than a quarter of a period away from it, averaged over eight seeded trains.

What jitter does to the only cue a click train has. The contrast of the autocorrelation peak against the lags around it, against the jitter applied to the click times, at 200 clicks a second and averaged over 8 seeded trains. Through a filter centred at 4000 Hz — a ring of 2.2 ms — the contrast halves at a jitter of 14 per cent of the period, 0.72 ms. Through a filter centred at 8000 Hz — a ring of 1.1 ms — the contrast halves at a jitter of 14 per cent of the period, 0.70 ms. The two filter widths differ by a factor of two and their half-contrast points differ by less than a tenth of that, so the collapse is a property of the jitter and not of the modelling choice. The shaded strip is the range in which the literature reports listeners losing the pitch, which is a measurement and not a computation.
Fig. 4 Peak contrast against jitter, for two filter widths that differ by a factor of two. Both halve at a jitter of about fourteen per cent of the period — 0.7 milliseconds at two hundred clicks a second — and the two curves differ from each other by less than a tenth of that, so the collapse is a property of the jitter and not of the modelling choice about the filter. The shaded strip is the range in which the literature reports listeners losing the pitch.

The threshold is stated as a fraction of the period rather than as a number of milliseconds, and that is a claim which one rate cannot support.

What jitter does to the only cue a click train has. The contrast of the autocorrelation peak against the lags around it, against the jitter applied to the click times, at 100 clicks a second and averaged over 8 seeded trains. Through a filter centred at 4000 Hz — a ring of 2.2 ms — the contrast halves at a jitter of 15 per cent of the period, 1.49 ms. Through a filter centred at 8000 Hz — a ring of 1.1 ms — the contrast halves at a jitter of 11 per cent of the period, 1.07 ms. The two filter widths differ by a factor of two and their half-contrast points differ by less than a tenth of that, so the collapse is a property of the jitter and not of the modelling choice. The shaded strip is the range in which the literature reports listeners losing the pitch, which is a measurement and not a computation.
Fig. 5 The same computation at half the rate. The contrast halves at 15 per cent of the period through the narrower filter and 11 through the wider one — the same fraction as at two hundred a second, to within the scatter of eight seeded trains — but the fraction is now 1.5 milliseconds rather than 0.7, because the period is twice as long. So the quantity the mechanism is sensitive to is a proportion of the interval it is measuring and not an absolute timing error, which is what a correlation predicts and what a fixed-resolution clock would not.

The model’s own answer is that the cue halves at a jitter of roughly a seventh of the period. The listener’s answer is a published measurement and it is quoted as a range rather than a number, because it depends on the rate, the level, the band and the task, and because reports differ between listeners by more than they differ between experiments. What the literature broadly says is that the pitch of a jittered train weakens progressively and becomes unusable somewhere between a tenth and a quarter of the period.

It is worth being explicit about what the argument survives that being wrong by. The claim here is not that the threshold is fourteen per cent. It is that a jitter small enough to leave the average rate and the long-term spectrum alone destroys the pitch, and that no spectral account has anything to say about why. A published threshold anywhere from five per cent to fifty per cent leaves that intact. A finding that the pitch survives jitter of a whole period would refute it, and nobody reports that.

There is one measurement the model makes that a published threshold could contradict rather than merely fail to confirm, and it is the scaling. The collapse happens at a fixed fraction of the period at both rates tested, which is what a correlation predicts; a mechanism with a fixed timing resolution — a clock rather than a correlator — would collapse at a fixed number of microseconds, and that is a different curve on the same axes. Two published thresholds at two rates would separate them, and the difference between the predictions at a hundred and two hundred clicks a second is a factor of two, which is well outside the scatter of anything here.

The window the experiment lives in is narrower than it looks

The stimulus needs three things to be true at once, and only the first is usually stated.

The partials in the band have to be unresolved, which needs a bandwidth wider than the repetition rate. The clicks have to be separate events in the filter’s output, which needs a bandwidth wider than about twice the rate: a filter that rings for longer than the gap between clicks turns them into a continuous tone and there is nothing left to correlate. And the nerve has to be locking to the waveform at all, which fails above about five kilohertz.

Where a 200-a-second train can have a pitch nothing spectral explains. The contrast of the autocorrelation peak for an unjittered 200-a-second click train, against the centre frequency of the band it is passed through. Three boundaries bound the experiment. Below 1624 Hz the harmonics are resolved and a spectral account is available. Below about 3477 Hz the filter rings for longer than the gap between clicks and its output is a tone rather than a series of events, which is why the curve is flat on the floor there. And above 5 kHz the nerve no longer locks to the waveform, which is a published measurement. What is left is 3477 to 5000 Hz, a window 0.52 octaves wide, and it is the region the classic high-harmonic experiments used.
Fig. 6 Peak contrast against the centre of the band the train is passed through. Below 1,624 Hz the harmonics are resolved and a spectral account is available, so the experiment proves nothing. Below 3,477 Hz the filter rings for longer than the gap between clicks, its output is a tone rather than a series of events, and the correlation has nothing to find — which is why the curve sits on the floor there. Above five kilohertz the nerve stops locking. What survives all three is less than half an octave wide.

Between 3.5 and 5 kilohertz, and nowhere else, a two-hundred-a-second train has a pitch that no spectral mechanism can be reading and a temporal mechanism can. That is where the classic high-harmonic experiments were run, and the arithmetic says it was not a free choice.

It also says why two hundred a second is close to the highest rate anybody used. Half an octave of usable band is not much room for a filter with skirts on it, and the next section finds where the room runs out entirely — which is close enough to two hundred that the choice of rate in the published work looks less like a convention than like the edge of what the stimulus allows.

Where a 100-a-second train can have a pitch nothing spectral explains. The contrast of the autocorrelation peak for an unjittered 100-a-second click train, against the centre frequency of the band it is passed through. Three boundaries bound the experiment. Below 698 Hz the harmonics are resolved and a spectral account is available. Below about 1624 Hz the filter rings for longer than the gap between clicks and its output is a tone rather than a series of events, which is why the curve is flat on the floor there. And above 5 kHz the nerve no longer locks to the waveform, which is a published measurement. What is left is 1624 to 5000 Hz, a window 1.62 octaves wide, and it is the region the classic high-harmonic experiments used.
Fig. 7 The same three boundaries at a hundred clicks a second. Everything moves down together except the phase-locking limit, which is a property of the nerve and not of the stimulus, so the window opens: it runs from 1,624 to 5,000 Hz and is 1.62 octaves wide against 0.52 at twice the rate. The experiment is comfortable at this rate and cramped at the last one, and the reason is that only two of its three boundaries follow the rate.

And it closes altogether at 283 a second

Two boundaries rising and one fixed is a window that must shut, and solving for where it does is one line.

rate separate events above window to 5 kHz
50 698 Hz 2.84 octaves
100 1,624 1.62
200 3,477 0.52
250 4,403 0.18
283 5,000 0
300 5,330 closed

Above about 283 clicks a second there is no band at all in which the experiment can be run. Below that frequency the filter rings for longer than the gap between clicks and the correlation has nothing to find; above it the nerve is not locking. The two conditions cross at a repetition rate of 283 hertz, which is a semitone below D4.

That is a bound on the evidence rather than on the phenomenon, and the distinction matters. Nothing here says a temporal mechanism stops working above 283 hertz; the phase-locking limit is at five kilohertz and a residue pitch is reported up to about 1.4. What it says is that the stimulus that isolates the temporal account from the spectral one does not exist above 283 hertz, so every direct demonstration of timing-based pitch is a demonstration in the bottom two octaves of the musical range, and the account is extended upward by assumption.

Which puts a shape on how far this rung’s conclusion travels. The bass, where the previous rung found that a note’s resolved harmonics collapse, is exactly the register in which the temporal account can be shown to work — and it is also, by that rung’s arithmetic, the register in which the spectral account has least to work with. The two results are the same register seen from opposite sides, and the coincidence is not one: both are consequences of the critical band being wide compared with a low fundamental.

Above middle C the position reverses, and neither result applies. There the partials are resolved and the spectral account is comfortable, and the experiment that would test the alternative cannot be built.

The window also closes with rate. At a hundred clicks a second it runs from 1.6 kilohertz to 5 and is comfortable; at three hundred it has almost shut; above about two hundred and fifty a second there is no band in which both conditions hold, because the bandwidth needed to keep the clicks separate is above the phase-locking limit. The experiment is possible only for repetition rates in the bottom two octaves of the musical range — which is, not by accident, the register where the residue is used and where resolvability is worst.

Where the model stops, and it stops on this ladder’s own preferred account

Two things have to be said here, and the second is uncomfortable for the position this anchor has been holding.

The first is that autocorrelation is a model and not a mechanism, exactly as the least-squares template fit is. Nothing in the auditory system computes s(t)s(t+τ)\sum s(t)s(t+\tau). Neural versions of the idea are built out of coincidence detectors and delay lines, the delay lines have never been found anatomically, and the whole family is argued about. The figures here state what a correlation predicts and not what happens.

The second is that the experiment which made the pattern account necessary does not separate the pattern account from this one.

The repeat rate of 3 steady partials. The autocorrelation of a sum of steady tones at 1840.0, 2040.0, 2240.0 Hz, evaluated in closed form as the amplitude-weighted sum of cosines. Its highest peak away from zero lag reaches 0.995 at 4.9 milliseconds — the period of 203.97 Hz, a frequency that is not among the tones — because a sum of harmonics of f repeats f times a second whichever of them are present, and a set that fits a series only approximately repeats only approximately. Nothing about how well the ear could use that repeat is in this drawing; the correlation is of the pressure, and the ear correlates the output of a filter.
Fig. 8 The shifted-residue stimulus — the three partials whose pitch an earlier essay uses to rule out difference tones — put through the same correlation. Its highest peak is at 4.9 milliseconds, which is 204 hertz, not the 200 hertz of the partials’ spacing. That is the shift listeners report, and it is what the least-squares template fit also predicts, to within a quarter of a hertz. The experiment refutes Helmholtz. It does not choose between Schouten’s pattern and a correlation.

The correlation gets the shifted residue right, and it gets the ambiguity right too: its local maxima in that window sit at 226, 204, 186 and 170 hertz, which are exactly the four fundamentals a template search over the same partials returns. Two models that were introduced as rivals turn out to agree on the case that was supposed to arbitrate them, and to agree on the errors as well.

Run the template search on the click train’s own partials — six of them, from the fifteenth harmonic up — and on paper it works: the series fits 200 hertz exactly and nothing else comes close. On paper is the whole difficulty, because the search is being handed partials the ear never separated.

So the jittered train is doing work the shifted residue cannot. It is the case where the two accounts come apart, because a template match over the long-term spectrum is completely indifferent to click timing and a correlation is not.

What the picture cannot show

The correlation drawn is of the pressure in the air. A real mechanism would correlate the output of each auditory filter separately and sum the results, which is a stronger and messier object — it weights each channel by how well that channel resolves anything, and it is what the published summary-autocorrelation models compute.

The octave problem is drawn and not solved. A correlation has maxima at every multiple of the period, so nothing in it prefers 200 hertz to 100. Every model in this family needs a further rule to pick the first peak, and the rules are ad hoc. That the same octave ambiguity shows up in the template account, and in an organ’s two pipes, is a point in favour of the ambiguity being real rather than of any one model.

Level is absent from all of it, and level moves both boundaries. Auditory filters broaden as the input grows, so a loud train is unresolved over a wider range and its filters ring for a shorter time — which widens the usable window at both ends. A figure drawn at one level is a figure about one level, and the same caution applies here as to every masking measurement on this site.

And the jitter is Gaussian and independent, which no real timing noise is. A drifting rate, a rate modulated by something else, and a rate with occasional missing clicks all leave different marks on a correlation, and all three occur in sounds people actually hear.

Above three kilohertz the partials stop being separable, and in this site’s usual units the reason is stark: adjacent harmonics of a 200 Hz train up at four kilohertz are eighty-four cents apart — less than a semitone — while the ear’s band up there is several times that wide. Nothing in the periphery is resolving them.

The same ceiling, met from a completely different direction

The five-kilohertz limit in these figures is not a fact about pitch. It is a fact about the auditory nerve: above roughly four to five kilohertz a nerve fibre’s firings stop being locked to the phase of the waveform and carry only its envelope and its average rate.

That property has a second consequence somewhere else entirely on this site. Locating a sound by the difference in arrival time at the two ears requires comparing the phase of two waveforms, and it fails at the top of the spectrum for the same reason — which is why localisation switches from a timing cue to a level cue partway up, and why the duplex theory has two halves rather than one.

So the ceiling on residue pitch and the crossover in binaural localisation are the same measurement seen twice, in two fields that do not otherwise touch. Neither is a property of pitch or of space; both are the point at which a nerve fibre stops being a clock.

Whose sound, and where it occurs

A band-passed click train is a laboratory object, and its musical relatives are real enough.

Vocal fry, the creaky register at the bottom of a speaking voice, is a train of glottal pulses at rates between about thirty and eighty a second, and its pulses are irregular. That irregularity is what makes fry sound rough rather than pitched, and it is the same manipulation as the figures above, applied by a larynx rather than by a random-number generator. A singer descending into fry crosses the boundary this essay computes.

A rattle, a snare buzz or a badly seated reed is the same object again at a rate too low or too irregular to have a pitch at all — as is an untuned drum, from the other end of the argument: a membrane’s partials support no common fundamental, so its waveform has no repeat for a correlation to find and no series for a template to fit, and both accounts return nothing together. And going the other way, a rhythm accelerated past about twenty events a second becomes a pitch, which is this essay’s stimulus approached from the bottom of the same continuum rather than from the top.

The engineering case is the one the anchor keeps returning to. A bass note through a small loudspeaker arrives with every resolved partial removed and a set of unresolved upper harmonics left, and what those harmonics have in common is that their sum repeats at the rate of the note. The next rung computes that case — and it is the one place where the argument of this essay is worth money.

Where the ladder goes next

This rung was written to supply evidence for the temporal account, and it does, and the honest summary of the phase is that it bounds the ladder’s own preferred model from the other side. The pattern account is not wrong. It is under-determined: it is one of at least two models that predict the shifted residue, the octave errors and the dominance of the low harmonics equally well, and the case that separates them is a stimulus nobody would call music.

What is left is a division of labour rather than a winner. Where partials are resolved, a template can be matched and the pitch is strong and definite. Where they are not — at the bottom of the bass, through a small speaker, in an organ’s resultant, in a click train above three kilohertz — a template has nothing to work on and the timing is all there is, and the pitch that results is exactly the weak, ambiguous, octave-prone thing the literature describes.

The last rung of this anchor takes that division into an engineering decision: what a three-inch driver leaves of a forty-hertz note, why the note is heard anyway, and what a manufacturer is really buying when a device adds harmonics of a bass it cannot reproduce.

Part 7 of 9

One essay in the series on missing fundamental. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AutocorrelationCritical bandwidthJitterMissing fundamentalPeriodicityResidue pitchResolvabilityTemporal coding