Intervals and chords

The note that is not there

A telephone reproduces nothing below about three hundred hertz, and a bass line down at eighty comes through it perfectly. The pitch that is heard is not a frequency present in the sound, and one nineteenth-century experiment settles what it is instead.

Assumes: A string does everything at once · The ear hears the list, not the shape

The lowest string of a bass guitar is tuned to about 41 hertz. A laptop speaker of the ordinary kind reproduces essentially nothing below 200 hertz, and a phone’s earpiece nothing below about 300. Bass lines are nevertheless perfectly audible on both, at the right pitch, and nobody notices anything missing.

The usual explanation is that the speaker is better than it looks. It is not. The energy at 41 hertz is genuinely absent, and it can be filtered out deliberately, with a steep filter, and the pitch does not move.

A note with its first partial removed. The spectrum of a 220 Hz tone with the lowest partial deleted, and the wave that remains. The wave still repeats 220 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 1 The spectrum of a 220 Hz note with its lowest partial deleted, and the wave that remains. The crossed-out bar is gone from the sound entirely. The wave underneath still repeats 220 times a second, because the repeat rate of a sum of harmonics is set by the spacing between them rather than by whether the lowest is present.

The argument

A note from any harmonic instrument is a sum of partials at whole-number multiples of a fundamental. Remove the first of them and the remainder is still a sum of harmonics of that fundamental — the second, third, fourth and so on — and a sum of harmonics of f0f_0 repeats f0f_0 times a second whatever subset of them is used.

The period is a property of the set, not of any member of it. Deleting the lowest changes the shape of the wave and does not change how often the shape recurs.

So the perceived pitch of a note is not a frequency in the signal. It is something the auditory system computes from the whole set, and it does not need the member the note is named after.

Which computation produced the number

The wave in that figure is drawn by evaluating

p(t)=nSansin(2πnf0t)p(t) = \sum_{n \in S} a_n \sin(2\pi n f_0 t)

over two periods of f0f_0, once with S={1,2,,8}S = \{1,2,\dots,8\} and once with S={2,,8}S = \{2,\dots,8\}. The two curves are drawn on the same axis and the vertical rules mark t=0t = 0, t=1/f0t = 1/f_0 and t=2/f0t = 2/f_0. Both curves cross those rules in the same state, because every term in the sum has a whole number of cycles in that interval by construction.

That is the entire mechanism, and it is arithmetic rather than physiology. The physiology is the question of how the auditory system arrives at the period, and that is where the interesting experiment lives.

Two explanations, and a way to choose

Two accounts of the missing fundamental were available by the middle of the twentieth century, and they make different predictions.

The distortion account. The ear is a nonlinear system, and nonlinear systems fed two tones produce energy at the difference between them. Partials at 440 and 660 would generate a real, physical 220 hertz inside the cochlea, and the ear would then hear it in the ordinary way. Helmholtz argued something close to this.

The pattern account. No energy is generated anywhere. The auditory system receives a set of resolved partials and asks which harmonic series they best belong to, reporting that series’ fundamental as the pitch. Schouten proposed a version of this in 1940 and called what was heard the residue.

The two accounts agree about every ordinary case, which is why they coexisted. They disagree about one contrived case, and the case is easy to build.

A note with its first 8 partials removed. The spectrum of a 200 Hz tone with the lowest 8 partials deleted, and the wave that remains. The wave still repeats 200 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 2 Three partials at 1840, 2040 and 2240 hertz — equally spaced 200 hertz apart, and not harmonics of 200. The spacing is unchanged, so any account that works from difference frequencies predicts a pitch of exactly 200 hertz. The harmonic series that fits them best has a fundamental of 204, and 204 is what listeners report.

Take partials at 1800, 2000 and 2200 — the 9th, 10th and 11th harmonics of 200 hertz — and shift all three up by 40. The result is 1840, 2040, 2240. The spacing is still exactly 200. Nothing about a difference-frequency mechanism has changed at all.

But these are no longer harmonics of 200. They are close to the 9th, 10th and 11th harmonics of 204, and the fit is much better. Fitting a fundamental by least squares over the harmonic numbers present:

f0=nifini2=9(1840)+10(2040)+11(2240)81+100+121=61600302=203.97.f_0 = \frac{\sum n_i f_i}{\sum n_i^2} = \frac{9(1840) + 10(2040) + 11(2240)}{81 + 100 + 121} = \frac{61600}{302} = 203.97.

Listeners hear the pitch go up by about four hertz. This is the pitch shift of the residue, first reported by Schouten and colleagues, and it is decisive: the difference frequency did not change and the pitch did. Whatever the ear is doing, it is not simply listening to a difference tone.

Licklider closed the other door in 1954 by masking the region around the fundamental with noise loud enough to swamp anything generated there. The residue pitch survived. There is nothing at f0f_0 to hear, and removing the possibility of hearing it changes nothing.

A note with its first 4 partials removed. The spectrum of a 261.626 Hz tone with the lowest 4 partials deleted, and the wave that remains. The wave still repeats 261.626 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 3 The same operation on a different source, because the effect is about the pattern and not about the instrument. A clarinet’s list is odd partials, so deleting the lowest four leaves a spectrum whose lowest energy is at five times the fundamental — and the wave that remains still repeats 261.626 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not moved. Neither the difference tone nor the lowest partial nor the loudest is the pitch. The fit is.

What this means for the word “pitch”

The consequence is worth stating carefully, because it is easy to overstate.

Pitch is not the fundamental frequency. Nor is it the lowest frequency present, nor the loudest, nor the spacing between partials. It is the fundamental of the best-fitting harmonic series, and “best-fitting” is doing real work: the shifted-residue case is one where the spacing and the fit disagree, and the fit wins.

That makes pitch a derived quantity, computed from a pattern, and it explains several things that otherwise look unrelated. It explains why a note’s timbre and its pitch are separable, and why the envelope can be changed without touching either — the pattern’s spacing and the pattern’s amplitudes are independent. It explains why an inharmonic instrument has an ambiguous pitch: if no harmonic series fits well, there is no clear answer to return. And it is the reason a bell, whose partials fit nothing, is described as having a “strike note” that is often not present in its spectrum at all.

A note with its first 2 partials removed. The spectrum of a 220 Hz tone with the lowest 2 partials deleted, and the wave that remains. The wave still repeats 220 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 4 Two partials removed from the note the rest of this essay takes apart, drawn with twelve in place of eight so the resolvable and unresolvable ends are both visible. The pitch a listener names is the fundamental of this pattern, and any subset that preserves the spacing carries the same name. What the figure also shows is where the mechanism runs out: the partials at the right-hand end are less than a semitone apart, and the pattern-matching that produces the name can only use the ones it can separate.

How many can be taken away

The deletion does not have to stop at one.

A note with its first 3 partials removed. The spectrum of a 165 Hz tone with the lowest 3 partials deleted, and the wave that remains. The wave still repeats 165 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 5 A 165 Hz note with its first three partials removed, leaving nothing below 660 hertz. The wave still repeats 165 times a second and the pitch is still E, two octaves below the lowest partial present. This is what a telephone does to a voice.

Removing the first three partials of a 165 hertz note leaves the lowest energy at 660 — two octaves above the pitch that is heard. That is a larger gap than the telephone’s, and the pitch survives it, because the spacing is unchanged and the fit is still unambiguous.

There is a limit, and it is not where intuition puts it. The limit is not how far below the lowest partial the fundamental sits; it is whether the surviving partials are resolvable — whether the ear’s frequency analysis can separate them from each other.

The critical band, measured in semitones. The width of the ear's frequency-analysis band at each pitch, converted from hertz into semitones. Two intervals drawn as horizontal lines cross the curves: below the crossing the interval fits inside one band and its notes are not resolved from each other, and above it they are.
Fig. 6 The width of the ear’s analysis band at each pitch, in semitones. Two partials closer together than one band are not separated, and above about the eighth partial of any note the harmonics are closer than that. Those are the ones the pattern-matching cannot use individually.

Partials 2, 3 and 4 of a note are a fifth and a fourth apart — far wider than any critical band, and comfortably resolved. Partials 20, 21 and 22 are less than a semitone apart and fall inside one band, where they produce roughness rather than three separable pitches. A residue built from the high, unresolved end of the series still produces a pitch, weaker and vaguer, and probably by a different mechanism — which is the main reason the field carries two theories rather than one.

Why the partials fuse at all

There is a prior question underneath all of this, and it is the more surprising one. A note from a string arrives at the ear as eight or more separate frequencies, spread over three octaves, each of which the cochlea resolves individually. Why is that heard as one note rather than as a chord?

The answer is that harmonicity is itself a grouping cue. The auditory system treats a set of frequencies that fit one harmonic series as evidence of a single source, because in the physical world that is overwhelmingly what such a set means — one vibrating object. Common onset time and common modulation do the same job. What arrives together, and moves together, and fits one series, is assigned to one thing.

The cleanest demonstration is to break the cue and watch the fusion fail. Take a harmonic complex and mistune a single partial — the fourth, say — by two or three per cent. Below about one per cent nothing happens. Above it, that partial stops belonging to the note and is heard as a separate whistling tone standing alongside it, and the pitch of the remaining complex shifts slightly to reflect a fit that no longer includes it. Moore, Glasberg and Peters reported the effect systematically in 1986, and the threshold is sharp enough to be a useful number.

220 Hz against 223 Hz. Two tones 3 hertz apart, added. The rapid oscillation is their average; the slow swelling is their difference, heard as 3 beats a second and used by every tuner who has ever worked by ear.
Fig. 7 Two tones a few hertz apart, and their sum. A partial mistuned by a small fraction of its frequency beats against nothing — there is no second tone at the correct frequency to beat with — so what a listener notices is not a pulsing but a change in what belongs to what.

This is why the missing fundamental is not an exotic case. Fusion, pitch extraction and the missing fundamental are the same operation seen from three sides: a decision about which frequencies belong to one source, and what to call the result.

The errors it predicts

A model that is doing real work makes wrong predictions in specific places, and the pattern account does.

Take partials at 800, 1000 and 1200 hertz. They are the 4th, 5th and 6th harmonics of 200 — and also the 8th, 10th and 12th harmonics of 100, and the 2nd, 2.5th and 3rd of 400, which is not a harmonic fit at all. The best fit is 200, and 200 is what most listeners report. But the fit at 100 is exactly as good, in the sense that every partial is an exact harmonic of it, and a minority of listeners hear the octave below.

Octave ambiguity of this kind is a standing feature of residue pitch, and it is exactly what a template-matching account predicts: an incomplete harmonic series is consistent with more than one fundamental, and which one is reported depends on how many partials are present, where they sit, and on the listener. A theory that returned the difference frequency would have no room for the ambiguity at all, and would be wrong.

What breaks the tie, and when it stops working

“Exactly as good” is true of the fit and not of the whole judgement, and the collection’s own residue matcher says what the difference is. It scores a candidate on two things: how far each partial sits from its assigned harmonic, and how many harmonic slots inside the span are left empty.

partials best candidate runner-up what separates them
800, 1000, 1200 4, 5, 6 on 200 8, 10, 12 on 100 fit identical; 2 empty slots
plus 1400 4, 5, 6, 7 on 200 8, 10, 12, 14 on 100 3 empty slots
600, 800, 1000, 1200 3, 4, 5, 6 on 200 6, 8, 10, 12 on 100 3 empty slots
1000, 1200 only 5, 6 on 200 6, 7 on 169 0 empty slots; 28 cents of fit
2000, 2200, 2400 10, 11, 12 on 200 11, 12, 13 on 183 0 empty slots; 15 cents of fit

The octave below is refused on parsimony, not on accuracy. Every partial of 800, 1000, 1200 is an exact harmonic of 100 as well as of 200, so the fit cannot choose; what chooses is that reading them as the 8th, 10th and 12th leaves the 9th and 11th unaccounted for, and reading them as the 4th, 5th and 6th leaves nothing out.

That makes the essay’s “depends on how many partials are present” a quantity. Adding a partial widens the gap — one more harmonic takes the octave-below reading from two empty slots to three — so a fuller series is progressively harder to hear an octave down, which is the direction listeners report.

The last two rows are where the account gets uncomfortable, and they are the interesting ones. Given only two partials, a candidate with no empty slots at all sits 28 cents away at a completely unrelated fundamental — so parsimony has nothing left to say and the ambiguity is not an octave ambiguity but a general one. And a residue built from high consecutive partials is in the same position: 10, 11, 12 on 200 is beaten by nothing on parsimony and by only fifteen cents on fit.

So the tie-breaker that makes an ordinary residue unambiguous stops working exactly where the essay says the effect gets weak — few partials, or high ones. The vagueness of a high residue is not a separate fact about the ear; it is the parsimony term going to zero and leaving the fit to decide alone.

Where the model stops

The effect needs resolvable partials, and there are not many. The auditory system separates partials up to roughly the eighth or tenth; above that they fall inside one critical band and are not individually available. A residue built entirely from the 20th to 25th harmonics produces a much weaker and less definite pitch, which is a real limit of the pattern account and one of the arguments for a second, temporal mechanism running alongside it.

Pitch strength is not the same as pitch. A note with its fundamental removed has the same pitch and a thinner, more nasal quality, and it is easier to mishear by an octave. The missing fundamental is a robust effect and not an indistinguishable one, and any claim that a small speaker is “just as good” is a claim about pitch, not about sound.

The parsimony term is a model too, and a cruder one than the fit. Counting empty harmonic slots is a plausible way to prefer 4, 5, 6 over 8, 10, 12 and it is not a measurement of anything: nothing says a listener weighs one empty slot against a cent of mistuning at the exchange rate the matcher uses. What the table above establishes is the structure of the judgement — that the octave ambiguity is settled by something other than fit, and that whatever settles it has nothing left to say when the partials are few or high — rather than the size of either term.

The least-squares fit is a model, not a mechanism. Nothing in the auditory system computes nifi/ni2\sum n_i f_i / \sum n_i^2. The fit is a compact way of predicting the reported pitch, and it does predict the shifted-residue result correctly, but the real machinery — whether it is autocorrelation on the auditory nerve, or a template match, or both — is still argued about. The figure states what fits, not what happens.

Existence has a region. The residue effect fades out above about 5 kHz and below about 50 hertz, and it depends on how many partials are present and how loud. The literature calls this the existence region, and the boundaries of it are not sharp.

Everything above assumes harmonicity. The argument runs on partials at exact whole-number multiples. Real strings are stiff and their partials run progressively sharp, so the harmonic series that fits a real piano note is a slightly stretched one, and the pitch the ear reports for it is correspondingly slightly displaced.

Whose music, and when

Organ builders were exploiting this three centuries before anybody explained it.

A 32-foot pipe is nine metres long, costs a fortune and does not fit in most buildings. The resultant or acoustic bass stop produces the effect of one by sounding a 16-foot pipe and a 10⅔-foot pipe together — a fundamental and its fifth, which are the second and third partials of a note an octave below. The listener supplies the missing 32-foot note. The technique is documented from the early eighteenth century, and the acoustic explanation arrived in the nineteenth. The pipe pair is the second and third partials of the note the listener supplies, which are the octave and the twelfth — the two intervals a vibrating body produces first, and the two that fuse most readily.

Two related observations from the same trade. The quintaton stop is designed to have a strong third partial and a weak second, which gives it its distinctive hollow twelfth-above quality — the same pattern-fitting machinery, running on a spectrum that is deliberately ambiguous. And organ mixtures, which sound several upper partials of each note at once, work because those partials fuse into the note rather than being heard as separate pitches. That fusion is the same phenomenon in the opposite direction.

A note with its first partial removed. The spectrum of a 130.813 Hz tone with the lowest partial deleted, and the wave that remains. The wave still repeats 130.813 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 8 The organ builder’s trick reduced to its arithmetic, and drawn three octaves above where an organ does it so that the buttons are audible on an ordinary speaker: a note with its fundamental deleted, leaving only the second and third partials — an octave and a twelfth above a note nobody sounds. On the organ the same three numbers are 16.35, 32.7 and 49 hertz, which is a 32-foot C and the 16-foot and 10⅔-foot pipes that stand in for it. Two pipes, and the listener supplies the note. The residual wave still repeats 32.7 times a second, and the two intervals involved are the octave and the twelfth, which are the two a vibrating body produces first and the two that fuse most readily. The same drawing is the quintaton’s deliberately weak second partial, and it is why an organ mixture is heard as a colour rather than as a chord.

The telephone case is the twentieth-century one. The standard telephone channel passes roughly 300 to 3400 hertz, which excludes the fundamental of every adult male voice and most female ones. Speech remains fully intelligible and speakers remain recognisable, because the pitch is carried by the harmonics the channel does pass. That bandwidth was chosen on cost grounds by engineers who had the psychoacoustic literature available, and it has shaped what a voice sounds like at a distance for a hundred years.

The same trade-off is made every time a small loudspeaker is designed. Bass enhancement in phones and laptops frequently works not by reproducing the low fundamental — which the driver physically cannot do — but by adding upper harmonics of it, so the pattern-matching machinery has something to work from. The device manufactures the evidence for a note it cannot play.

The ladder from here

Later rungs on this anchor: the existence region, and what happens at its edges. Pitch strength as a measurable quantity, and how it is measured. Virtual pitch in inharmonic sounds — bells, gongs and drums, where several fits compete and the answer is genuinely ambiguous. The octave errors that pattern matching predicts and that listeners actually make. Autocorrelation as a rival account, and the experiments that separate it from template matching. And the resultant stop as a piece of instrument design, alongside the other places where a spectrum is engineered rather than merely produced.

A speaker that cannot make a sound below 200 hertz plays a bass line at 41, and the note is not a reconstruction or an illusion in any useful sense. It is what pitch has always been.

Part 1 of 9

One essay in the series on missing fundamental. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 23.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Harmonic seriesMissing fundamentalPeriodicityResidue pitchSpectrum