Intervals and chords

The ear makes its own sound, and it is not the missing fundamental

Play two loud tones and a third pitch appears that is in neither of them. The ear is not a passive analyser: it is nonlinear, it generates frequencies of its own, and it emits sound back out of the ear canal. None of which explains the missing fundamental — the products land in the wrong place, and finding out where they land is the experiment that made the residue theory necessary.

Assumes: The note that is not there

Play two steady tones, reasonably loud, a few semitones apart. A third pitch appears, well below both, and it is not in the signal. Giuseppe Tartini described it in 1714, taught violin students to use it as an intonation check, and gave it a name it still carries in Italian: il terzo suono, the third sound. Georg Sorge described the same thing independently a few years earlier, which is a small reminder that a phenomenon obvious enough to be noticed by two people in a decade had nevertheless gone unremarked for the whole history of two-part playing before it.

1000 and 1200 hertz, and what the ear addsTwo tones presented to a listener, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions: they are products of the pair, at f₂ − f₁ = 200 Hz, 2f₁ − f₂ = 800 Hz, 3f₁ − 2f₂ = 600 Hz, 2f₂ − f₁ = 1400 Hz. The cubic difference tone sits just below the lower primary and is audible at modest levels; the quadratic one is far below both and needs a loud pair.f₂ − f₁2003f₁ − 2f₂6002f₁ − f₂800f₁1000f₂12002f₂ − f₁1400hertzin the air:two tonesin the ear:4 moref₂/f₁ = 1.20the cubic product islargest near 1.22
Fig. 1 Two tones at 1,000 and 1,200 Hz, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions. The quadratic difference tone at 200 Hz is Tartini’s, sits far below both primaries and needs a loud pair; the cubic tone at 800 Hz sits just below the lower primary, is much closer, and survives at levels where Tartini’s has vanished.

Where the extra frequencies come from

Any system whose output is not exactly proportional to its input generates frequencies that were not present. Put two sinusoids through a device with a squared term and the output contains their sum and their difference; put them through one with a cubed term and it contains 2f₁ − f₂ and 2f₂ − f₁ as well. This is ordinary distortion, familiar from any overdriven amplifier, and the ear does it.

The arithmetic is worth doing once, because it is what distinguishes the two products.

The quadratic difference tone, f₂ − f₁, comes from the squared term. Here that is 200 Hz — a long way below both primaries, in a completely different part of the spectrum. Its level grows steeply with the primaries’ level, so it is loud when they are loud and gone when they are not.

The cubic difference tone, 2f₁ − f₂, comes from the cubed term. Here that is 800 Hz — just below the lower primary. Its level grows only slowly with the primaries’, so it survives at modest levels, and it is the one that matters clinically.

The cubic tone is also strongest at a particular frequency ratio. As f₂/f₁ is increased from 1, its level rises to a maximum near 1.22 and then falls away — a fact with no obvious explanation from the arithmetic alone, and one of the pieces of evidence that the nonlinearity is a specific mechanism rather than generic distortion.

1000 and 1220 hertz, and what the ear adds. Two tones presented to a listener, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions: they are products of the pair, at f₂ − f₁ = 220 Hz, 2f₁ − f₂ = 780 Hz, 3f₁ − 2f₂ = 560 Hz, 2f₂ − f₁ = 1440 Hz. The cubic difference tone sits just below the lower primary and is audible at modest levels; the quadratic one is far below both and needs a loud pair.
Fig. 2 The same pair at the ratio where the cubic product is largest, 1.22. The quadratic tone has moved up to 220 Hz and the cubic down to 780. Both positions follow from the arithmetic and neither is where a naive listener expects a “difference tone” to be — the cubic one in particular sits between the primaries and their octave below, close enough to the lower primary to be partly masked by it, which is one reason it went unnoticed for two centuries after Tartini described the other.

The two products behave so differently that treating them as one phenomenon is the standard confusion in this area. A useful summary: the quadratic tone is what Tartini heard and needs volume; the cubic tone is what a clinic measures and needs sensitivity. They have different frequencies, different level dependences, different masking situations, and different uses.

The evidence that it is really in the ear

Two demonstrations settle that these tones are generated inside the listener rather than in the air, in the loudspeaker, or anywhere else.

Cancellation. Add a third real tone at the combination tone’s frequency, and adjust its amplitude and phase. At the right setting the combination tone disappears. It can be cancelled, which means it is a physical vibration somewhere with a definite amplitude and phase — not an imagined pitch. And the cancelling tone that removes it is a real tone, so if the combination tone were merely a perceptual inference the two would not annihilate.

Emission. Put a sensitive microphone in the ear canal, present two tones, and the microphone records a tone at 2f₁ − f₂ coming out of the ear. This is a distortion-product otoacoustic emission, it is measurable in most healthy ears, and it is direct evidence that something inside is generating acoustic energy rather than only absorbing it.

The emission result is the more remarkable of the two and it changed the subject. A passive filter bank cannot emit; something in the cochlea is an amplifier, running on metabolic energy, with enough gain to push sound back out through the middle ear. That amplifier — the outer hair cells, changing length in response to voltage — is what gives the ear its sensitivity at low levels and its sharp tuning, and its nonlinearity is the price.

The clinical consequence is that a newborn’s hearing can be screened without asking the newborn anything. Present two tones, listen for the emission, and its presence indicates a working cochlear amplifier. That test is now routine in many countries and it exists because the ear distorts.

1000 and 1200 hertz, and what the ear addsTwo tones presented to a listener, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions: they are products of the pair, at f₂ − f₁ = 200 Hz, 2f₁ − f₂ = 800 Hz, 3f₁ − 2f₂ = 600 Hz, 2f₂ − f₁ = 1400 Hz. The cubic difference tone sits just below the lower primary and is audible at modest levels; the quadratic one is far below both and needs a loud pair.f₂ − f₁2003f₁ − 2f₂6002f₁ − f₂800f₁1000f₂12002f₂ − f₁1400hertzin the air:two tonesin the ear:4 moref₂/f₁ = 1.20the cubic product islargest near 1.22
Fig. 3 The same pair with the residue drawn beside the products, which is what makes the two accounts look alike. Every adjacent pair of a harmonic series on 200 hertz differs by 200, so a difference tone between any two neighbours lands on the fundamental — and the ear’s own f₂ − f₁ lands there too. Two mechanisms, one prediction, from the same stimulus. That coincidence is why an account based on difference tones held the field for so long, and it is why separating them takes a stimulus that is not a harmonic series.

Why this does not explain the missing fundamental

Here is the connection that makes this a rung of the missing-fundamental ladder rather than a curiosity of its own.

Helmholtz’s account of residue pitch was exactly this mechanism. Present partials at 800, 1,000 and 1,200 Hz with no 200 Hz component, and the ear hears a pitch at 200 Hz. Helmholtz proposed that the ear’s nonlinearity generates a 200 Hz component from the differences between adjacent partials, and that the listener then hears the real, internally generated tone. Elegant, mechanical, and testable.

It is wrong, and the test that showed it is the one the previous rung is built on.

A note with its first 8 partials removed. The spectrum of a 200 Hz tone with the lowest 8 partials deleted, and the wave that remains. The wave still repeats 200 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 4 The residue experiment: a set of partials with the lower ones deleted. The pitch heard corresponds to the spacing of the partials rather than to any frequency present. The distortion-product account predicts this correctly — the difference between adjacent partials is exactly the spacing — which is why it survived as an explanation for so long.

Now shift every partial up by the same amount. Take partials at 1,840, 2,040 and 2,240 Hz: their spacing is still 200 Hz, so a distortion product at the difference frequency is still at 200 Hz exactly. The distortion account predicts the pitch does not move.

A note with its first 8 partials removed. The spectrum of a 200 Hz tone with the lowest 8 partials deleted, and the wave that remains. The wave still repeats 200 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound.
Fig. 5 The same partials shifted up by 40 Hz each. The spacing is unchanged, so any difference tone is unchanged — and the pitch heard does move, to about 204 Hz. That is a small shift and it is decisive: it is what a pattern-matching account predicts, fitting the best harmonic series to the partials present, and it is flatly incompatible with a difference tone at the spacing.

The measured shift matches the best-fitting fundamental rather than the spacing. The pitch is being computed from the pattern of partials, not read off a frequency the ear manufactured, and the distortion products — which really are there — are a separate phenomenon that happens to have the right value in the unshifted case and the wrong one as soon as the case is varied.

That is a clean example of a good hypothesis being killed by a variation rather than by a direct test. Helmholtz’s account and the pattern account agree on every unshifted stimulus, and they disagree by four hertz on a shifted one. Four hertz at 200 is about 34 cents, which is well above anything a listener could miss — so the experiment is not delicate, it simply had to be thought of.

The shift is not one number, and that is the second half of the test

Four hertz is the answer for the particular partials drawn above, and the more discriminating fact is that it is not the answer for any others. Keep the shift at forty hertz and the spacing at two hundred, and move the partials up and down the series:

harmonics partials best-fit fundamental shift
3, 4, 5 640, 840, 1040 209.6 Hz 81 cents
4, 5, 6 840, 1040, 1240 207.8 66
6, 7, 8 1240, 1440, 1640 205.6 48
9, 10, 11 1840, 2040, 2240 204.0 34
16, 17, 18 3240, 3440, 3640 202.4 20

Every row has the same spacing, so the distortion account predicts two hundred hertz in all five, and the pattern account predicts five different pitches spanning sixty cents. The shift falls as one over the harmonic number — a least-squares fit is dominated by the highest partials, so the same absolute displacement is a smaller fractional one the further up the series it sits, and the classical approximation of the displacement divided by the mean harmonic number reproduces the whole column to within two per cent.

That turns the experiment from a coin toss into a curve. A single shifted stimulus gives one number and could be explained away — mistuned equipment, a listener’s bias, a poorly matched comparison tone. Five shifted stimuli give a slope, and the slope is one over n, which no account based on the spacing of the partials can produce at all, since the spacing is identical in every condition.

It also says which stimulus to use, and the answer is not the one usually drawn. The effect is largest at the bottom of the series — eighty-one cents at harmonics three to five, against thirty-four at nine to eleven — so an experimenter wanting the clearest separation should use low harmonics, and the reason the literature’s canonical stimulus sits higher is that low harmonics are resolved individually and a listener may report one of them instead of the residue.

Beats, difference tones, and the thing they are not

Three phenomena get confused with each other constantly and they are cleanly separable, which is worth doing once here because two of them already have essays.

440 and 444 hertz, and what the ear adds. Two tones presented to a listener, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions: they are products of the pair, at f₂ − f₁ = 4 Hz, 2f₁ − f₂ = 436 Hz, 3f₁ − 2f₂ = 432 Hz, 2f₂ − f₁ = 448 Hz. The cubic difference tone sits just below the lower primary and is audible at modest levels; the quadratic one is far below both and needs a loud pair.
Fig. 6 The same figure with the two tones four hertz apart, which is where the first of the three phenomena lives. The ear’s arithmetic still runs — 2f₁ − f₂ at 436, 2f₂ − f₁ at 448 — but f₂ − f₁ is 4 hertz, and 4 hertz is not a pitch. It is a throb, and it is in the air rather than in the ear: a microphone records it, and it is arithmetic on the sum of two waves with no listener required. So the same drawing produces beating at one spacing and a difference tone at another, and nothing about the mechanism changed between them except the gap.

Beating is a property of the sum of two waves and exists outside any listener. It is heard when the difference is under about 20 Hz.

A difference tone is a frequency generated inside the ear and does not exist in the air. It is heard as a pitch, so it requires the difference to be above about 20 Hz, and it requires the primaries to be loud.

And a residue pitch is neither: no frequency anywhere, in the air or in the ear, corresponds to it. It is the output of a pattern-matching computation, and the shifted-partials experiment is what proves it.

The three form a neat progression from most physical to least, and a listener presented with two tones and asked what they hear may be reporting any of them depending on how far apart the tones are and how loud they are. That is the practical reason the literature took two centuries to sort out: the same experimental setup produces all three, and which one is being described was frequently left to the reader.

Which computation produced the numbers

The combination tone frequencies are arithmetic and the figure computes them: f₂ − f₁, 2f₁ − f₂, 2f₂ − f₁, 3f₁ − 2f₂. Nothing about their positions is quoted.

The residue figures compute the shifted fundamental by least-squares fit over the partials present, which is where the 204 Hz comes from — and that computation has its own recorded history on this site. The fit needs eleven partials while the named timbres are eight long, so an earlier version sliced an empty set of present partials and produced a fitted fundamental of NaN, printed into a caption. The series now continues as one over n past the end of the table.

The levels of the combination tones are not computed and the figure does not draw them to scale. Predicting how loud a distortion product is requires a model of the cochlear amplifier, which is a substantial piece of physiology and not arithmetic over two frequencies. What the figure asserts is where the products fall, which is exactly what the essay’s argument needs.

What an active cochlea explains that a passive one does not

The nonlinearity is a side effect of something the ear needs, and the list of what it buys is long enough to make the trade obviously worth it.

Sensitivity. A passive cochlea would have a threshold tens of decibels higher than the real one. The amplifier supplies most of the gain at the quietest levels, which is why the threshold of hearing sits where it does and why it rises so sharply when outer hair cells are damaged.

Sharp tuning. The mechanical tuning of a dead cochlea is broad. A living one is sharp, and the sharpening is active — which is why the critical band is as narrow as it is, and why bandwidth broadens with level and with hearing loss.

And compression. The amplifier’s gain falls as level rises, compressing a 120-decibel range of input into a much smaller range of mechanical response. That is why masking spreads further upward at high levels: the compression flattens the tuning of each place as the input grows.

Every one of those is a consequence of a mechanism that is nonlinear by construction, and a nonlinear mechanism generates combination tones whether anybody wants it to or not. Tartini’s third sound is the audible by-product of the machinery that makes quiet sounds audible at all — which is a good example of a general pattern in perception, where the illusions and the artefacts are usually the exhaust of something useful.

What the picture cannot show

Level, which decides everything about audibility. Whether a listener hears Tartini’s tone at all depends on how loud the primaries are, and the two products behave completely differently as level changes. A figure drawn at one level tells nothing about the others.

The frequency ratio’s effect. The cubic tone’s dependence on f₂/f₁ — rising to a peak near 1.22 and falling away — is one of the most characteristic features of the phenomenon and none of it is in a single-ratio drawing.

The five-row table is a prediction and not a measurement. The pattern account’s column is a least-squares fit computed here; the distortion account’s is two hundred by construction; and no listener was asked. What the published work reports is the first row of the argument rather than the slope — the shift exists and goes the right way — and the one over n dependence is stated in the literature as an approximation rather than measured across a range as wide as the one drawn. The table’s value is that it says what a discriminating experiment would look like, which is five conditions rather than one.

And the emission itself. The strongest evidence in this essay is a measurement made with a microphone in an ear canal, and no drawing of a spectrum shows it. What can be said is what the figure does say: the frequencies are in the wrong places to be anything but products of the pair.

The listener’s own ear is also the apparatus, which is the one respect in which the sound buttons on these figures are doing something the figures cannot. A reader who plays the two primaries loudly on a good pair of headphones may hear the third sound; a reader who plays them quietly, or on a small loudspeaker whose own distortion is larger than the ear’s, will hear something else entirely. This is the only place on this site where a sound button’s result depends on the reproduction chain in a way that could actively mislead, and it is why the buttons here play the products separately for comparison rather than asking the reader to take an unverifiable experience as evidence.

Whose practice, and how far the violin lesson goes

Tartini taught the third sound as an intonation aid: play a double stop, listen for the low tone, and adjust until it is where it should be. The technique is still taught and it works — for loud, sustained, just-tuned double stops in the middle of the instrument’s range, which is a narrow and specific condition.

Outside it the technique fails, and the failures follow from the figure. At quiet dynamics the quadratic tone is not generated at any useful level. On tempered rather than just intervals the product is a few cents off and much harder to identify. And on an instrument or in a register where the primaries are far apart, the difference tone lands so low it is below the audible range.

196 and 245 hertz, and what the ear adds. Two tones presented to a listener, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions: they are products of the pair, at f₂ − f₁ = 49 Hz, 2f₁ − f₂ = 147 Hz, 3f₁ − 2f₂ = 98 Hz, 2f₂ − f₁ = 294 Hz. The cubic difference tone sits just below the lower primary and is audible at modest levels; the quadratic one is far below both and needs a loud pair.
Fig. 7 Tartini’s own case, drawn: a just major third on the violin’s G and the note a 5:4 above it, at 196 and 245 hertz. The quadratic difference tone lands on 49 hertz — two octaves and a bit below the lower string, at the very bottom of the instrument’s implied bass — and the cubic one at 147, which is a fourth below the lower note and sits comfortably in the audible middle. That is why the technique works on this interval and fails as the primaries separate: widen the double stop and f₂ − f₁ climbs into the texture, narrow the register and it drops out of hearing altogether.

So a practice that looks like folklore turns out to be an exact piece of applied psychoacoustics with a stated domain of validity, and violinists have been operating inside the domain without a way of writing it down. The essay’s honest position is that the technique is real, the mechanism is understood, and the claim that a player is tuning by difference tones in ordinary playing is much stronger than the evidence supports.

Where the ladder goes next

The missing-fundamental anchor now has two rungs, and the second was written to rule out an explanation of the first. What is left open is the mechanism that does the work: pattern matching over resolved partials is one account, autocorrelation of the timing of nerve firings is another, and the two make different predictions for spectra whose partials are too close together to resolve.

The competing accounts also differ in a way the difference limen can arbitrate, which is worth saying because it is unusual for two theories of pitch to make predictions a few cents apart and unusually convenient that they do — the shifted-residue experiment works because the disagreement is thirty-four cents rather than three.

Sideways, this rung supplies the machinery for a different argument entirely. An ear that generates frequencies of its own is an ear whose output spectrum is not its input spectrum, which means every masking measurement has an extra term in it — a listener detecting a probe beside a loud masker may be detecting a distortion product instead, and the experiments have to be designed around it.

Part 2 of 9

One essay in the series on missing fundamental. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 11.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Cochlear nonlinearityCombination toneDifference toneDistortionOtoacoustic emissionResidue pitch