The ear makes its own sound, and it is not the missing fundamental
Assumes: The note that is not there
Play two steady tones, reasonably loud, a few semitones apart. A third pitch appears, well below both, and it is not in the signal. Giuseppe Tartini described it in 1714, taught violin students to use it as an intonation check, and gave it a name it still carries in Italian: il terzo suono, the third sound. Georg Sorge described the same thing independently a few years earlier, which is a small reminder that a phenomenon obvious enough to be noticed by two people in a decade had nevertheless gone unremarked for the whole history of two-part playing before it.
Where the extra frequencies come from
Any system whose output is not exactly proportional to its input generates frequencies that were not present. Put two sinusoids through a device with a squared term and the output contains their sum and their difference; put them through one with a cubed term and it contains 2f₁ − f₂ and 2f₂ − f₁ as well. This is ordinary distortion, familiar from any overdriven amplifier, and the ear does it.
The arithmetic is worth doing once, because it is what distinguishes the two products.
The quadratic difference tone, f₂ − f₁, comes from the squared term. Here that is 200 Hz — a long way below both primaries, in a completely different part of the spectrum. Its level grows steeply with the primaries’ level, so it is loud when they are loud and gone when they are not.
The cubic difference tone, 2f₁ − f₂, comes from the cubed term. Here that is 800 Hz — just below the lower primary. Its level grows only slowly with the primaries’, so it survives at modest levels, and it is the one that matters clinically.
The cubic tone is also strongest at a particular frequency ratio. As f₂/f₁ is increased from 1, its level rises to a maximum near 1.22 and then falls away — a fact with no obvious explanation from the arithmetic alone, and one of the pieces of evidence that the nonlinearity is a specific mechanism rather than generic distortion.
The two products behave so differently that treating them as one phenomenon is the standard confusion in this area. A useful summary: the quadratic tone is what Tartini heard and needs volume; the cubic tone is what a clinic measures and needs sensitivity. They have different frequencies, different level dependences, different masking situations, and different uses.
The evidence that it is really in the ear
Two demonstrations settle that these tones are generated inside the listener rather than in the air, in the loudspeaker, or anywhere else.
Cancellation. Add a third real tone at the combination tone’s frequency, and adjust its amplitude and phase. At the right setting the combination tone disappears. It can be cancelled, which means it is a physical vibration somewhere with a definite amplitude and phase — not an imagined pitch. And the cancelling tone that removes it is a real tone, so if the combination tone were merely a perceptual inference the two would not annihilate.
Emission. Put a sensitive microphone in the ear canal, present two tones, and the microphone records a tone at 2f₁ − f₂ coming out of the ear. This is a distortion-product otoacoustic emission, it is measurable in most healthy ears, and it is direct evidence that something inside is generating acoustic energy rather than only absorbing it.
The emission result is the more remarkable of the two and it changed the subject. A passive filter bank cannot emit; something in the cochlea is an amplifier, running on metabolic energy, with enough gain to push sound back out through the middle ear. That amplifier — the outer hair cells, changing length in response to voltage — is what gives the ear its sensitivity at low levels and its sharp tuning, and its nonlinearity is the price.
The clinical consequence is that a newborn’s hearing can be screened without asking the newborn anything. Present two tones, listen for the emission, and its presence indicates a working cochlear amplifier. That test is now routine in many countries and it exists because the ear distorts.
Why this does not explain the missing fundamental
Here is the connection that makes this a rung of the missing-fundamental ladder rather than a curiosity of its own.
Helmholtz’s account of residue pitch was exactly this mechanism. Present partials at 800, 1,000 and 1,200 Hz with no 200 Hz component, and the ear hears a pitch at 200 Hz. Helmholtz proposed that the ear’s nonlinearity generates a 200 Hz component from the differences between adjacent partials, and that the listener then hears the real, internally generated tone. Elegant, mechanical, and testable.
It is wrong, and the test that showed it is the one the previous rung is built on.
Now shift every partial up by the same amount. Take partials at 1,840, 2,040 and 2,240 Hz: their spacing is still 200 Hz, so a distortion product at the difference frequency is still at 200 Hz exactly. The distortion account predicts the pitch does not move.
The measured shift matches the best-fitting fundamental rather than the spacing. The pitch is being computed from the pattern of partials, not read off a frequency the ear manufactured, and the distortion products — which really are there — are a separate phenomenon that happens to have the right value in the unshifted case and the wrong one as soon as the case is varied.
That is a clean example of a good hypothesis being killed by a variation rather than by a direct test. Helmholtz’s account and the pattern account agree on every unshifted stimulus, and they disagree by four hertz on a shifted one. Four hertz at 200 is about 34 cents, which is well above anything a listener could miss — so the experiment is not delicate, it simply had to be thought of.
The shift is not one number, and that is the second half of the test
Four hertz is the answer for the particular partials drawn above, and the more discriminating fact is that it is not the answer for any others. Keep the shift at forty hertz and the spacing at two hundred, and move the partials up and down the series:
| harmonics | partials | best-fit fundamental | shift |
|---|---|---|---|
| 3, 4, 5 | 640, 840, 1040 | 209.6 Hz | 81 cents |
| 4, 5, 6 | 840, 1040, 1240 | 207.8 | 66 |
| 6, 7, 8 | 1240, 1440, 1640 | 205.6 | 48 |
| 9, 10, 11 | 1840, 2040, 2240 | 204.0 | 34 |
| 16, 17, 18 | 3240, 3440, 3640 | 202.4 | 20 |
Every row has the same spacing, so the distortion account predicts two hundred hertz in all five, and the pattern account predicts five different pitches spanning sixty cents. The shift falls as one over the harmonic number — a least-squares fit is dominated by the highest partials, so the same absolute displacement is a smaller fractional one the further up the series it sits, and the classical approximation of the displacement divided by the mean harmonic number reproduces the whole column to within two per cent.
That turns the experiment from a coin toss into a curve. A single shifted stimulus gives one number and could be explained away — mistuned equipment, a listener’s bias, a poorly matched comparison tone. Five shifted stimuli give a slope, and the slope is one over n, which no account based on the spacing of the partials can produce at all, since the spacing is identical in every condition.
It also says which stimulus to use, and the answer is not the one usually drawn. The effect is largest at the bottom of the series — eighty-one cents at harmonics three to five, against thirty-four at nine to eleven — so an experimenter wanting the clearest separation should use low harmonics, and the reason the literature’s canonical stimulus sits higher is that low harmonics are resolved individually and a listener may report one of them instead of the residue.
Beats, difference tones, and the thing they are not
Three phenomena get confused with each other constantly and they are cleanly separable, which is worth doing once here because two of them already have essays.
Beating is a property of the sum of two waves and exists outside any listener. It is heard when the difference is under about 20 Hz.
A difference tone is a frequency generated inside the ear and does not exist in the air. It is heard as a pitch, so it requires the difference to be above about 20 Hz, and it requires the primaries to be loud.
And a residue pitch is neither: no frequency anywhere, in the air or in the ear, corresponds to it. It is the output of a pattern-matching computation, and the shifted-partials experiment is what proves it.
The three form a neat progression from most physical to least, and a listener presented with two tones and asked what they hear may be reporting any of them depending on how far apart the tones are and how loud they are. That is the practical reason the literature took two centuries to sort out: the same experimental setup produces all three, and which one is being described was frequently left to the reader.
Which computation produced the numbers
The combination tone frequencies are arithmetic and the figure computes them: f₂ − f₁, 2f₁ − f₂, 2f₂ − f₁, 3f₁ − 2f₂. Nothing about their positions is quoted.
The residue figures compute the shifted fundamental by least-squares fit over the partials present, which is where the 204 Hz comes from — and that computation has its own recorded history on this site. The fit needs eleven partials while the named timbres are eight long, so an earlier version sliced an empty set of present partials and produced a fitted fundamental of NaN, printed into a caption. The series now continues as one over n past the end of the table.
The levels of the combination tones are not computed and the figure does not draw them to scale. Predicting how loud a distortion product is requires a model of the cochlear amplifier, which is a substantial piece of physiology and not arithmetic over two frequencies. What the figure asserts is where the products fall, which is exactly what the essay’s argument needs.
What an active cochlea explains that a passive one does not
The nonlinearity is a side effect of something the ear needs, and the list of what it buys is long enough to make the trade obviously worth it.
Sensitivity. A passive cochlea would have a threshold tens of decibels higher than the real one. The amplifier supplies most of the gain at the quietest levels, which is why the threshold of hearing sits where it does and why it rises so sharply when outer hair cells are damaged.
Sharp tuning. The mechanical tuning of a dead cochlea is broad. A living one is sharp, and the sharpening is active — which is why the critical band is as narrow as it is, and why bandwidth broadens with level and with hearing loss.
And compression. The amplifier’s gain falls as level rises, compressing a 120-decibel range of input into a much smaller range of mechanical response. That is why masking spreads further upward at high levels: the compression flattens the tuning of each place as the input grows.
Every one of those is a consequence of a mechanism that is nonlinear by construction, and a nonlinear mechanism generates combination tones whether anybody wants it to or not. Tartini’s third sound is the audible by-product of the machinery that makes quiet sounds audible at all — which is a good example of a general pattern in perception, where the illusions and the artefacts are usually the exhaust of something useful.
What the picture cannot show
Level, which decides everything about audibility. Whether a listener hears Tartini’s tone at all depends on how loud the primaries are, and the two products behave completely differently as level changes. A figure drawn at one level tells nothing about the others.
The frequency ratio’s effect. The cubic tone’s dependence on f₂/f₁ — rising to a peak near 1.22 and falling away — is one of the most characteristic features of the phenomenon and none of it is in a single-ratio drawing.
The five-row table is a prediction and not a measurement. The pattern account’s column is a least-squares fit computed here; the distortion account’s is two hundred by construction; and no listener was asked. What the published work reports is the first row of the argument rather than the slope — the shift exists and goes the right way — and the one over n dependence is stated in the literature as an approximation rather than measured across a range as wide as the one drawn. The table’s value is that it says what a discriminating experiment would look like, which is five conditions rather than one.
And the emission itself. The strongest evidence in this essay is a measurement made with a microphone in an ear canal, and no drawing of a spectrum shows it. What can be said is what the figure does say: the frequencies are in the wrong places to be anything but products of the pair.
The listener’s own ear is also the apparatus, which is the one respect in which the sound buttons on these figures are doing something the figures cannot. A reader who plays the two primaries loudly on a good pair of headphones may hear the third sound; a reader who plays them quietly, or on a small loudspeaker whose own distortion is larger than the ear’s, will hear something else entirely. This is the only place on this site where a sound button’s result depends on the reproduction chain in a way that could actively mislead, and it is why the buttons here play the products separately for comparison rather than asking the reader to take an unverifiable experience as evidence.
Whose practice, and how far the violin lesson goes
Tartini taught the third sound as an intonation aid: play a double stop, listen for the low tone, and adjust until it is where it should be. The technique is still taught and it works — for loud, sustained, just-tuned double stops in the middle of the instrument’s range, which is a narrow and specific condition.
Outside it the technique fails, and the failures follow from the figure. At quiet dynamics the quadratic tone is not generated at any useful level. On tempered rather than just intervals the product is a few cents off and much harder to identify. And on an instrument or in a register where the primaries are far apart, the difference tone lands so low it is below the audible range.
So a practice that looks like folklore turns out to be an exact piece of applied psychoacoustics with a stated domain of validity, and violinists have been operating inside the domain without a way of writing it down. The essay’s honest position is that the technique is real, the mechanism is understood, and the claim that a player is tuning by difference tones in ordinary playing is much stronger than the evidence supports.
Where the ladder goes next
The missing-fundamental anchor now has two rungs, and the second was written to rule out an explanation of the first. What is left open is the mechanism that does the work: pattern matching over resolved partials is one account, autocorrelation of the timing of nerve firings is another, and the two make different predictions for spectra whose partials are too close together to resolve.
The competing accounts also differ in a way the difference limen can arbitrate, which is worth saying because it is unusual for two theories of pitch to make predictions a few cents apart and unusually convenient that they do — the shifted-residue experiment works because the disagreement is thirty-four cents rather than three.
Sideways, this rung supplies the machinery for a different argument entirely. An ear that generates frequencies of its own is an ear whose output spectrum is not its input spectrum, which means every masking measurement has an extra term in it — a listener detecting a probe beside a loud masker may be detecting a distortion product instead, and the experiments have to be designed around it.
Part 2 of 9
One essay in the series on missing fundamental. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 11.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Cochlear nonlinearityCombination toneDifference toneDistortionOtoacoustic emissionResidue pitch
- Three harmonics of the bass arrive before the bass combination tone, difference tone, residue pitch
- A combination-tone bass needs a forte combination tone, difference tone
- The bass line under a passage in thirds combination tone, difference tone
- The played notes already name the ghost bass combination tone, residue pitch