The ghost bass drops a twelfth at a forte
Assumes: Three harmonics of the bass arrive before the bass · A combination-tone bass needs a forte
The essay that found the crowd naming the bass first found two crowds and priced them separately. The difference tones of a just interval’s partials are an exact harmonic series on the fundamental the interval implies, and they need a forte. The cubic products of the same partials are an exact harmonic series on times that fundamental — a twelfth above it for a major third — and they are there at a piano.
Nothing in the ear separates them, and nothing in the air does either — both are made inside the listener. Both crowds arrive together, and since every member of both is a whole multiple of the same , what a listener receives at any dynamic is a subset of one harmonic series. Which subset is what the dynamic decides, and different subsets support different pitches.
The subset is the whole question
Consider a just major third, , whose implied fundamental is C2 at 65.4 hertz. At fifty-five decibels the audible members are the sixth, ninth, twelfth, fifteenth and eighteenth multiples, and nothing else.
Those are not a haphazard collection. Divided by three they are 2, 3, 4, 5 and 6 — a perfect run of harmonics of , with no gaps at all. A pitch template fitted to them names 196 hertz, three times the implied bass, and has nothing to explain away.
At eighty decibels the difference tones arrive as well, and they bring the first, second, third, fourth and fifth multiples. Divided by three those are not whole numbers, so the reading on is destroyed: it would have to explain five components that sit a third of a harmonic apart. The only template that survives names itself — with seven empty slots in it, which is a worse fit by the usual measure, and it is the only fit available.
So the pitch drops by a twelfth, and it drops because something arrived rather than because something left. Every component that was audible at fifty-five decibels is still audible at eighty. The soft reading is destroyed by additions to the evidence, not by subtractions from it.
The evidence improves and the fit gets worse
The mechanism above contains a small paradox and it is worth separating out, because it is the reason the switch is sharp rather than gradual.
A residue template is judged by how many harmonic slots it predicts that nothing occupies. Which harmonics carry the pitch is the essay that established why: a candidate fundamental an octave below another explains the same components equally well in arithmetic and has to invent twice as many absent ones, so the slot count is the only thing that separates them.
By that measure the soft reading is the best fit the essays here have ever computed on anything. Five components, assigned harmonic numbers two through six of 196 hertz, zero empty slots, zero cents of error. It is a textbook harmonic series with its fundamental missing — the object the essays about a pitch with no fundamental under it — and there is no competing candidate within thirty cents.
The loud reading is much worse. Fifteen components on 65.4 hertz, exactly fitted, with seven empty slots: the sixth-through-eighth region is patchy and there is nothing at all between the twelfth and the fifteenth or between the eighteenth and the twenty-second. If the two readings were ever competing on the same evidence the soft one would win comfortably.
They never compete on the same evidence, and that is the point. The loud crowd contains components — the first, second, fourth and fifth multiples — that the reading on cannot place anywhere at all: a third of a harmonic is not a harmonic. So the soft template does not lose a comparison; it is eliminated, and what wins is whatever is left. A worse fit that can explain everything beats a perfect fit that cannot explain four things.
That is what makes the transition a step rather than a slope. The quantity that changes is not a score but a feasibility, and feasibility changes at the decibel where the fourth multiple clears its threshold.
Where it happens
A major third’s ghost sits on up to seventy decibels and on above it, a drop of nineteen semitones. A minor third’s sits on up to sixty-two and on above, a drop of twelve. A fourth’s sits on up to seventy-two and on above, a drop of twelve.
A fifth does not drop at all, and that is the control the previous essay supplied: is one for a fifth, so both its crowds aim at the same note and there is nothing for a dynamic to change. That mirrors what the essay that found the two products changing places at the fifth found for pure tones, where the fifth was also the interval at which the two products changed places. The intervals whose ghost moves are exactly the intervals whose two crowds disagree, which is every common interval except the fifth and the octave.
The level at which it happens falls with register, by about four decibels an octave: a third whose implied bass is C1 switches at seventy-four decibels, one on C2 at seventy, on C3 at sixty-six and on C4 at sixty-two. The mechanism is the threshold of hearing, which is expensive at the low multiples and cheap at the high ones, and which gets cheaper for everything as the whole figure moves up.
The band where the passage is in two places
A diatonic scale harmonised in thirds alternates major and minor thirds, and those two switch at different levels — seventy decibels and sixty-two. Between them lies a band in which the two kinds of step have done different things.
Inside that eight-decibel band the inner line is not a line at all. Some of its notes are a twelfth above their own implied bass and the others are an octave above theirs, and the resulting sequence is neither the loud line nor the soft one. A crescendo through a passage in thirds therefore does not move the ghost line; it rewrites it, one step at a time, in an order set by which thirds are major.
That is a strange enough prediction that it is worth being clear what would make it wrong, and there are two candidates. If a listener tracks a ghost pitch across steps the way a listener tracks a melody — and there is every reason to think pitch tracking is not independent step by step — then a line half of which has jumped may simply not be heard as a line, and the percept inside the band would be an absence rather than a scramble. And if the switch is gradual rather than sharp, because the template competition is graded rather than winner-takes-all, then the band is a region of weak or bistable pitch instead of a region of wrong ones.
Both of those are more likely than not. What the arithmetic supports is the narrower claim: the evidence a listener is given changes category at a stated dynamic, and it changes at a different dynamic for the two qualities of third. The size of the jump is fixed by the interval and not by the ear: nineteen semitones is exactly the ratio between and , which is the same arithmetic the essay that measured the third sound’s gearing runs on.
What the effect is not
Three nearby phenomena share the shape of this one and none of them is it, and separating them is most of what makes the claim a claim.
It is not the difference tone getting louder. A component that rises through threshold as a passage grows is the ordinary case and it is what the essay that priced the two level laws: the bass appears where it was absent. What happens here is that a pitch which was already present, stable and unambiguous is replaced by a different one. A listener attending to the ghost does not hear something arrive under it; they hear the thing they were attending to move.
It is not the pitch shift of the residue. A residue pitch does shift when its components are moved bodily, by a well-known and quite small amount — the pitch that moves the wrong distance is the essay that measured it, and the shifts there are a few per cent. Nothing here moves any component at all. Every frequency in the crowd is exactly where it was; only the membership of the set changes, and the jump is an exact octave or twelfth rather than a few per cent.
And it is not a matter of one mechanism handing over to another. Both crowds are made by the same nonlinearity in the same place, and both are harmonics of the same fundamental. There is one set of evidence at every dynamic and one template fitted to it. The switch is internal to a single account rather than a boundary between two.
What it most resembles is the ambiguity the residue has always had. Two tones a fifth apart are harmonics two and three of one note, four and six of the note an octave below, six and nine of the one a twelfth below, and nothing in the pair chooses between them. The crowd here is that ambiguity with a dial on it: the dynamic decides how many of the low harmonics are supplied, and supplying them is what forces the choice.
What a player would have to do to hear it
The prediction is testable on an ordinary instrument and the protocol is short.
Two players hold a just major third — the upper player tuning by ear until the beating stops, which puts the interval within a couple of cents of . They begin at a dynamic they can both hold steadily and quietly, and grow it slowly to a full forte over twenty or thirty seconds, without vibrato, on an instrument with a strong upper spectrum. The register wanted is low, because the switch is sharpest where the difference tones are most expensive: a major third in the bottom octave of a cello’s compass puts the implied fundamental near C1, where the predicted switch is at seventy-four decibels and the drop is nineteen semitones.
What the arithmetic predicts they will hear is a faint third pitch that is stable through the first part of the crescendo, and that at some point during it falls by an octave and a fifth — not fades, falls — while the two played notes go on getting louder.
The most likely negative result is that nothing audible happens at all, and that would be informative rather than disappointing. Every level in this essay comes from two published laws applied thirty-six times over inside one ear, with constants swept rather than measured, and the strongest claim the model can make is about where a boundary is rather than about whether anything is above threshold at all.
Which computation produced the numbers
Both crowds are computed as the previous essay computes them: each note given six partials at one over , every pair of partials treated as a pair of primaries at the level of the weaker, the quadratic difference tone at and the cubic product at a nearly fixed distance below the primaries widening steeply with the ratio. A component is audible when it clears both the threshold of hearing at its own frequency and the masked threshold each primary casts there, and where several pairs make one frequency the loudest is taken.
The two sets are then merged and the template is fitted to the lowest eight members. That cap is a choice and it is doing real work, so it is worth stating why it is there rather than at twenty: a pitch mechanism does not have twenty resolved components to work with, and the enumeration over every assignment of harmonic numbers to twenty frequencies is combinatorially large as well as physiologically wrong. Raising the cap does not move the switch, because the switch is decided by the arrival of the low multiples and those are always inside the lowest eight.
The switch level is found by stepping the playing level in two-decibel increments and taking the first level at which the fitted multiple changes, so every figure above quotes a level to two decibels and no better.
Where the model stops
The template is winner-takes-all and hearing is not. The residue model ranks candidates and this essay reads the top one. A real pitch mechanism produces a salience over candidates, and near the switch two candidates a twelfth apart are close in that ranking — which is precisely the ambiguity the octave and the twelfth have in every residue experiment ever run.
And the crowds are computed as though the ear were the only nonlinearity. A loudspeaker, a reed, a horn’s air column at high amplitude and a microphone all distort, and their products are in the air where the ear’s are not. Anything that adds real acoustic difference tones adds them to the low multiples first, which would move the switch downward — the hazard a pitch with nothing to match describes from the other direction — so a loud recording played through an imperfect chain may show the effect at a lower level than a live performance.
The passage has no tempo. Each step is computed as a held interval in equilibrium, and the previous essay recorded that the clock is what this whole account is missing.
What the picture cannot show
It cannot show what the two played notes’ own partials do. They arrive alongside the crowd, and in the just case they are harmonics of the same — the fourth, fifth, eighth, tenth, twelfth and fifteenth for a major third. Adding them to the evidence adds low multiples at every dynamic, which would push the switch downward and might abolish it. Whether a listener’s pitch mechanism pools the products with the played notes or segregates them is the question a spectrum that will not fuse is about, and it decides whether this essay describes anything.
Nor a tempered interval. the previous essay found that tempering splits the crowd into a coherent series seventy cents flat and a scatter around it. A switch between two templates when one of them is already seventy cents wrong is a different object, and probably a much weaker one.
And it cannot show the ghost’s loudness. Everything here is about which note is named. A pitch that is named by five components a few decibels above threshold and a pitch named by fifteen components well above it are not equally present, and nothing in the template fit carries that difference.
Still open: whether the played notes abolish the switch
The single most consequential simplification above is that the crowd is fitted on its own, with the two played notes’ partials set aside. They are not set aside in an ear.
For a just interval, partial of the lower note sits at and partial of the upper at — both multiples of the same fundamental as every product. So the full evidence a listener receives is the union of three sets of multiples of one number, and the played notes contribute the low-numbered members that the soft crowd lacks. A major third’s own partials supply multiples 4, 5, 8, 10, 12, 15, 16 and 20, and 4 and 5 are exactly the kind of component that destroys a reading on .
Running the fit on the union of all three would say whether the soft reading exists at all, and it would settle the question this essay is built on rather than refining it. The reason it is not done here is that it requires a decision the essays here have not made and cannot make from arithmetic: whether a pitch mechanism treats the two loud played notes and the faint products as one object to be explained or as a figure and a ground. Both answers are defensible, they give opposite results, and the difference between them is the whole of what a listener would report.
Part 7 of 8
One essay in the series on combination tone. The essays either side of this one:
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Combination toneDifference toneDynamicsHarmonic seriesHearing thresholdJust intonationResidue pitch
- The bass line under a passage in thirds combination tone, difference tone, just intonation
- A bell has no fundamental harmonic series, residue pitch
- A fifth on a piano is not a fifth a second later harmonic series, hearing threshold
- The register where the series becomes a scale harmonic series, just intonation
- The series is not a chord harmonic series, just intonation
- The top that falls while the note lasts harmonic series, hearing threshold