What makes two partials one note
Assumes: The ear builds objects, and sometimes offers a choice
A single note on a cello is between ten and thirty simultaneous pure tones. They are audibly present — a trained listener can pick out the first few by ear — and they are nevertheless heard as one object with one pitch. Something is deciding that they belong together, and finding out what requires making it fail.
Three per cent of 660 Hz is about twenty hertz — half a semitone, at that frequency. The note’s pitch does not move. Its loudness does not change. Nine of the ten partials are untouched. And a listener who could not hear the third partial at all a moment ago can now hear it plainly, because it is no longer part of anything.
The obvious explanation, and why it is not enough
The obvious explanation is harmonicity: the ear looks for a set of frequencies that are whole-number multiples of a common fundamental, finds them, and reports one note. Anything not fitting the pattern is left over.
This account is not wrong. It is doing real work — the mistuning experiment above is direct evidence for it, since what changed was precisely the arithmetic — and a pattern-matching model of pitch is what the missing fundamental forced on the subject in the first place. But it is not sufficient, and the demonstration that it is not sufficient is the whole of this essay’s argument.
A partial that is exactly where the harmonic template says it should be, at exactly the right amplitude, is nevertheless not part of the note if it did not start with the note. That is decisive: the grouping decision is not made on the spectrum alone. It is made on the spectrum as it develops.
Where the thresholds are
Three numbers do most of the work, and all three come from a body of experiments in the 1980s, principally Brian Moore, Robert Peters and Brian Glasberg.
About one per cent of mistuning is enough to hear a partial out. For low-numbered partials, a departure of one per cent from the harmonic position makes the partial separately audible. That is roughly seventeen cents at the third partial — smaller than a syntonic comma, and comfortably above the difference limen, which is the only reason it can be a threshold at all.
About three per cent and it stops contributing to pitch. The mistuned partial not only leaves the note; it stops voting on what the note’s pitch is. Below three per cent, mistuning a partial shifts the perceived pitch of the whole complex slightly — the partial is still a member and its opinion is counted, at a reduced weight. Above it, the partial is excluded and the pitch is decided by the rest.
About thirty milliseconds of asynchrony does the same job. An onset difference of thirty milliseconds is enough to segregate a partial. Below about ten it is ineffective, and there is nothing sharp about the transition.
Two of those numbers are about the note’s arithmetic and one is not, which is the point. And the third is much smaller than a musical event: thirty milliseconds is a tenth of a quaver at a moderate tempo, and it is nothing anybody would notate.
It is worth comparing the two mistuning numbers, because the gap between them is where the interesting behaviour is. One per cent hears a partial out and three per cent stops it voting on the pitch, so there is a factor of three in which the partial is audible as separate and still counted. That is not a transitional muddle; it is a state with two properties that a simple sorting model cannot produce, and it is what the pitch-pulling below is evidence of.
The pitch-pulling below threshold is worth dwelling on, because it says something about how the decision is made. A partial mistuned by half a per cent shifts the pitch of the whole complex by a fraction of the amount it was itself moved — as though the ear were taking a weighted average over the partials and this one still had a vote. Above the threshold the vote is withdrawn entirely. The system is not sorting components into two piles at the start and then processing each; it is estimating a fundamental and deciding membership at the same time, with each answer conditioning the other.
That is the harmonicity cue applied as a census rather than to one manipulated partial. The threshold is about one per cent, and what the census adds is that a spectrum is not fused or unfused as a whole — the count of partials that make it into the series is the quantity, and it varies from spectrum to spectrum rather than switching.
Why the ear works this way
The rule the ear is applying is not “find the harmonic series”. It is closer to things that came from one physical event behave alike, and harmonicity is one of several consequences of that, not the principle itself.
A vibrating object driven periodically produces harmonics. That is a fact about physics, and betting on it is sensible. But the same physical event also produces components that start together, stop together, swell and fade together, and come from the same direction, and each of those is an independent piece of evidence for the same conclusion. A system that used only one of them would be discarding four.
There is also a reason to weight onset heavily, and it is worth putting a number on because the number does not support it in the case this essay cares about. Harmonicity is a coincidence a mixture can produce accidentally: two instruments an octave apart share most of their partials, and read as a spectrum at one instant, their combination is a perfectly good harmonic series on the lower fundamental. Simultaneous onset is supposed to be much harder to fake.
For independent sources it is. For music it is not, because music is the one situation in which independent sources deliberately start together. Taking each part’s onsets as an independent random process and asking how often one falls within thirty milliseconds of another:
| notes a second, per part | two parts | four | twelve |
|---|---|---|---|
| 1 | 6% | 16% | 48% |
| 2 | 11% | 30% | 73% |
| 4 | 21% | 51% | 93% |
At an ordinary note rate in a small ensemble, half the onsets have a companion inside the fusion window by chance alone — and that is for parts assumed independent. Parts playing to a shared metre coincide by construction, so the real figure is not fifty per cent but a hundred.
So onset is an excellent cue for telling a bird from a car and a poor one for telling one instrument from another in an ensemble, and the essay’s own orchestration section is the evidence. Rehearsing the attacks works precisely because the cue can be faked: a section aligning its onsets to a few milliseconds is manufacturing exactly the evidence a listener uses for a common source, and it succeeds. If onset were the reliable signature the paragraph above claims, no amount of rehearsal could make twelve players sound like one instrument.
The window at which independent parts at four notes a second would coincide less than one time in twenty is six milliseconds, which is a fifth of the fusion threshold and about the precision a good section actually achieves. That is not a coincidence worth leaning on, but it is the right order to notice: an ensemble is aiming at the tolerance at which the cue would have been informative.
What this predicts about orchestration, and whether it holds
The prediction is that the blend of an ensemble is controlled by attack alignment more than by anything to do with intonation, and it is a strong enough claim to be checkable against practice.
Ensembles work at the attacks. Rehearsal time in choirs, wind sections and string quartets goes overwhelmingly into starting together and releasing together, and almost none of it into tuning steady tones against each other. That allocation looks like superstition from a purely acoustic point of view and is exactly right from this one: an ensemble whose attacks are aligned within a few milliseconds will fuse into one instrument even if its intonation is mediocre, and an ensemble with immaculate intonation and scattered attacks will not.
The organ’s mixture stops. An organ mixture is a rank of pipes tuned to upper harmonics of the note being played — the twelfth, the fifteenth, the seventeenth — sounding together with the fundamental. They are separate pipes, and they fuse into one composite timbre rather than being heard as a chord, because they speak together and share the same wind and the same articulation. Whether a listener hears a mixture as a chord or as a bright tone colour is entirely a fusion question, and the answer for a well-regulated instrument is “as a colour”.
And it predicts where blend fails. A note that is bowed and a note that is plucked will not blend however well tuned, because their onsets are a hundred milliseconds apart in shape even when they are simultaneous in time. That case is the one where the cue is genuinely hard to fake: a player can align the moment two notes begin to a millisecond and cannot align the shape of two attacks at all, because the shape is set by the mechanism. So the half of the onset cue that an ensemble can manufacture is the timing and the half it cannot is the envelope, which is a sharper statement of what an orchestration rule about combining instruments is a rule about. Orchestration textbooks record this as a rule about which instruments combine, and the underlying quantity is the attack rather than the steady spectrum.
Whose practice. All three are claims about Western ensemble music from roughly the eighteenth century onward and about the instruments it uses. The strongest of them — that blend is bought at the attack — is likely to generalise, because it follows from the mechanism rather than from a convention. The specific ones do not: a gamelan deliberately detunes paired instruments so that they beat, which is the opposite aesthetic decision made with the same machinery, and it works because the pairs still start together.
The cues disagree, and the figure is about what happens then. A partial can be harmonic and decaying wrongly, or inharmonic and decaying with the others, and the answer is not given by either cue alone — which is why “what makes two partials one note” has no single-cue answer and needs an arbitration rule.
Which computation produced the numbers
The figures here compute two things and take two others from the literature, and the boundary between them should be visible.
Computed: the position of every partial, and the beat rate the mistuned one produces. A partial at three per cent above its harmonic position beats against where it should have been at a rate equal to the mistuning in hertz — 19.8 Hz for the third partial of a 220 Hz tone — which is fast enough to be heard as roughness rather than as a slow throb, and that is why the mistuned tone sounds rough as well as separate. Both the figure and its sound button are built from the same list of ratios, so the drawing and the noise cannot disagree about where the partial is.
Taken from the literature: the one-per-cent and three-per-cent thresholds and the thirty-millisecond asynchrony. All three are averages over listeners in a laboratory task, all three depend on which partial is being asked about — high-numbered partials are much harder to hear out, because they are crowded together inside one critical band — and all three are drawn as vertical lines when they are really soft transitions.
What the picture cannot show
Fusion is not binary and the figure draws it as if it were. A partial mistuned by half a per cent is partly out: audible if attended to, invisible otherwise, and still contributing to the note’s pitch. Most real cases are in this state.
Partial number matters enormously and the figure fixes it. The first five or six partials are resolved by the ear into separate critical bands and can be individually heard out. Above about the eighth, several partials share a band, none of them can be heard out at all, and mistuning one produces roughness rather than segregation. The thresholds quoted here are for low partials and are not thresholds for high ones.
Nothing here is about a mixture. The experiment mistunes one partial of one note. Real polyphony presents several complete harmonic series at once, overlapping, and the assignment problem is far harder than deciding whether one component is a member. Whether the ear solves the polyphonic case with the same machinery is an open question, and the honest answer is that the models which work on isolated notes do not work on orchestras.
The generalisation, and the sound it explains
The cue that decides fusion is common fate, and once that is the principle, several unrelated-looking facts turn out to be one fact.
Vibrato binds a voice together. Vibrato applies the same frequency modulation to every partial of a note at once. Nothing else in a texture is wobbling in that particular way at that particular rate, so the modulation acts as a tag: these components belong together, and they belong to this source rather than to the one beside it. This is why a singer with vibrato is easy to follow through an orchestra and why a synthesised tone with no modulation at all sounds artificially inert and blends into everything.
A chorus effect works by breaking it. Slightly detuning and delaying several copies of a sound gives each copy a different fate, and the result is heard as several sources rather than as one louder source. It is the mistuned-partial experiment done deliberately and at scale.
And a synthesiser’s oldest problem is exactly this. Additive synthesis from a perfect harmonic series with a common envelope produces a sound that fuses too well: it has no internal life, because every partial is doing precisely the same thing. Every convincing synthetic instrument gives its partials slightly different envelopes and slightly different micro-detunings, which is to say it deliberately weakens fusion in order to sound like an object rather than like a formula.
What this has to do with tuning
There is a connection to the rest of this site that is easy to miss and worth making explicit.
Every essay on the tuning ladder treats a note as a thing with a frequency, and asks how two of those frequencies should be related. This essay says a note is not a thing with a frequency; it is a coalition of twenty things, held together by a set of coincidences, and the coalition can be broken.
That is exactly why a tuning system has consequences for timbre and not only for pitch. Two notes an interval apart share partials to the extent that their frequency ratio is a ratio of small integers — a just fifth shares every third partial of the lower with every second of the upper — and shared partials are components with a genuinely ambiguous membership. Tune the fifth a little narrow and the shared partials separate into pairs a few hertz apart, each pair beating, and the two notes become audibly two objects. Tune it pure and they fuse.
A pure interval is therefore not merely a smoother interval; it is one whose notes are harder to hear as two. That is the fact the ban on parallel fifths is really about, and it is the same mechanism as this essay’s, working at the level of a chord instead of a note.
Where the ladder goes next
This rung and the one before it are the same question asked on two axes: which simultaneous components belong together, and which successive events belong together. The cues overlap heavily and the answers interact — a partial that has been segregated from its note is then available to be streamed with something else, which is how the compositional devices in the previous essay get their material.
From here the ladder can go two ways. Towards the room: components that arrive from the same direction fuse more readily, which is the localisation ladder supplying a cue to this one. And towards the instrument: what a listener does with a spectrum that will not fuse — a bell, a gong, a detuned pair — is a question about inharmonicity, which already has an essay on the physics and not yet one on what it sounds like.
Part 2 of 8
One essay in the series on auditory scene. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 31.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Common fateFusionHarmonicityHearing outMistuned partialSimultaneityTimbre
- Eleven partials is one too many fusion, harmonicity
- The body is the filter harmonicity, timbre