The partials that do not die together
Assumes: The spectrum that will not fuse · What makes two partials one note
The spectrum that will not fuse took five spectra this collection already carried — an ideal string, a piano string, a bell, a bar, a kettledrum — and asked how many of each one’s partials sit near enough to a whole multiple of a single fundamental to be heard as one object. That is harmonicity, which every textbook calls the strongest fusion cue, and a list of ratios is all it needs.
Its last paragraph named the cue after it and why the census could not have it. Components that move together belong together, and every spectrum in that figure is a still photograph. A bell’s partials decay at rates differing by an order of magnitude, so they are emphatically not moving together; a vibrato applied to a harmonic tone moves every partial in exact proportion. Running a second census over the same five spectra with time in it would very likely disagree with the first.
It does. It nearly reverses it.
What the second census measures
Common fate is Gestalt vocabulary and it means one thing: components undergoing the same change at the same time are grouped. In hearing it has two limbs, and they are physically different quantities.
Onset and offset synchrony. Components that start together and stop together are one object. This is the limb the five spectra can be asked about, because every one of them is a struck or plucked sound with a decay, and every partial of a struck sound has its own decay rate for a physical reason.
Coherent modulation. Components whose frequencies move together are one object. This is the limb that needs a source that modulates at all, which none of the five does, and which is the second half of this essay.
The first limb needs a decay time per partial, and the reason there is one is different for each spectrum. A string’s higher modes shed energy faster, so its partial decay times fall steeply with mode number. A bell’s hum note rings for half a minute while the partials above it are gone in seconds. A kettledrum is damped almost entirely by radiation into the air, which loads every mode nearly equally — which is why a timpano’s note is short and why every part of it is short.
The inversion
Ranked by harmonicity: ideal string, piano string, bell, kettledrum, bar. Ranked by common fate: kettledrum, bar, ideal string, piano string, bell.
The best spectrum on the first cue is third on the second. The worst is second. The bell is third on the first and last on the second. Nothing survives the change of cue except the piano string’s place in the middle, and that is a coincidence of two different middles.
The extreme case is the one worth stating on its own. An ideal string is a perfect harmonic series and holds only a fifth of its partials together by decay. Its harmonicity is exactly one — every partial is exactly a whole multiple, by construction, because that is what an ideal string is — and its tenth partial dies ten times faster than its first. On the second cue an ideal string is a set of components going their separate ways.
That is not a defect in either census. It is the reason a plucked string sounds the way it does. The note that gets duller as it dies is the essay about exactly that slope, and its finding is that the rate at which the colour drains is an identity cue in its own right — one number that sorts instruments the attack cannot. So the same fact is a fusion failure under one grouping principle and an identity signature under another, which suggests the two principles are not competing to describe one thing.
The case that separates them cleanly
The five spectra disagree about which cue is better satisfied. A stronger test is a stimulus where one cue says one object and the other says two, unambiguously, and it is easy to build.
Take a tone at 220 hertz and a tone an octave above it. Every partial of the upper tone is also a partial of the lower — the union of the two sets is exactly the harmonic series of 220 — so harmonicity has nothing to work with at all. The census fuses sixteen of sixteen and reports one object. That is not an artefact: it is why an octave is the interval that fuses, which what makes two partials one note established from the other side.
Now put a fifty-cent vibrato on the lower tone. Eight partials move and eight do not, and the eight that move do so in exact proportion — fifty cents at the first partial and fifty cents at the eighth, which is six hertz and fifty-one hertz. Common fate separates them on the first cycle.
Why the excursion is in cents and the cue is in hertz
There is an arithmetical point hiding in that figure and it is what makes coherent modulation work at all.
A vibrato is a change of fundamental, so every partial moves by the same fraction. On a logarithmic axis they move rigidly — the whole comb slides up and down without changing shape. In hertz they do not: the eighth partial’s excursion is eight times the first’s, fifty-one hertz against six.
That matters because the auditory system’s frequency analysis is not linear either. A partial’s movement has to be large enough relative to the width of the filter it sits in to be detected as movement, and critical bandwidths widen with frequency but less than proportionally. So the higher partials of a vibrato tone are the ones whose movement is most detectable in the units the ear works in, which is the opposite of the usual assumption that the low partials carry everything — which harmonics carry the pitch puts the dominance region for pitch between the third and fifth partials, and the dominance region for this cue is somewhere else entirely.
Computing it is division. Take each partial’s excursion in hertz and divide it by the width of the band it sits in, for a fifty-cent vibrato:
| fundamental | partial where the ratio is largest | the frequency that is |
|---|---|---|
| 110 Hz | the sixteenth | 1,760 Hz |
| 220 Hz | the eighth | 1,760 Hz |
| 440 Hz | the fourth | 1,760 Hz |
The dominance region for coherent modulation is a fixed frequency, not a fixed partial number, and on the Bark bandwidth it sits at about 1.8 kilohertz whatever the fundamental. That is the structural difference from the pitch-dominance region, which is the third to the fifth partial and therefore moves with the note: a singer going up an octave keeps the same partials carrying their pitch and hands the modulation cue to partials half as far up the series.
The two bandwidth models disagree about whether there is a maximum at all, and the disagreement is worth reporting rather than resolving. On the Bark formula the ratio rises to 0.196 at 1,760 hertz and falls away above it. On the equivalent rectangular bandwidth it rises monotonically to a ceiling of 0.271 and is within one per cent of that ceiling from about two kilohertz upward — a plateau rather than a peak. Both put the region well above the pitch one and neither puts it anywhere near the low partials, which is all the argument needs; where exactly it sits depends on which of two fits to listening data is used, and the two are known to disagree by a factor of two low down.
One number from that table is worth carrying on its own. At its most detectable, a fifty-cent vibrato moves a partial by a fifth to a quarter of a critical bandwidth — not across a filter, but a fraction of the way along one. Whatever mechanism reads coherent modulation is reading a change of level inside a channel rather than a partial crossing between channels, and fifty cents is a large vibrato. That is a constraint on any account of how the cue works, and it is the kind of thing a census over five spectra could not have produced.
The census the third rung ran, for the comparison
The point of a second census is that it is the same five spectra and the same shape of question, so the two numbers can be put side by side without an argument about what is being compared. It is worth having the first one in view.
Setting the two beside each other also makes a smaller point that is easy to miss. The harmonicity census has a free parameter that is searched — the fundamental — and the common-fate census does not. There is no “best decay rate” to be found; the rates are what they are, and the count is a count. So the second cue is in one sense a stronger measurement and in another a weaker one: it cannot be gamed by a search, and it also cannot discover anything the way the search discovers a bell’s strike note.
What a mistuned partial does to each cue
The ladder’s own first rung is the smallest version of this question: one partial of an otherwise harmonic tone moved off its whole multiple, and the mistuning at which it is heard out of the complex.
That stimulus is a clean test of harmonicity alone, and it is why the number it produces — a partial is heard out at a mistuning of a few per cent — is quoted everywhere. Under common fate it is not a test at all: the mistuned partial starts and stops with all the others, so the cue says it belongs and goes on saying so however far the mistuning is pushed.
That is the practical form of the disagreement, and it is the reason the two cues are usually studied with two different stimuli. It is also why the census in the hero figure is worth having: five real spectra, one axis, and no experiment designed to isolate anything.
Which computation produced the numbers
The harmonicity census is fusionCensus, unchanged from the third rung: a search over candidate fundamentals for the one that puts the most partials within the mistuning at which a partial is heard out of its complex, with the search capped so that a fundamental twenty-five harmonics below the top partial cannot be nominated.
The common-fate census gives each spectrum a decay law of the form T60 over n to a power, evaluates it for ten partials, and counts how many fall within a factor of two of the strongest partial’s rate. The tolerance is stated rather than fitted and the result is not sensitive to it: the two rankings are nearly reversed at a factor of two and still nearly reversed at four.
The exponents are asserted and ordinal. What is published, and reproducible across instruments and laboratories, is that a bell’s partial decay times span more than an order of magnitude and a timpano’s barely span a factor of two. The five numbers here reproduce that ordering by construction. They are not a measurement of any instrument, and this collection has now recorded the same caution enough times that it should be read as a standing property of these figures rather than as a disclaimer.
The modulation coherence is the correlation between each partial’s frequency track and the modulating waveform, over whole vibrato cycles. Steady partials have zero variance and are reported as zero rather than as undefined.
Where the model stops
Decay is a single exponential per partial and it is not. Real struck strings and bells show two-stage decays — a fast initial rate and a slower tail — because energy moves between polarisations and between coupled strings. Three strings and the note that comes back is the essay about the mechanism on a piano, and it means a single rate per partial is a summary of something with structure in it.
Onsets are not in the census at all. Onset synchrony is the more powerful half of the first limb — components that start together fuse very strongly — and every spectrum here is treated as starting at one instant. A real strike excites the modes at slightly different times and with very different rise times, which is a whole cue this figure omits.
And a census is not a listener. Counting how many partials satisfy a cue is not a model of grouping; it is a way of making two cues commensurable. Real auditory scene analysis weighs cues against each other, and what happens when harmonicity says one object and common fate says two is an empirical question this arithmetic cannot settle.
What the picture cannot show
It cannot show which cue wins. The octave figure sets up the conflict and stops. What a listener hears when a vibrato is applied to one of two octave-related tones is a published result — the tones do separate — but the degree of separation, and how it trades against the vibrato’s extent, is a listening experiment and not a calculation.
Nor can it show a real vibrato. A singer’s vibrato is not a pure sinusoidal frequency modulation: it carries an amplitude modulation with it, because the vocal tract’s resonances do not move, so partials sweeping past a formant get louder and quieter as they go. That is a second common-fate cue riding on the first, and a note that is never at its pitch is the ladder that computes it.
It cannot show onset asynchrony as a continuum. A partial delayed by thirty milliseconds is heard out of a complex at a mistuning far below the harmonicity threshold, and the ladder’s own first rung has the machinery to draw it and did not.
And it cannot say what a kettledrum sounds like. Coming top of a fusion census is not the same as sounding like one note, and a timpano does not: its pitch is famously weak and famously debated. The census says its partials arrive and leave together, which is a real property and evidently not sufficient.
Whose instruments, and when
The five spectra are the collection’s own and they are a mixture of a definition, two published tables and two derivations. The ideal string is a definition. The piano string’s inharmonicity comes from a stiffness coefficient. The bell’s partial ratios are a founder’s target arrived at by shaving metal off a casting for six hundred years, and the kettledrum’s are Bessel zeros with an air load on them.
The historical reading is about the bell, and it is a nice one. Bell founders spent centuries pushing the partial ratios toward 2, 3 and 4 — improving the harmonicity of the spectrum, deliberately and with instruments they built for the purpose — and did nothing whatever about the decay rates, which are set by the mode shapes and the way the bell is hung. So six hundred years of tuning improved a bell’s score on one of these two cues and left it last on the other, and the sound of a bell is what that combination produces: a strong pitch and a set of components that plainly do not belong to it.
Where this ladder goes next
Four rungs. The ear builds objects and sometimes offers a choice; the same question on the simultaneous axis; a spectrum’s inharmonicity read as a perceptual count; and now the same count under the cue that needs time.
What is owed is the arbitration. This ladder now has two cues that disagree, and every figure in it reports each cue separately because there is no principled way here to weigh one against the other. The published apparatus for that is a competition between grouping hypotheses with a cost per cue — the same shape as the constrained optimisation the orchestration anchor had to adopt to avoid inventing an exchange rate — and the honest next rung is to say what such a competition would need, and which of its parameters this collection could supply.
Part 4 of 8
One essay in the series on auditory scene. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Auditory scene analysisCommon fateEnvelopeFusionInharmonicityPartialTimbreVibrato
- The cue that settles it auditory scene analysis, common fate, fusion, inharmonicity, partial
- A bell has no fundamental fusion, inharmonicity, partial, timbre
- Eleven partials is one too many auditory scene analysis, fusion, inharmonicity, partial
- A hammer is not an impulse envelope, partial, timbre
- A spectrum chooses its own scale inharmonicity, partial, timbre
- The other wolf inharmonicity, partial, timbre