Perception and the listener

The partials that do not die together

The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.

Assumes: The spectrum that will not fuse · What makes two partials one note

The spectrum that will not fuse took five spectra this collection already carried — an ideal string, a piano string, a bell, a bar, a kettledrum — and asked how many of each one’s partials sit near enough to a whole multiple of a single fundamental to be heard as one object. That is harmonicity, which every textbook calls the strongest fusion cue, and a list of ratios is all it needs.

Its last paragraph named the cue after it and why the census could not have it. Components that move together belong together, and every spectrum in that figure is a still photograph. A bell’s partials decay at rates differing by an order of magnitude, so they are emphatically not moving together; a vibrato applied to a harmonic tone moves every partial in exact proportion. Running a second census over the same five spectra with time in it would very likely disagree with the first.

It does. It nearly reverses it.

Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together.
Fig. 1 The two censuses over one set of spectra. Hollow is harmonicity — the fraction of partials near enough a whole multiple of one fundamental. Filled is common fate — the fraction whose decay rate is within a factor of two of the strongest partial’s. The arrows are long and they mostly point left, and the two rankings share almost nothing.

What the second census measures

Common fate is Gestalt vocabulary and it means one thing: components undergoing the same change at the same time are grouped. In hearing it has two limbs, and they are physically different quantities.

Onset and offset synchrony. Components that start together and stop together are one object. This is the limb the five spectra can be asked about, because every one of them is a struck or plucked sound with a decay, and every partial of a struck sound has its own decay rate for a physical reason.

Coherent modulation. Components whose frequencies move together are one object. This is the limb that needs a source that modulates at all, which none of the five does, and which is the second half of this essay.

The first limb needs a decay time per partial, and the reason there is one is different for each spectrum. A string’s higher modes shed energy faster, so its partial decay times fall steeply with mode number. A bell’s hum note rings for half a minute while the partials above it are gone in seconds. A kettledrum is damped almost entirely by radiation into the air, which loads every mode nearly equally — which is why a timpano’s note is short and why every part of it is short.

How long each partial of each spectrum lasts. The decay time of every partial of the five spectra, on a logarithmic axis. A set of components that is moving together is a flat line; one that is not is a slope. an ideal string spans a factor of 10.0, a piano string spans a factor of 5.0, a bell spans a factor of 14.9, a bar spans a factor of 2.4, a kettledrum spans a factor of 1.6. The exponents behind these are stipulated and ordinal — what is published is that a bell's partial decay times span more than an order of magnitude and a kettledrum's barely span a factor of two, and these reproduce that ordering rather than measuring it. The physical reasons differ: a string's higher modes shed energy faster, a bell's hum outlasts everything above it, and a kettledrum is damped almost entirely by radiation into the air, which loads every mode nearly equally.
Fig. 2 Why the second census says what it says: the decay time of every partial of every spectrum, on a logarithmic axis. A set of components moving together is a flat line and one that is not is a slope. The bell falls by a factor of fifteen across ten partials and the kettledrum by less than two.

The inversion

Ranked by harmonicity: ideal string, piano string, bell, kettledrum, bar. Ranked by common fate: kettledrum, bar, ideal string, piano string, bell.

The best spectrum on the first cue is third on the second. The worst is second. The bell is third on the first and last on the second. Nothing survives the change of cue except the piano string’s place in the middle, and that is a coincidence of two different middles.

The extreme case is the one worth stating on its own. An ideal string is a perfect harmonic series and holds only a fifth of its partials together by decay. Its harmonicity is exactly one — every partial is exactly a whole multiple, by construction, because that is what an ideal string is — and its tenth partial dies ten times faster than its first. On the second cue an ideal string is a set of components going their separate ways.

That is not a defect in either census. It is the reason a plucked string sounds the way it does. The note that gets duller as it dies is the essay about exactly that slope, and its finding is that the rate at which the colour drains is an identity cue in its own right — one number that sorts instruments the attack cannot. So the same fact is a fusion failure under one grouping principle and an identity signature under another, which suggests the two principles are not competing to describe one thing.

The case that separates them cleanly

The five spectra disagree about which cue is better satisfied. A stronger test is a stimulus where one cue says one object and the other says two, unambiguously, and it is easy to build.

Take a tone at 220 hertz and a tone an octave above it. Every partial of the upper tone is also a partial of the lower — the union of the two sets is exactly the harmonic series of 220 — so harmonicity has nothing to work with at all. The census fuses sixteen of sixteen and reports one object. That is not an artefact: it is why an octave is the interval that fuses, which what makes two partials one note established from the other side.

Now put a fifty-cent vibrato on the lower tone. Eight partials move and eight do not, and the eight that move do so in exact proportion — fifty cents at the first partial and fifty cents at the eighth, which is six hertz and fifty-one hertz. Common fate separates them on the first cycle.

An octave, with a vibrato on the lower note. Sixteen partials over a second: eight belonging to a tone at 220 hertz with a 50-cent vibrato at 6 hertz, and eight belonging to a tone an octave above it that is perfectly steady. Harmonicity cannot separate them at all — every one of the sixteen is a whole multiple of 220, so the census fuses 16 of 16 into a single series and reports one object. Common fate separates 8 of the 16 on the first cycle: the moving partials correlate with the modulation at better than 0.9 and the steady ones at zero. The excursion is the same fifty cents for every moving partial and a very different number of hertz — 6.4 for the first and 51 for the eighth — which is why the cue works on a log axis and not on a linear one.
Fig. 3 Sixteen partials over one second. The moving ones belong to the lower tone and the steady ones to the octave above it, and every one of the sixteen is a whole multiple of the same fundamental. Harmonicity has no purchase. The correlation between each partial’s track and the modulation is better than 0.9 for the moving set and zero for the other, so the two are separable by a cue that a list of ratios cannot express.

Why the excursion is in cents and the cue is in hertz

There is an arithmetical point hiding in that figure and it is what makes coherent modulation work at all.

A vibrato is a change of fundamental, so every partial moves by the same fraction. On a logarithmic axis they move rigidly — the whole comb slides up and down without changing shape. In hertz they do not: the eighth partial’s excursion is eight times the first’s, fifty-one hertz against six.

That matters because the auditory system’s frequency analysis is not linear either. A partial’s movement has to be large enough relative to the width of the filter it sits in to be detected as movement, and critical bandwidths widen with frequency but less than proportionally. So the higher partials of a vibrato tone are the ones whose movement is most detectable in the units the ear works in, which is the opposite of the usual assumption that the low partials carry everything — which harmonics carry the pitch puts the dominance region for pitch between the third and fifth partials, and the dominance region for this cue is somewhere else entirely.

Computing it is division. Take each partial’s excursion in hertz and divide it by the width of the band it sits in, for a fifty-cent vibrato:

fundamental partial where the ratio is largest the frequency that is
110 Hz the sixteenth 1,760 Hz
220 Hz the eighth 1,760 Hz
440 Hz the fourth 1,760 Hz

The dominance region for coherent modulation is a fixed frequency, not a fixed partial number, and on the Bark bandwidth it sits at about 1.8 kilohertz whatever the fundamental. That is the structural difference from the pitch-dominance region, which is the third to the fifth partial and therefore moves with the note: a singer going up an octave keeps the same partials carrying their pitch and hands the modulation cue to partials half as far up the series.

The two bandwidth models disagree about whether there is a maximum at all, and the disagreement is worth reporting rather than resolving. On the Bark formula the ratio rises to 0.196 at 1,760 hertz and falls away above it. On the equivalent rectangular bandwidth it rises monotonically to a ceiling of 0.271 and is within one per cent of that ceiling from about two kilohertz upward — a plateau rather than a peak. Both put the region well above the pitch one and neither puts it anywhere near the low partials, which is all the argument needs; where exactly it sits depends on which of two fits to listening data is used, and the two are known to disagree by a factor of two low down.

One number from that table is worth carrying on its own. At its most detectable, a fifty-cent vibrato moves a partial by a fifth to a quarter of a critical bandwidth — not across a filter, but a fraction of the way along one. Whatever mechanism reads coherent modulation is reading a change of level inside a channel rather than a partial crossing between channels, and fifty cents is a large vibrato. That is a constraint on any account of how the cue works, and it is the kind of thing a census over five spectra could not have produced.

The census the third rung ran, for the comparison

The point of a second census is that it is the same five spectra and the same shape of question, so the two numbers can be put side by side without an argument about what is being compared. It is worth having the first one in view.

How much of each spectrum a listener can assemble into one noteEach partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there.each partial, at the harmonic number it is nearestan ideal string10 of 10 fusea piano string9 of 10 fuse+1%a bell7 of 8 fuse+20%a bar2 of 6 fuse-78%-38%+21%+4%a kettledrum3 of 5 fuse-2%-4%filled: fuses into the notehollow: heard as a sound of its own
Fig. 4 The earlier figure: each spectrum’s partials against the whole multiples of the fundamental that fits most of them, with the ones near enough to fuse marked. Everything about a spectrum that this picture can express is in a list of ratios, and a decay rate is not a ratio.

Setting the two beside each other also makes a smaller point that is easy to miss. The harmonicity census has a free parameter that is searched — the fundamental — and the common-fate census does not. There is no “best decay rate” to be found; the rates are what they are, and the count is a count. So the second cue is in one sense a stronger measurement and in another a weaker one: it cannot be gamed by a search, and it also cannot discover anything the way the search discovers a bell’s strike note.

What a mistuned partial does to each cue

The ladder’s own first rung is the smallest version of this question: one partial of an otherwise harmonic tone moved off its whole multiple, and the mistuning at which it is heard out of the complex.

That stimulus is a clean test of harmonicity alone, and it is why the number it produces — a partial is heard out at a mistuning of a few per cent — is quoted everywhere. Under common fate it is not a test at all: the mistuned partial starts and stops with all the others, so the cue says it belongs and goes on saying so however far the mistuning is pushed.

Partial 3, mistuned by 3%A 10-partial tone on 220 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 3.0 per cent, which is 19.8 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune.19.8 Hz off12345678910partial numberamplitude1% is enoughto hear the partialout of the note;3% and it stopscontributing to thepitch of the wholeMoore, Peters &Glasberg, 1985
Fig. 5 The original stimulus: the third partial moved three per cent sharp of where it belongs. Harmonicity hears it leave. Common fate does not, and would not at thirty per cent either — which is why the published way of making a partial pop out of a complex is to give it a different onset, and why doing so works at mistunings far smaller than this one.

That is the practical form of the disagreement, and it is the reason the two cues are usually studied with two different stimuli. It is also why the census in the hero figure is worth having: five real spectra, one axis, and no experiment designed to isolate anything.

Which computation produced the numbers

The harmonicity census is fusionCensus, unchanged from the third rung: a search over candidate fundamentals for the one that puts the most partials within the mistuning at which a partial is heard out of its complex, with the search capped so that a fundamental twenty-five harmonics below the top partial cannot be nominated.

The common-fate census gives each spectrum a decay law of the form T60 over n to a power, evaluates it for ten partials, and counts how many fall within a factor of two of the strongest partial’s rate. The tolerance is stated rather than fitted and the result is not sensitive to it: the two rankings are nearly reversed at a factor of two and still nearly reversed at four.

The exponents are asserted and ordinal. What is published, and reproducible across instruments and laboratories, is that a bell’s partial decay times span more than an order of magnitude and a timpano’s barely span a factor of two. The five numbers here reproduce that ordering by construction. They are not a measurement of any instrument, and this collection has now recorded the same caution enough times that it should be read as a standing property of these figures rather than as a disclaimer.

The modulation coherence is the correlation between each partial’s frequency track and the modulating waveform, over whole vibrato cycles. Steady partials have zero variance and are reported as zero rather than as undefined.

Where the model stops

Decay is a single exponential per partial and it is not. Real struck strings and bells show two-stage decays — a fast initial rate and a slower tail — because energy moves between polarisations and between coupled strings. Three strings and the note that comes back is the essay about the mechanism on a piano, and it means a single rate per partial is a summary of something with structure in it.

Onsets are not in the census at all. Onset synchrony is the more powerful half of the first limb — components that start together fuse very strongly — and every spectrum here is treated as starting at one instant. A real strike excites the modes at slightly different times and with very different rise times, which is a whole cue this figure omits.

And a census is not a listener. Counting how many partials satisfy a cue is not a model of grouping; it is a way of making two cues commensurable. Real auditory scene analysis weighs cues against each other, and what happens when harmonicity says one object and common fate says two is an empirical question this arithmetic cannot settle.

What the picture cannot show

It cannot show which cue wins. The octave figure sets up the conflict and stops. What a listener hears when a vibrato is applied to one of two octave-related tones is a published result — the tones do separate — but the degree of separation, and how it trades against the vibrato’s extent, is a listening experiment and not a calculation.

Nor can it show a real vibrato. A singer’s vibrato is not a pure sinusoidal frequency modulation: it carries an amplitude modulation with it, because the vocal tract’s resonances do not move, so partials sweeping past a formant get louder and quieter as they go. That is a second common-fate cue riding on the first, and a note that is never at its pitch is the ladder that computes it.

It cannot show onset asynchrony as a continuum. A partial delayed by thirty milliseconds is heard out of a complex at a mistuning far below the harmonicity threshold, and the ladder’s own first rung has the machinery to draw it and did not.

Partial 3, 40 ms earlyA 10-partial tone on 220 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 0.0 per cent, which is 0.0 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune.19.8 Hz off12345678910partial numberamplitude1% is enoughto hear the partialout of the note;3% and it stopscontributing to thepitch of the wholeMoore, Peters &Glasberg, 1985
Fig. 6 The other half of the first limb, on that same stimulus: the third partial not mistuned at all — its ratio is exactly three — but started forty milliseconds out of step with the rest. Harmonicity is untouched and the partial is heard out anyway. Onset is the strongest form of common fate, and it is the one this essay’s five spectra cannot express, because a struck sound’s partials all start at the strike.

And it cannot say what a kettledrum sounds like. Coming top of a fusion census is not the same as sounding like one note, and a timpano does not: its pitch is famously weak and famously debated. The census says its partials arrive and leave together, which is a real property and evidently not sufficient.

Whose instruments, and when

The five spectra are the collection’s own and they are a mixture of a definition, two published tables and two derivations. The ideal string is a definition. The piano string’s inharmonicity comes from a stiffness coefficient. The bell’s partial ratios are a founder’s target arrived at by shaving metal off a casting for six hundred years, and the kettledrum’s are Bessel zeros with an air load on them.

The historical reading is about the bell, and it is a nice one. Bell founders spent centuries pushing the partial ratios toward 2, 3 and 4 — improving the harmonicity of the spectrum, deliberately and with instruments they built for the purpose — and did nothing whatever about the decay rates, which are set by the mode shapes and the way the bell is hung. So six hundred years of tuning improved a bell’s score on one of these two cues and left it last on the other, and the sound of a bell is what that combination produces: a strong pitch and a set of components that plainly do not belong to it.

Where this ladder goes next

Four rungs. The ear builds objects and sometimes offers a choice; the same question on the simultaneous axis; a spectrum’s inharmonicity read as a perceptual count; and now the same count under the cue that needs time.

What is owed is the arbitration. This ladder now has two cues that disagree, and every figure in it reports each cue separately because there is no principled way here to weigh one against the other. The published apparatus for that is a competition between grouping hypotheses with a cost per cue — the same shape as the constrained optimisation the orchestration anchor had to adopt to avoid inventing an exchange rate — and the honest next rung is to say what such a competition would need, and which of its parameters this collection could supply.

Part 4 of 8

One essay in the series on auditory scene. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Auditory scene analysisCommon fateEnvelopeFusionInharmonicityPartialTimbreVibrato