Perception and the listener

The exchange rate nobody has

There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.

Assumes: The partials that do not die together · The spectrum that will not fuse

The partials that do not die together gave this ladder a second cue and a problem. Harmonicity asks whether a partial is near enough a whole multiple of a fundamental to belong; common fate asks whether it decays at the same rate as the rest. On some spectra the two agree and on others they emphatically do not, and every figure in the ladder reports them side by side because there is nothing here that says which to believe.

Its last paragraph named the apparatus:

The published apparatus for that is a competition between grouping hypotheses with a cost per cue — the same shape as the constrained optimisation the orchestration anchor had to adopt to avoid inventing an exchange rate.

Adopting it does not supply the exchange rate. What it does is say how much the answer depends on one, and that turns out to be the useful reading.

What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged.
Fig. 1 Each spectrum with its two cue readings and the grouping a competition chooses. Beside each is how many decades of exchange rate leave that choice unchanged, which is the quantity the missing number is being priced against.

What a competition is

The apparatus is standard in the auditory scene analysis literature and it is worth stating in full, because its shape is what makes the exchange rate visible.

A hypothesis is a way the partials could divide. The simplest family is a cut: partials one to c belong to one stream and the rest to another, plus the null hypothesis that everything is one stream. That is the division every published demonstration of these cues uses, and it is what a listener reports when a mistuned partial “pops out”.

Each hypothesis pays a cost for every partial it gets wrong according to each cue, and it can be wrong in two ways. It is wrongly grouped when a cue rejects a partial and the hypothesis keeps it with the first stream; it is wrongly separated when a cue accepts a partial and the hypothesis puts it in the second. Charging only for the first makes the competition degenerate, because a cut after the first partial then has almost nothing left to be charged for.

There is a third cost, a fixed charge per stream, which is what stops the competition from splitting every partial into its own object. Its value is asserted and it matters much less than the exchange rate does.

The exchange rate is the number of units of harmonicity evidence one unit of common-fate evidence is worth. It is the thing this collection cannot supply, and the whole design of the rung is to make its absence measurable.

Two cues, five spectra, considerable disagreement

The five spectra are the ones the ladder carries, and the two cues part company on most of them.

spectrum harmonicity common fate
an ideal string 100% 20%
a piano string 90% 20%
a bell 88% 13%
a bar 33% 67%
a kettledrum 60% 100%

The first three are cases where harmonicity says one object and common fate says many. A piano string’s partials are near enough whole multiples to fuse — that is what makes a piano a pitched instrument — and they decay at very different rates, because higher partials lose energy faster. So a piano note is, by the common-fate criterion, a set of eight or nine separate events that happen to start together.

The last two invert. A bar’s partials are nowhere near a harmonic series — the free–free ratios are 1, 2.76, 5.40 — so harmonicity has almost nothing to hold them together, and they decay at similar rates, so common fate does. A kettledrum is the extreme: every partial decays alike and none of them fits a series.

That is a real conflict rather than a modelling artefact. Both cues are well attested, both are computed here by criteria this ladder set out in earlier rungs, and they say opposite things about the most ordinary musical sound there is.

Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together.
Fig. 2 The second cue on its own, computed earlier: how many of each spectrum’s partials decay at a rate close enough to the fundamental’s to be heard as belonging with it. This is the column that disagrees with harmonicity on a piano string.

What the competition decides, and over how wide a range

Running the competition on a bell and sweeping the exchange rate over three orders of magnitude gives a picture with exactly two answers in it.

Below an exchange rate of about 0.7 — that is, when a unit of common-fate evidence is worth less than seven tenths of a unit of harmonicity evidence — one stream wins. Above about 1.0, a cut after the first partial wins: the strike tone separates from everything above it.

The change of mind happens at an exchange rate of about 0.8, which is to say very nearly where a listener weighing the two cues equally would sit. And each of the two answers is stable over about one and a half decades on its own side.

That is the reading worth having. The arbitration is not delicately poised on a number nobody has: it has two answers, they are separated by a boundary at roughly equal weighting, and small errors in the exchange rate do not move it.

It also means the answer is genuinely undetermined for a bell, in a specific and honest sense. A listener who trusts harmonicity slightly more hears one object; a listener who trusts common fate slightly more hears a strike tone and a hum. Both of those are things people report about bells.

The cost of each grouping of a bell, against the exchange rate. The horizontal axis is what one partial of common-fate evidence is worth in units of one partial of harmonicity evidence — the number nobody in this collection can supply. Each line is one hypothesis about how a bell's partials divide, and the lowest line at a given rate is what a listener with that rate would hear. one stream wins from 0.03 to 0.71; a cut after partial 2 wins from 0.84 to 0.84; a cut after partial 1 wins from 1.00 to 31.62. The widest run covers 1.5 decades, so the arbitration is not delicately balanced on a number nobody has — it changes its mind once, near an exchange rate of one, which is exactly where a listener weighing two cues equally would sit.
Fig. 3 The cost of each grouping of a bell against the exchange rate. The lowest line at a given rate is what a listener with that rate hears, and the crossing sits near equal weight, with each answer stable over about a decade and a half.

The piano is where the answer is interesting

The bell is a hard case that has always been described as ambiguous. The piano string is not, and the competition has something to say about it.

Harmonicity gives a piano string 90 per cent and common fate 20, and the competition’s answer over its widest stable range is a cut after the second partial. That is not “one note”; it is a fundamental and a first partial together, and everything above them as a second object.

A listener does not report that. A piano note is heard as one thing, emphatically, and nobody hears the upper partials of a struck string as a separate sound.

So either the exchange rate is far out at the harmonicity end for this stimulus, or the model is missing a cue that binds a piano note together — and there is an obvious candidate. Onset synchrony is the strongest grouping cue there is, and every partial of a struck string starts at the same instant. Neither of the two cues in this ladder has time-of-onset in it at all: harmonicity is a frequency criterion and common fate as computed here is a decay criterion.

That is the ladder’s own gap arriving as a wrong answer, which is the best way for a gap to arrive.

How much of each spectrum a listener can assemble into one noteEach partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there.each partial, at the harmonic number it is nearestan ideal string10 of 10 fusea piano string9 of 10 fuse+1%a bell7 of 8 fuse+20%a bar2 of 6 fuse-78%-38%+21%+4%a kettledrum3 of 5 fuse-2%-4%filled: fuses into the notehollow: heard as a sound of its own
Fig. 4 The first cue on its own, computed earlier: how many partials of each spectrum sit near enough a whole multiple to fuse. This is the column that says a piano note is one object.

One of the cues is very much stronger than the others, and its strength is worth a number because it sets the scale everything else is traded against.

The partials start together by a factor of four hundred. How far each partial of a struck piano string is from the first in its onset, computed from the string's own dispersion: a stiff string's high partials travel faster, so component n reaches the bridge ahead of component 1 by t₁(1 − 1/√(1 + Bn²)). The largest offset here is 111 microseconds at the sixteenth partial. The line across the top is 20 milliseconds, which is the asynchrony at which a partial stops being heard as part of the note. The margin is a factor of 180. So the cue every account of grouping calls the strongest is, for this source, unanimous: every partial votes to fuse, and nothing in the physics of the string comes near to changing that.
Fig. 5 How far each partial of a struck piano string is from the first in its onset, computed from the string’s own dispersion: a stiff string’s high partials travel faster, so component n reaches the bridge ahead of the fundamental. The largest offset here is 111 microseconds.

The partials start together by a factor of about four hundred against the two-millisecond window inside which onsets are heard as simultaneous. So onset synchrony is not a cue a piano note is passing narrowly; it is one it passes by orders of magnitude, which is why every other cue in this essay is being asked to do the discriminating.

The two extremes agree, and that is a check

The five spectra are not equally interesting and the two that are least interesting are the ones worth checking first.

A kettledrum has both cues pointing one way: 60 per cent harmonicity and 100 per cent common fate, and the competition says one stream over the widest range of any spectrum here, 1.65 decades. That is right, and it is right for a reason the ladder already has — what a drum is doing instead shows that a timpani’s kettle pulls its modes toward a series precisely so that the instrument has a pitch, and a listener hears one.

An ideal string has the cues pointing opposite ways in the most extreme way possible — 100 per cent against 20 — and the competition still says one stream. That is the stream cost doing its job: with harmonicity perfect, no cut can improve the harmonicity term at all, so a cut has to pay for itself entirely out of common fate and cannot.

Those two are what a working model has to get right before its answers on the hard cases mean anything. Getting the drum right rules out a competition that splits everything, and getting the ideal string right rules out one that splits nothing until the cues are unanimous.

The middle three are where the arbitration is doing work, and the piano is the one that goes wrong.

How sharp each partial of a real string is. The departure of each partial from the whole-number multiple it is supposed to be, in cents, for three strings — an ideal string, middle of a small upright, top octave of the same piano. The coefficient of each is computed from the stiffness of a steel wire of a stated length and gauge; nothing is fitted. The sharpest of them is 497 cents sharp by the eighth partial.
Fig. 6 The spectra the argument is about, as ratios. Two of them have almost nothing for harmonicity to work with, which is why common fate carries the whole verdict there and why the exchange rate barely matters at those two.

What this shares with the orchestration anchor

The shape of this rung was borrowed deliberately, and the resemblance is worth making explicit because it is a technique this collection now uses twice.

Who plays what and how loud is one question faced the same problem from the other direction: two quantities in different units — a loudness in sones and a roughness in whatever roughness is — with no exchange rate between them. Its solution was to make one of them a constraint and the other an objective, so that no rate was needed: every candidate is held at the same loudness by construction and only the roughness is minimised.

That option is not available here, because neither cue is naturally a requirement. A listener is not obliged to satisfy harmonicity and then optimise common fate; both are evidence, and evidence combines by weighting.

So the second-best move is the one taken: keep the rate as a free parameter, sweep it, and report how much of the answer survives not knowing it. That converts a missing number from a hole in the model into a measured sensitivity, which is a weaker result and an honest one.

What a missing number is worth reporting as

There is a habit here worth naming, because this is the third place in the collection it has been needed.

A model with a parameter nobody can supply has three honest options. It can avoid the parameter, which is what the orchestration anchor did by making one quantity a constraint. It can assert a value and say so, which is what most of this collection does with published constants. Or it can sweep it and report the sensitivity, which is what this rung does.

The third is the weakest and the most informative when it works, because what comes out is not an answer but a shape: how much of the conclusion is a consequence of the evidence and how much is a consequence of the guess. A parameter whose sweep changes nothing was never worth arguing about; one whose sweep changes everything is the whole model in disguise.

Here it comes out in the middle. Four of the five spectra keep their answer over a decade and a half of exchange rate, and the boundary on the fifth sits at almost exactly equal weighting — so the arbitration is doing real work and is not resting on the missing number.

The same discipline has now been applied to a tempo, to a profile exponent and to an exchange rate on three different ladders. Two of the three sweeps said the parameter did not decide the answer, and saying so is what makes the third one’s boundary worth believing.

Which computation produced the numbers

The harmonicity verdict per partial is the third rung’s: a partial is heard out if it is more than about one per cent from the best whole-number series through the spectrum, which is the same criterion the mistuned-partial rung used for a single partial applied to all of them at once.

The common-fate verdict per partial is the fourth rung’s decay law, one criterion at a time: a partial belongs if its sixty-decibel decay time is within a factor of two of the fundamental’s. The fourth rung reported that as a count; the competition needs the verdict per partial, which is the same criterion evaluated individually.

The cost of a hypothesis is the number of harmonicity mistakes, plus the exchange rate times the number of common-fate mistakes, plus a fixed charge per additional stream. The sweep is forty-one exchange rates over three orders of magnitude, and the reported stability is the widest run of consecutive rates over which one hypothesis wins.

Where the model stops

Onset synchrony is not in it, and the piano result above says how much that costs. Adding it would need a third cue with its own exchange rate, which is the direction this apparatus grows in and is not obviously an improvement.

The hypotheses are cuts. A real grouping can put any subset of partials together, and there is no reason a listener’s alternative to “one object” should be “everything above partial c”. Cuts are what the demonstrations use and they are a small corner of the space. What makes two partials one note is the rung about the simplest non-cut case — one partial removed from the middle — which this family cannot represent at all.

And the spectra are ratios rather than sounds. A bell has no fundamental shows that a bell’s perceptual pitch is supplied by a listener rather than present in the spectrum, so the object the cues are trying to group is not the object a listener reports hearing.

The stream cost is asserted and it decides how readily anything splits at all. It matters less than the exchange rate because it enters once rather than per partial, and it is not free of consequence.

And a cue is not a binary. Both verdicts here are thresholds applied to continuous quantities — how far from a whole multiple, how far from the fundamental’s decay time — and a real competition would weight by how badly each partial fails rather than by whether it fails. That is a straightforward extension and it would make the boundaries softer rather than moving them.

What the picture cannot show

It cannot show a listener. Everything here is a scoring of hypotheses; whether a listener performs anything like this optimisation, or arrives at a percept some other way, is exactly the question the apparatus was invented to model and does not answer.

Nor can it show the alternation. The ear builds objects is about a case where a listener’s grouping flips over time with the stimulus unchanged, which is the strongest evidence that grouping is a competition — and a competition with one winner cannot produce it. Bistability needs two nearly-equal costs and a dynamics, and this has the first and not the second.

And it cannot show context. Which grouping a listener hears depends on what came before: a partial that popped out on the previous note pops out on this one. Nothing in a per-spectrum competition has a memory, and a voice is a stream is the essay about the sequential version of the same problem, where memory is the whole subject.

Nor can it show the room. Every spectrum here is a source in isolation, and a listener in a hall receives it with the reflections a position and a width describes laid over it — which comb the spectrum and change every amplitude the two cues are computed from.

A 7-semitone sequence at 100 ms a tone. Tones drawn as pitch against time, one bar per tone. The events are the same in both readings of this pattern; what changes is whether a listener assigns them to one line that leaps back and forth or to two lines that each stay put. Nothing in the drawing decides which, and nothing in the sound does either.
Fig. 7 The other axis of the same question, opened earlier: a sequence fast enough and far enough apart splits into two streams. Everything in this essay is the simultaneous version, and the two have never been given one apparatus.

Whose sounds, and when

The five spectra are the collection’s standing set: an ideal string, a piano string with its measured stiffness, a bell’s founder’s ratios, a free–free bar and a tuned kettledrum. Three of the five are instruments a Western orchestra contains and two — the bell and the bar — are the standard laboratory cases for inharmonic spectra.

The competition itself belongs to the auditory scene analysis literature of the 1970s onwards, where it is usually stated informally as a set of grouping principles that “compete”, and formalised in various later models as a cost or a probability. The version here is the plainest form of it, and its shape rather than its details is what has been borrowed.

The disagreement it is applied to is older than either. That a bell has a strike note and a hum that can be attended to separately is a description going back centuries in founding, and it is exactly the ambiguity the sweep locates.

Where this ladder goes next

Five rungs. The ear builds objects and sometimes offers a choice; the same question on the simultaneous axis; a spectrum’s inharmonicity read as a perceptual count; the same count under the cue that needs time; and now the two counts arbitrated, with the arbitration’s dependence on a missing number measured rather than hidden.

What is owed after this is the onset. The competition’s wrong answer about a piano string points straight at it: every partial of a struck string begins at the same instant, that is the strongest grouping cue in the literature, and this ladder has no term for it. The excitation ladder has the arithmetic — a hammer is not an impulse computes when each partial is excited and by how much — so the input exists here already, and adding it would be the first cue in this anchor that comes from the physics of the instrument rather than from the shape of the spectrum.

Part 5 of 8

One essay in the series on auditory scene. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Auditory scene analysisCommon fateDecayFusionHarmonicityOptimisationPartialStreaming