Perception and the listener

The cue that settles it

Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.

Assumes: The exchange rate nobody has · What makes two partials one note

The exchange rate nobody has built this ladder’s arbitration: a competition between grouping hypotheses, each paying for the harmonicity it violates, for the common fate it violates, and for every stream it posits. It could not supply the rate at which one cue is traded against the other, so it swept it and reported how wide a range of rates leaves the winner unchanged.

Its last paragraph names the cue that is missing entirely:

The competition’s wrong answer about a piano string points straight at it: every partial of a struck string begins at the same instant, that is the strongest grouping cue in the literature, and this ladder has no term for it.

It also says where the input is. The excitation ladder computes what happens when a hammer meets a string, and a struck string’s partials are not quite simultaneous — a stiff string’s high partials travel faster, so the corner’s high components reach the bridge first.

How much first is a number, and it is the whole rung.

The partials start together by a factor of four hundred. How far each partial of a struck piano string is from the first in its onset, computed from the string's own dispersion: a stiff string's high partials travel faster, so component n reaches the bridge ahead of component 1 by t₁(1 − 1/√(1 + Bn²)). The largest offset here is 111 microseconds at the sixteenth partial. The line across the top is 20 milliseconds, which is the asynchrony at which a partial stops being heard as part of the note. The margin is a factor of 180. So the cue every account of grouping calls the strongest is, for this source, unanimous: every partial votes to fuse, and nothing in the physics of the string comes near to changing that.
Fig. 1 Each partial of a struck piano string, by how far ahead of the fundamental its onset arrives. The line across the top is the asynchrony at which a partial stops being heard as part of the note.

Forty-five microseconds against twenty milliseconds

A stiff string’s nth partial travels at a speed raised by the factor √(1 + Bn²), so it completes one traverse in t₁/√(1 + Bn²) rather than t₁. On a piano’s wire at A3 that puts the sixteenth partial a hundred and eleven microseconds ahead of the first.

The asynchrony at which a listener stops hearing a partial as part of a complex is about twenty milliseconds, which this collection has had since the rung about what makes two partials one note.

The margin is a factor of a hundred and eighty.

So the strongest cue in the field is, for a struck string, unanimous: every partial votes to fuse, and there is nothing in the physics of a string that comes anywhere near changing that. That is not a small term to add to a competition. It is a term that pays nothing for the one-stream hypothesis and the full count for every cut.

Three of five verdicts change, all the same way

The cue that settles it. Every spectrum to hand, arbitrated by the earlier competition and then again with the onset cue added at equal weight. 3 of the 5 change their verdict, and all 3 change the same way — from splitting into two streams to staying as one: a piano string, a bell, a bar. Nothing changes the other way, because the onset cue on a struck source votes for fusion on every partial and can only ever push toward one stream. The bell is the case worth naming: its partials are wildly inharmonic and it is heard as one sound, which is a fact the harmonicity cue alone cannot produce.
Fig. 2 Every spectrum to hand, arbitrated by the two earlier cues and then again with the onset cue added at equal weight.

Adding it at equal weight to the other two changes the verdict on a piano string, a bell and a bar, and leaves an ideal string and a kettledrum where they were. Every change is in the same direction — from splitting into two streams to staying as one — and it has to be, because a cue that votes to fuse on every partial can only push one way.

The bell is the case worth naming. A bell’s partials are wildly inharmonic: the harmonicity cue alone rejects most of them and the fifth rung’s competition duly cuts a bell’s spectrum after its first partial. A bell is heard as one sound. Everybody knows this and the two-cue model cannot produce it.

With the onset cue it comes out. A bell’s partials all start when the clapper hits, the cue is unanimous, and one stream wins at every exchange rate in the swept range.

That is the strongest piece of evidence this rung has, and it is worth being clear that it is evidence rather than a prediction: the answer was known and the model was failing to give it.

And then the exchange rate stops mattering

The fifth rung’s whole apparatus exists because there is a number nobody has. Its report is how many decades of exchange rate leave the winner unchanged, which is a statement about how much the arbitration depends on the missing number.

With the third cue in, the piano string and the bar are settled: one hypothesis wins across the entire swept range, three decades wide, and the exchange rate between the first two cues has no influence at all.

Twenty milliseconds, and it stops being one note. The same spectrum with an asynchrony imposed on its upper partials, from nothing to sixty milliseconds. Up the axis is how many different hypotheses win over the swept range of exchange rates — three means the answer depends on a number nobody has, and one means it does not. Below 20 milliseconds the winner is one stream at every exchange rate that matters; at and above it the winner is a cut, and the arbitration becomes completely settled. The step is as sharp as it is because the cue is modelled as a threshold rather than as a graded quantity, which is the shape of the published criterion and not a finding. What is a finding is where it sits: 20 milliseconds is the spread between the attacks of two orchestral instruments told to play together.
Fig. 3 The same spectrum with an asynchrony imposed on its upper partials. Up the axis is how many hypotheses win somewhere in the swept range; one means the exchange rate does not matter.

A cue strong enough to be unanimous is a cue that removes a free parameter, and that is a different kind of contribution from a third term with a third weight. The fifth rung’s honest complaint — that its answer depends on a quantity nobody has measured — is answered not by measuring it but by adding something that overwhelms it.

Twenty milliseconds, and it is a number this collection already has

Impose an asynchrony on the upper partials — as though they were coming from a different source — and the fusion holds until exactly the threshold and then goes.

Below twenty milliseconds the winner is one stream at every exchange rate. At twenty and above the winner is a cut, and the arbitration is again completely settled, the other way.

The step is as sharp as it is because the cue is modelled as a threshold rather than as a graded quantity, which is the shape of the published criterion and is not a finding of this rung’s. Where it sits is the finding, because twenty milliseconds is not an arbitrary number in this collection.

It is very nearly the spread between the attacks of two orchestral instruments told to start together. The perceptual-centre ladder puts a trumpet at nine and a half milliseconds after its physical onset and a bowed violin at twenty-eight and a half, and the difference is nineteen.

So two instruments playing one note are within a millisecond of the boundary at which the ear stops fusing them. A doubled unison is not comfortably one sound or comfortably two; it is at the edge, which is why doubling is a technique with a knack to it rather than an operation with a result.

That connects two ladders that were built four phases apart on completely different objects, and neither of them was aiming at the other.

What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged.
Fig. 4 The earlier competition, without the third cue. Everything on this page is one more column added to this cost.

What a hammer contributes, and it is not the asynchrony

There is a second thing the excitation ladder has that belongs here, and it is worth separating from the first because it points the other way.

A hammer is not an impulse computes the force pulse a felt hammer delivers, and its duration — one to two milliseconds in the middle of a piano — low-passes the excitation. What it does not do is delay anything: every partial is excited by the same pulse at the same instant, and a finite contact changes how much of each partial is set going rather than when.

So the two things the excitation ladder knows about a struck note make opposite contributions to this arbitration. The contact time removes energy from the high partials, which weakens them and makes them easier for a listener to overlook. The synchrony binds them, and binds them by a factor of nearly two hundred.

The cue and the amplitude are pulling in the same direction on a piano and for unrelated reasons, which is probably why nobody has needed to separate them: a struck note’s upper partials are both quiet and welded to the fundamental.

A3: the pulse computed and the pulse assumedAbove, the force the hammer delivers to the string at A3, integrated forward against the felt's nonlinear force and the string's returning corner, drawn against the half-sine of 1.60 milliseconds that every earlier figure assumed. The computed contact lasts 2.38 milliseconds and the corner comes home 4.2 times inside it. Below, the excitation each pulse gives to each partial. They agree at the bottom and part company higher up — worst at partial 23, by 24 decibels — because the assumed pulse has nulls the computed one does not.<,c,l,i,p,P,a,t,h, ,i,d,=,",q,a,f,i,g,u,i,d,q,a,a,-,p,l,o,t,",>,<,r,e,c,t, ,x,=,",7,2,", ,y,=,",3,0,", ,w,i,d,t,h,=,",5,8,2,", ,h,e,i,g,h,t,=,",1,2,8,", ,/,>,<,/,c,l,i,p,P,a,t,h,>,<,c,l,i,p,P,a,t,h, ,i,d,=,",q,a,f,i,g,u,i,d,q,a,b,-,p,l,o,t,",>,<,r,e,c,t, ,x,=,",7,2,", ,y,=,",1,6,", ,w,i,d,t,h,=,",5,8,2,", ,h,e,i,g,h,t,=,",1,7,0,", ,/,>,<,/,c,l,i,p,P,a,t,h,>00.511.522.53milliseconds of contactcomputedassumed half-sine24681012141618202224-60-40-20partial numberexcitation, decibels
Fig. 5 The pulse that does the exciting, from the excitation model. Its length decides how much of each partial is set going and its instant decides when — and the second of those is what this essay is about.

What the cue is not

Two things this rung does not claim.

It does not claim the onset cue is right. A cost model with three terms and two free weights is a description of an arbitration, not a theory of one. What it does is make one term’s magnitude a computed quantity instead of an assertion, and the magnitude turns out to be extreme.

And it does not claim struck strings are special. Every impulsively excited source has this property — a bell, a bar, a plucked string, a drum — because a single blow starts everything at once. What differs is a blown or bowed source, where the partials build at their own rates: a wind instrument’s higher modes establish over tens of milliseconds and its onset cue is not unanimous at all.

That is a real prediction and it is the one this rung would most like tested. A struck note should be harder to hear out a mistuned partial from than a blown note of the same spectrum, because the onset cue is defending the fusion in one case and abstaining in the other.

Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together.
Fig. 6 An earlier cue for comparison: whether the partials die together, which is the other computed term in the competition and is far less decisive than this one.
How much of each spectrum a listener can assemble into one noteEach partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there.each partial, at the harmonic number it is nearestan ideal string10 of 10 fusea piano string9 of 10 fuse+1%a bell7 of 8 fuse+20%a bar2 of 6 fuse-78%-38%+21%+4%a kettledrum3 of 5 fuse-2%-4%filled: fuses into the notehollow: heard as a sound of its own
Fig. 7 The harmonicity cue on its own, computed earlier: how many partials of each spectrum are far enough from a harmonic to be heard out. The bell’s row is the one the onset cue overturns.

The size of the margin is what makes this a result

It is worth dwelling on the factor of a hundred and eighty, because a smaller one would have made this a different essay.

Had the dispersion put the sixteenth partial five milliseconds ahead rather than a tenth of one, the cue would have been graded: some partials in and some out, a verdict depending on where in the compass the note is, and a third weight to argue about. The arbitration would have got more complicated rather than simpler.

It does not, and what is striking is how little it varies. The spread goes as t₁·Bn²/2, and as the pitch rises t₁ falls while B rises by very nearly as much — a treble string is much shorter and much stiffer for its length — so the two nearly cancel. At A3 the sixteenth partial is 111 microseconds early and at A6 it is 130, three octaves up.

Nothing a string can do brings its own partials near the threshold, anywhere in the compass, and the constancy is a coincidence of a piano’s own scaling rather than a law. That is what makes the cue unanimous rather than merely strong, and unanimity is what removes the free parameter.

Which computation produced the numbers

The onset offsets are the string’s dispersion and nothing else: one traverse takes t₁ = 1/(2f₀) and partial n takes t₁/√(1 + Bn²), with B from this collection’s own piano-string design.

The twenty-millisecond threshold is the ladder’s own from its second rung, and it is a rounded figure from a literature that reports anything between about ten and forty depending on the stimulus.

The competition is the fifth rung’s, unchanged: hypotheses are the ways of cutting a spectrum into two streams at a partial boundary, each pays one unit for every partial it groups against a cue’s verdict and one for every partial it separates against it, plus a cost per extra stream. The onset cue is a third column of verdicts, added at the same weight as harmonicity.

Equal weight is a choice and is stated as one. Any positive weight gives the same answer on these spectra, because the cue is unanimous and the other two are not — which is what makes the choice unimportant here and would not in a case where it was divided.

What it says about the anchor’s own opening question

The ear builds objects opened this anchor by observing that the same physical signal sometimes offers a listener a choice, and that the choice is a real one — a mistuned partial can be heard in the note or out of it, and the same listener can do both.

Three cues later, this rung says something about when the choice is available. It is available exactly when the cues disagree and no one of them is unanimous: on an ideal string the harmonicity cue is unanimous and there is no choice; on a struck bell the onset cue is unanimous and there is none either. The interesting cases are the ones in between, and the competition’s own report — how many hypotheses win somewhere in the swept range of exchange rates — is a measure of exactly that.

So the ambiguity the anchor was opened to describe is a property of a spectrum’s cues disagreeing, and it can be counted. Adding the third cue reduces the count on four of five spectra, which is another way of saying that a listener has fewer real choices than the two-cue model allowed.

Where the model stops

Dispersion is the only asynchrony in the string. A real struck note’s partials also start at different times because the hammer has a finite contact width and because the soundboard’s own modes ring up at their own rates — both of which this collection has and neither of which is in this number. Both are milliseconds rather than microseconds, so they would shrink the four-hundred-fold margin and not close it.

A threshold is not a cue. Real onset-asynchrony effects are graded: a partial five milliseconds early is partly heard out. Modelling that as a sharp line is what makes the breakpoint figure a step.

The weights are still free. Two of the three cues have no measured exchange rate and the third’s weight is asserted. What has changed is that the answer no longer depends on them.

And the sources are all impulsive. Four of the five spectra here are struck or plucked, so the cue’s unanimity is nearly guaranteed by the way the ladder chose its examples.

What the picture cannot show

It cannot show a sequence. Everything here is one event’s partials at one instant, and auditory streaming — the other half of this anchor — is about how events over time are assigned to sources. An onset cue has nothing to say about that.

Nor can it show the room. Reflections arrive tens of milliseconds after the direct sound, which is precisely the scale of the threshold, and a listener in a hall receives every partial twice.

It cannot show attention. A listener told to listen for a mistuned partial hears it far more readily than one listening to a note, which is the standard finding and is not a property of any cost model.

It cannot show a mistuning that is also late. The two cues are added as separate columns, and a real heard-out partial is usually both mistuned and asynchronous — which is how the demonstrations are built. Whether the two combine additively is exactly the kind of exchange-rate question this rung claims to have made unnecessary, and it has only made it unnecessary in the cases where one cue is unanimous.

And it cannot show two of anything. The competition splits one spectrum into two streams; two instruments playing at once is a different problem with a different arithmetic, and the twenty-millisecond coincidence above is a hint at it rather than a treatment of it. Two players on one note is the version of that question this collection has, and it is about spectrum rather than about time.

Whose sounds, and when

The spectra are the ladder’s own idealisations — an ideal string, a piano string, a bell, a bar, a kettledrum — and none of them is a measurement.

The bell is the one with any history in it. Bell founders have spent six centuries shaving metal to bring the partials of a casting toward a set of target ratios that are not a harmonic series, and the instrument is heard as one sound with a definite pitch throughout. That is a strong argument that fusion does not require harmonicity, it was made long before anybody had a cost model, and it was made in bronze.

Where this ladder goes next

Six rungs. The ear builds objects and sometimes offers a choice; the same question on the simultaneous axis; a spectrum’s inharmonicity read as a perceptual count; the same count under the cue that needs time; the two counts arbitrated with the exchange rate swept; and now the cue that makes the exchange rate stop mattering.

What is owed after this is the blown note. Everything here rests on the partials starting together, and on a wind instrument they do not: the ladder that computes settling times has each impedance peak establishing over a number of periods set by its own Q, and those Qs differ up the compass by nearly a factor of two. So a blown note’s onset cue is graded rather than unanimous, its partials arrive over tens of milliseconds rather than microseconds, and the arbitration should go back to depending on a rate nobody has. That would say which families of instrument the ear fuses easily and which it has to work at — a claim about orchestration made entirely from onset physics, and one the two ladders concerned could compute between them tomorrow.

Part 6 of 8

One essay in the series on auditory scene. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Auditory scene analysisCommon fateFusionInharmonicityOnsetPartialStreaming