The cue that settles it
Assumes: The exchange rate nobody has · What makes two partials one note
The exchange rate nobody has built this ladder’s arbitration: a competition between grouping hypotheses, each paying for the harmonicity it violates, for the common fate it violates, and for every stream it posits. It could not supply the rate at which one cue is traded against the other, so it swept it and reported how wide a range of rates leaves the winner unchanged.
Its last paragraph names the cue that is missing entirely:
The competition’s wrong answer about a piano string points straight at it: every partial of a struck string begins at the same instant, that is the strongest grouping cue in the literature, and this ladder has no term for it.
It also says where the input is. The excitation ladder computes what happens when a hammer meets a string, and a struck string’s partials are not quite simultaneous — a stiff string’s high partials travel faster, so the corner’s high components reach the bridge first.
How much first is a number, and it is the whole rung.
Forty-five microseconds against twenty milliseconds
A stiff string’s nth partial travels at a speed raised by the factor √(1 + Bn²), so it completes one traverse in t₁/√(1 + Bn²) rather than t₁. On a piano’s wire at A3 that puts the sixteenth partial a hundred and eleven microseconds ahead of the first.
The asynchrony at which a listener stops hearing a partial as part of a complex is about twenty milliseconds, which this collection has had since the rung about what makes two partials one note.
The margin is a factor of a hundred and eighty.
So the strongest cue in the field is, for a struck string, unanimous: every partial votes to fuse, and there is nothing in the physics of a string that comes anywhere near changing that. That is not a small term to add to a competition. It is a term that pays nothing for the one-stream hypothesis and the full count for every cut.
Three of five verdicts change, all the same way
Adding it at equal weight to the other two changes the verdict on a piano string, a bell and a bar, and leaves an ideal string and a kettledrum where they were. Every change is in the same direction — from splitting into two streams to staying as one — and it has to be, because a cue that votes to fuse on every partial can only push one way.
The bell is the case worth naming. A bell’s partials are wildly inharmonic: the harmonicity cue alone rejects most of them and the fifth rung’s competition duly cuts a bell’s spectrum after its first partial. A bell is heard as one sound. Everybody knows this and the two-cue model cannot produce it.
With the onset cue it comes out. A bell’s partials all start when the clapper hits, the cue is unanimous, and one stream wins at every exchange rate in the swept range.
That is the strongest piece of evidence this rung has, and it is worth being clear that it is evidence rather than a prediction: the answer was known and the model was failing to give it.
And then the exchange rate stops mattering
The fifth rung’s whole apparatus exists because there is a number nobody has. Its report is how many decades of exchange rate leave the winner unchanged, which is a statement about how much the arbitration depends on the missing number.
With the third cue in, the piano string and the bar are settled: one hypothesis wins across the entire swept range, three decades wide, and the exchange rate between the first two cues has no influence at all.
A cue strong enough to be unanimous is a cue that removes a free parameter, and that is a different kind of contribution from a third term with a third weight. The fifth rung’s honest complaint — that its answer depends on a quantity nobody has measured — is answered not by measuring it but by adding something that overwhelms it.
Twenty milliseconds, and it is a number this collection already has
Impose an asynchrony on the upper partials — as though they were coming from a different source — and the fusion holds until exactly the threshold and then goes.
Below twenty milliseconds the winner is one stream at every exchange rate. At twenty and above the winner is a cut, and the arbitration is again completely settled, the other way.
The step is as sharp as it is because the cue is modelled as a threshold rather than as a graded quantity, which is the shape of the published criterion and is not a finding of this rung’s. Where it sits is the finding, because twenty milliseconds is not an arbitrary number in this collection.
It is very nearly the spread between the attacks of two orchestral instruments told to start together. The perceptual-centre ladder puts a trumpet at nine and a half milliseconds after its physical onset and a bowed violin at twenty-eight and a half, and the difference is nineteen.
So two instruments playing one note are within a millisecond of the boundary at which the ear stops fusing them. A doubled unison is not comfortably one sound or comfortably two; it is at the edge, which is why doubling is a technique with a knack to it rather than an operation with a result.
That connects two ladders that were built four phases apart on completely different objects, and neither of them was aiming at the other.
What a hammer contributes, and it is not the asynchrony
There is a second thing the excitation ladder has that belongs here, and it is worth separating from the first because it points the other way.
A hammer is not an impulse computes the force pulse a felt hammer delivers, and its duration — one to two milliseconds in the middle of a piano — low-passes the excitation. What it does not do is delay anything: every partial is excited by the same pulse at the same instant, and a finite contact changes how much of each partial is set going rather than when.
So the two things the excitation ladder knows about a struck note make opposite contributions to this arbitration. The contact time removes energy from the high partials, which weakens them and makes them easier for a listener to overlook. The synchrony binds them, and binds them by a factor of nearly two hundred.
The cue and the amplitude are pulling in the same direction on a piano and for unrelated reasons, which is probably why nobody has needed to separate them: a struck note’s upper partials are both quiet and welded to the fundamental.
What the cue is not
Two things this rung does not claim.
It does not claim the onset cue is right. A cost model with three terms and two free weights is a description of an arbitration, not a theory of one. What it does is make one term’s magnitude a computed quantity instead of an assertion, and the magnitude turns out to be extreme.
And it does not claim struck strings are special. Every impulsively excited source has this property — a bell, a bar, a plucked string, a drum — because a single blow starts everything at once. What differs is a blown or bowed source, where the partials build at their own rates: a wind instrument’s higher modes establish over tens of milliseconds and its onset cue is not unanimous at all.
That is a real prediction and it is the one this rung would most like tested. A struck note should be harder to hear out a mistuned partial from than a blown note of the same spectrum, because the onset cue is defending the fusion in one case and abstaining in the other.
The size of the margin is what makes this a result
It is worth dwelling on the factor of a hundred and eighty, because a smaller one would have made this a different essay.
Had the dispersion put the sixteenth partial five milliseconds ahead rather than a tenth of one, the cue would have been graded: some partials in and some out, a verdict depending on where in the compass the note is, and a third weight to argue about. The arbitration would have got more complicated rather than simpler.
It does not, and what is striking is how little it varies. The spread goes as t₁·Bn²/2, and as the pitch rises t₁ falls while B rises by very nearly as much — a treble string is much shorter and much stiffer for its length — so the two nearly cancel. At A3 the sixteenth partial is 111 microseconds early and at A6 it is 130, three octaves up.
Nothing a string can do brings its own partials near the threshold, anywhere in the compass, and the constancy is a coincidence of a piano’s own scaling rather than a law. That is what makes the cue unanimous rather than merely strong, and unanimity is what removes the free parameter.
Which computation produced the numbers
The onset offsets are the string’s dispersion and nothing else: one traverse takes t₁ = 1/(2f₀) and partial n takes t₁/√(1 + Bn²), with B from this collection’s own piano-string design.
The twenty-millisecond threshold is the ladder’s own from its second rung, and it is a rounded figure from a literature that reports anything between about ten and forty depending on the stimulus.
The competition is the fifth rung’s, unchanged: hypotheses are the ways of cutting a spectrum into two streams at a partial boundary, each pays one unit for every partial it groups against a cue’s verdict and one for every partial it separates against it, plus a cost per extra stream. The onset cue is a third column of verdicts, added at the same weight as harmonicity.
Equal weight is a choice and is stated as one. Any positive weight gives the same answer on these spectra, because the cue is unanimous and the other two are not — which is what makes the choice unimportant here and would not in a case where it was divided.
What it says about the anchor’s own opening question
The ear builds objects opened this anchor by observing that the same physical signal sometimes offers a listener a choice, and that the choice is a real one — a mistuned partial can be heard in the note or out of it, and the same listener can do both.
Three cues later, this rung says something about when the choice is available. It is available exactly when the cues disagree and no one of them is unanimous: on an ideal string the harmonicity cue is unanimous and there is no choice; on a struck bell the onset cue is unanimous and there is none either. The interesting cases are the ones in between, and the competition’s own report — how many hypotheses win somewhere in the swept range of exchange rates — is a measure of exactly that.
So the ambiguity the anchor was opened to describe is a property of a spectrum’s cues disagreeing, and it can be counted. Adding the third cue reduces the count on four of five spectra, which is another way of saying that a listener has fewer real choices than the two-cue model allowed.
Where the model stops
Dispersion is the only asynchrony in the string. A real struck note’s partials also start at different times because the hammer has a finite contact width and because the soundboard’s own modes ring up at their own rates — both of which this collection has and neither of which is in this number. Both are milliseconds rather than microseconds, so they would shrink the four-hundred-fold margin and not close it.
A threshold is not a cue. Real onset-asynchrony effects are graded: a partial five milliseconds early is partly heard out. Modelling that as a sharp line is what makes the breakpoint figure a step.
The weights are still free. Two of the three cues have no measured exchange rate and the third’s weight is asserted. What has changed is that the answer no longer depends on them.
And the sources are all impulsive. Four of the five spectra here are struck or plucked, so the cue’s unanimity is nearly guaranteed by the way the ladder chose its examples.
What the picture cannot show
It cannot show a sequence. Everything here is one event’s partials at one instant, and auditory streaming — the other half of this anchor — is about how events over time are assigned to sources. An onset cue has nothing to say about that.
Nor can it show the room. Reflections arrive tens of milliseconds after the direct sound, which is precisely the scale of the threshold, and a listener in a hall receives every partial twice.
It cannot show attention. A listener told to listen for a mistuned partial hears it far more readily than one listening to a note, which is the standard finding and is not a property of any cost model.
It cannot show a mistuning that is also late. The two cues are added as separate columns, and a real heard-out partial is usually both mistuned and asynchronous — which is how the demonstrations are built. Whether the two combine additively is exactly the kind of exchange-rate question this rung claims to have made unnecessary, and it has only made it unnecessary in the cases where one cue is unanimous.
And it cannot show two of anything. The competition splits one spectrum into two streams; two instruments playing at once is a different problem with a different arithmetic, and the twenty-millisecond coincidence above is a hint at it rather than a treatment of it. Two players on one note is the version of that question this collection has, and it is about spectrum rather than about time.
Whose sounds, and when
The spectra are the ladder’s own idealisations — an ideal string, a piano string, a bell, a bar, a kettledrum — and none of them is a measurement.
The bell is the one with any history in it. Bell founders have spent six centuries shaving metal to bring the partials of a casting toward a set of target ratios that are not a harmonic series, and the instrument is heard as one sound with a definite pitch throughout. That is a strong argument that fusion does not require harmonicity, it was made long before anybody had a cost model, and it was made in bronze.
Where this ladder goes next
Six rungs. The ear builds objects and sometimes offers a choice; the same question on the simultaneous axis; a spectrum’s inharmonicity read as a perceptual count; the same count under the cue that needs time; the two counts arbitrated with the exchange rate swept; and now the cue that makes the exchange rate stop mattering.
What is owed after this is the blown note. Everything here rests on the partials starting together, and on a wind instrument they do not: the ladder that computes settling times has each impedance peak establishing over a number of periods set by its own Q, and those Qs differ up the compass by nearly a factor of two. So a blown note’s onset cue is graded rather than unanimous, its partials arrive over tens of milliseconds rather than microseconds, and the arbitration should go back to depending on a rate nobody has. That would say which families of instrument the ear fuses easily and which it has to work at — a claim about orchestration made entirely from onset physics, and one the two ladders concerned could compute between them tomorrow.
Part 6 of 8
One essay in the series on auditory scene. The essays either side of this one:
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Auditory scene analysisCommon fateFusionInharmonicityOnsetPartialStreaming
- The partials that do not die together auditory scene analysis, common fate, fusion, inharmonicity, partial
- Eleven partials is one too many auditory scene analysis, fusion, inharmonicity, partial
- A bell has no fundamental fusion, inharmonicity, partial
- A voice is a stream, and the ear decides which auditory scene analysis, fusion, streaming
- A bar's partials are the odd numbers, squared inharmonicity, partial
- A beat is never one beat inharmonicity, partial