Perception and the listener

A blown note does not start late, it starts slowly

Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.

Assumes: The cue that settles it · A note takes a number of periods to speak

The cue that settles it did something a cost model rarely gets to do: it removed a free parameter. The exchange rate nobody has had built a competition between grouping hypotheses that could not say how much a violation of harmonicity is worth against a violation of common fate, so it swept the rate and reported how wide a range left the winner unchanged. Adding onset synchrony settled three spectra outright, because on a struck string that cue is unanimous — every partial votes to fuse and the argument between the other two stops mattering.

Its last paragraph says where that stops working:

A blown or bowed source, where the partials build at their own rates: a wind instrument’s higher modes establish over tens of milliseconds and its onset cue is not unanimous at all.

That prediction is half right, and the half that is wrong is the important half.

A clarinet's partials, each on its own resonance. The 8 partials of the clarinet's chalumeau D that ride an impedance peak, each building toward its steady amplitude as 1 − exp(−t/τ) with τ = Q/πf from that peak's own Q. The time constants run from 11.9 milliseconds to 64.5, so the partials do not arrive at different times — they all begin the instant the reed does — and what differs is how fast each approaches its final level. The horizontal bars are how far apart the first and last are at three criteria: 5.5 ms at 10 per cent, 36.4 ms at 50 per cent, 121.0 ms at 90 per cent. A twenty-millisecond asynchrony is the threshold for hearing a partial out of a note, and this note crosses it at 32 per cent of steady amplitude — so whether a blown note's onset cue is unanimous or divided is decided entirely by how far along a partial has to be before it counts as having started.
Fig. 1 Eight partials of a clarinet’s chalumeau D, each riding its own impedance peak and building toward its steady level at that peak’s own rate. The bars are how far apart the first and last are at three criteria: 5.5 milliseconds at a tenth of the steady amplitude, 36 at a half and 121 at nine tenths.

The tens of milliseconds are real

A note takes a number of periods to speak establishes the whole machinery. A driven resonance builds toward its steady amplitude as one minus a decaying exponential, and the time constant of that exponential is the resonance’s own quality factor over pi times its frequency. The number of periods is the Q; the number of milliseconds is Q over f, and the two order instruments differently.

A clarinet’s chalumeau D sits on its first impedance peak at 126 hertz, and because a cylinder closed at one end supports only odd harmonics, its partials are the third, fifth, seventh and so on — each of which lands on the next peak up. The Qs of those peaks rise steeply, from 26 at the bottom to 72 at the eighth, while the frequencies rise faster still. So the time constants fall, from 64.5 milliseconds at the fundamental to 11.9 at the fifteenth harmonic.

At nine tenths of the steady amplitude those eight partials are spread over 121 milliseconds. On a saxophone it is 151, on a trombone 92, on a trumpet 53 and on a horn 49. Tens of milliseconds, exactly as predicted, against a fusion threshold of twenty.

And it does not matter, because a spread of settling times is not an asynchrony.

Nothing here starts late

Every one of those partials begins at the instant the reed does. A resonance driven from rest starts responding immediately; what its Q decides is how long it takes to get where it is going. There is no partial on that figure that is silent while another sounds, and no instant at which any of them has not begun.

The onset cue is a claim about beginnings. What makes two partials one note states the criterion: a partial displaced by about twenty milliseconds from the rest is heard out of the note as a whistle of its own, and the demonstrations that establish it delay one partial’s start against the others. That is not what is happening here.

So the right question is not how far apart the partials are when the note has arrived. It is how far apart they are at the beginning, and the answer depends entirely on where the beginning is put.

The spread at a criterion of p is the negative logarithm of one minus p, times the difference between the largest and smallest time constant. That is a single quantity — call it the spread of rates — multiplied by a number that grows without limit as p approaches one. At a tenth of the steady amplitude the multiplier is 0.105; at nine tenths it is 2.30, twenty-two times larger.

At a tenth of steady amplitude the clarinet’s eight partials are spread over 5.5 milliseconds, and a saxophone’s over 6.9, a trombone’s over 4.2, a trumpet’s over 2.4 and a horn’s over 2.2. Every one of those is comfortably inside twenty.

The partials start together by a factor of four hundred. How far each partial of a struck piano string is from the first in its onset, computed from the string's own dispersion: a stiff string's high partials travel faster, so component n reaches the bridge ahead of component 1 by t₁(1 − 1/√(1 + Bn²)). The largest offset here is 142 microseconds at the sixteenth partial. The line across the top is 20 milliseconds, which is the asynchrony at which a partial stops being heard as part of the note. The margin is a factor of 140. So the cue every account of grouping calls the strongest is, for this source, unanimous: every partial votes to fuse, and nothing in the physics of the string comes near to changing that.
Fig. 2 The struck string computed earlier, for comparison: dispersion puts its high partials a few tens of microseconds ahead of its fundamental, a margin of three hundred against the same threshold.

So the cue is still unanimous, and the margin is a hundredth of what it was

Against a struck string the difference is enormous. A stiff string’s dispersion puts its sixteenth partial 63 microseconds ahead of its first at the trumpet’s own written middle, which is a margin of three hundred against the twenty-millisecond threshold. A clarinet’s margin at the same criterion is 3.6.

Both are margins. Neither cue is divided; both vote to fuse on every partial; and the arbitration settles the same way for both.

So the prediction that the exchange rate would come back is refused. Adding the onset cue to the competition still removes the free parameter, on a blown note as on a struck one, at any criterion under about a quarter of the steady amplitude. The sixth rung’s result survives being taken to the family of instruments it expected to break it.

What has happened instead is that the safety margin has fallen by a factor of about a hundred. On a struck string the cue is unanimous by so much that nothing about the model’s assumptions could change the verdict; on a blown note it is unanimous by a factor of three or four, which is inside the uncertainty of the threshold itself — a literature that reports anything between ten and forty milliseconds depending on the stimulus.

And the missing number is a different missing number

This is the honest result of the rung and it is worth stating as plainly as possible.

The sixth rung removed a quantity nobody had — the exchange rate between two cues — by adding a third cue whose magnitude it could compute. The blown note does not put that quantity back. It introduces a different one: how far along a rising partial has to be before the auditory system counts it as having started. Nothing in this collection has that number and nothing in the literature the collection reads states it, because every published measurement of onset asynchrony uses stimuli that switch on.

Which families the ear has to work at, from onset physics alone. For each wind instrument in this collection, the fraction of the steady amplitude at which its partials are twenty milliseconds out of step — the threshold at which a partial stops being heard as part of the note. an alto saxophone reaches it at 26 per cent and an F horn at 61, and the ordering is the reverse of the spread of time constants: 66, 53, 40, 23, 21 milliseconds. The reeds are at the bottom and the horn at the top, and the reason is the bell: the horn's impedance peaks lose their Q at 1693 hertz, so only 4 partials of its written middle F ride a resonance at all and the rest are radiated straight out at whatever rate the lips impose. An instrument with few resonant partials has little to spread.
Fig. 3 For each wind instrument, the fraction of the steady amplitude at which its partials are twenty milliseconds out of step. Below its own bar an instrument’s onset cue is unanimous and above it the cue is divided, so the whole census is a statement about one unmeasured number.

The crossing is at 26 per cent for a saxophone, 30 for a clarinet, 34 for a trombone, 54 for a trumpet and 61 for a horn. A criterion below a quarter fuses every wind instrument; a criterion above two thirds divides all of them; and in between the answer is a list.

That is not a failure of the rung. It is the rung: a quantity that was hidden inside the phrase “the partials start together” turns out to be doing work, and it can be priced. Every claim below is stated as an ordering for that reason, because the ordering does not depend on where the criterion is put.

The ordering is an orchestration claim

The reeds come first and the brass last, and the horn is last of all.

An F horn's partials, each on its own resonance. The 4 partials of F4, the horn's written middle that ride an impedance peak, each building toward its steady amplitude as 1 − exp(−t/τ) with τ = Q/πf from that peak's own Q. The time constants run from 10.8 milliseconds to 32.1, so the partials do not arrive at different times — they all begin the instant the reed does — and what differs is how fast each approaches its final level. The horizontal bars are how far apart the first and last are at three criteria: 2.2 ms at 10 per cent, 14.8 ms at 50 per cent, 49.1 ms at 90 per cent. A twenty-millisecond asynchrony is the threshold for hearing a partial out of a note, and this note crosses it at 61 per cent of steady amplitude — so whether a blown note's onset cue is unanimous or divided is decided entirely by how far along a partial has to be before it counts as having started.
Fig. 4 The same computation on an F horn’s written middle F. Four partials rather than eight, time constants from 10.8 to 32.1 milliseconds rather than 11.9 to 64.5, and the whole spread a third of the clarinet’s.

The mechanism is not the one the ordering suggests, and it is worth separating them. It would be natural to guess that the brass have more uniform Qs. They do not: a trombone’s supported partials run from a time constant of 7.6 to 47.6, which is a wider ratio than the clarinet’s.

What the horn has is four partials on resonances instead of eight. Above 1,693 hertz its impedance peaks lose their sharpness — the Qs fall from 48 to about 2 within one peak spacing — and a partial above that is not accumulating in a resonance at all. It is being radiated straight out of the bell at whatever rate the lips impose on it.

That collapse is the bell cutoff measured rather than computed from the profile, and it is what the maker’s own cutoff test finds. So the chain is: a low cutoff means few resonant partials, few resonant partials means little spread of rates, and little spread means the partials stay in step further into the note.

The bell decides the blend. That is a claim about orchestration made entirely out of bore acoustics, and it points at the instrument every treatise names as the one that blends with everything.

The claim to be careful about is the reverse one. Nothing here says a saxophone fails to fuse — every one of these instruments is heard as one sound by anybody, at any criterion, which is the fact the whole ladder is trying to model. What the ordering says is which of them the ear has the least margin on, and therefore which would be the first to come apart if anything else went wrong: a second player slightly late, a room reflection at thirty milliseconds, a note begun softly.

The same instrument at a different note

One note per instrument is a thin basis for a claim about families, and the machinery to widen it is already computed: an impedance sweep contains every peak the bore has, and a player choosing a different note is choosing a different peak to sit on.

Up a clarinet, the two clocks disagree. Every impedance peak of a clarinet, with the settling time each implies. The Q rises up the ladder and so does the frequency, and the settling time is their ratio — so in milliseconds the wait falls from 148 at the B2 to 23 at the E♭7, while in periods it rises from 19 to 56. Both curves are monotone and they point opposite ways, which means a player going up the instrument gets notes that arrive sooner and take longer in their own terms. Nothing here is a measurement of an instrument: it is the transmission-line solve of a clarinet-shaped bore, whose peak Qs are sensitive to how finely the sweep is sampled at about five per cent.
Fig. 5 Every impedance peak of the clarinet whose lowest one carries the note above, with the settling time each implies read two ways. The Q rises up the series and the frequency rises faster, so the wait in milliseconds falls monotonically while the wait in periods rises.

That figure is the reason the spread of rates has the shape it does, and it also says what happens when the player goes up. The spread here is the difference between the fundamental’s time constant and the fastest partial’s, and the fundamental’s is by far the larger of the two — so the spread is very nearly the settling time of the note itself, and it is bounded by it exactly, since the spread is one time constant minus another and time constants are positive.

An instrument’s onset spread cannot exceed the settling time of its own lowest partial, which is a number this collection already prints for every wind instrument it has. On the clarinet that number falls from 148 milliseconds at the bottom of the compass to 23 at the top, so the spread falls with it, and a note played high on any of these instruments has partials that stay in step much further into it than the same instrument’s low register.

That gives the ordering a second reading. It is not only a ranking of five instruments: it is a ranking of registers, and the two overlap. A clarinet’s chalumeau and a horn’s middle F are at opposite ends of the census, and a clarinet’s altissimo would sit near the horn — which is a testable claim about when a wind note is easy to hear as one object, and which nothing in this rung has computed because it holds the note fixed to compare the bores.

What the differing rates actually do to a listener

There is something a blown note’s spread of time constants unmistakably does, and it is not a timing effect at all.

A blown attack is brighter than the note it becomes. Each partial of the clarinet's chalumeau D on a clarinet, at four instants of its onset, measured against the level that partial will settle at. Every partial is below its own steady level while it is still building, and the high ones are less far below because their resonances are faster: at 5 milliseconds the top partial is 13.2 decibels above the fundamental relative to steady, at 20 it is 9.7, and by 150 it is 0.9. That tilt is the whole of what a blown instrument's differing Qs do to a listener, and it is a change in the spectrum rather than a difference in when anything starts — which is why it is not an onset asynchrony however large it gets.
Fig. 6 The clarinet’s partials at four instants of its attack, each measured against the level it will settle at. Every partial is below its own steady value and the high ones are less far below, so the note is brighter at its beginning than it will be.

Five milliseconds into the note the fifteenth harmonic is 13.2 decibels above the fundamental relative to their steady levels; at twenty it is 9.7; at fifty, 5.2; by a hundred and fifty, under one. On a saxophone the same tilt is 19.3 decibels at five milliseconds.

A blown attack is brighter than the note it becomes, and that is exactly what the differing Qs produce. It is the ordinary description of a reed attack, it is audible without any equipment, and it is a change in the spectrum rather than a difference in when anything begins.

Which means it loads a different cue. The partials that do not die together built this ladder’s common-fate census on decay times, with a tolerance of a factor of two: partials whose fates differ by more than that are heard as belonging to different objects. Apply the same tolerance to rise rates and every wind instrument fails it — the ratio of largest time constant to smallest is 3.0 on a horn, 5.4 on a clarinet, 5.9 on a trumpet, 6.3 on a trombone and 13.9 on a saxophone.

So the arithmetic that removes one free parameter hands the problem to a cue with a free parameter of its own, and this time it is a tolerance that was set for decays and has never been checked against rises. Three cues, and now two of them have unmeasured constants in them where the sixth rung had one. That is not progress toward a closed model and it is a more accurate picture of the state of the question than the sixth rung’s was.

Partial 4, 7 ms earlyA 10-partial tone on 233 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 0.0 per cent, which is 0.0 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune.28.0 Hz off12345678910partial numberamplitude1% is enoughto hear the partialout of the note;3% and it stopscontributing to thepitch of the wholeMoore, Peters &Glasberg, 1985
Fig. 7 What the blown spread would look like if it were an asynchrony: one partial of a harmonic tone starting seven milliseconds before the rest, which is the saxophone’s own spread at a tenth of steady amplitude. The difference is that on the saxophone the partial has not started earlier — it has got further.

Which computation produced the numbers

The bores are the ones the settling ladder uses: a trumpet, an F horn and a tenor trombone as Bessel flares, a clarinet as a cylinder and an alto saxophone as a cone, each solved with the transmission-line method the bore ladder built, with losses, and swept to 2,600 hertz.

The note is each instrument’s written middle, chosen by frequency rather than by peak index for the reason that ladder gives: a peak-finder that gains or loses a pedal at the bottom of a sweep renumbers everything above it. Its partials are matched to the nearest impedance peak within half a peak spacing, and a peak whose Q has fallen below ten is not counted as a resonance — which is where the bell cutoff shows itself in the data rather than being imposed on it.

The time constant of each partial is that peak’s Q over pi times its frequency, which is the identity the onset ladder’s first rung derives and is exact rather than fitted.

The twenty-millisecond threshold is this ladder’s own from its second rung and is a rounded figure from a literature reporting between about ten and forty.

Where the model stops

One note per instrument. Every number here is the written middle, and both the Qs and the frequencies change up an instrument’s compass — a wind instrument’s own settling time varies by a factor of three to eight across its compass. The ordering of the five is stated at one note and has not been checked at others.

The regime is not a set of independent resonances. A sounding wind instrument is a nonlinear oscillator that has locked its partials into an exactly harmonic series; treating each partial as a linear resonance driven from rest is the standard first approximation to its onset and it is not the physics of a regime establishing itself. It is right about the order of magnitude and about the ordering and it is not a simulation.

And the excitation is not modelled at all. A player’s tongue, breath pressure and embouchure decide the shape of the drive, and a note begun sharply and one begun softly have different onsets for reasons that have nothing to do with any Q. The reed is a valve rather than a switch, and nothing about the valve is in these curves.

What the picture cannot show

It cannot show a bowed string. The sixth rung’s list of exceptions had two entries and this rung has addressed one. A bowed string’s onset is a capture rather than a build, and it has a duration for a completely different reason — the bow reaching a force inside its own window.

Nor can it show two players. Two instruments told to start together are spread by tens of milliseconds for reasons of human timing, and that spread genuinely is an asynchrony. Two players on one note is this collection’s version of that question and it is about spectrum rather than about time.

It cannot show the room. A reflection arriving thirty milliseconds after the direct sound is a second copy of every partial at the same relative levels, which is precisely the scale of the threshold being argued about.

And it cannot show attention. A listener asked to hear a clarinet’s twelfth out of its note can do it, at any criterion, and no cost model has ever explained that.

Whose instruments, and when

The bores are modern orchestral shapes solved at modern dimensions, and none of them is a measurement of an instrument.

The one historical observation available is about the horn and it is not acoustics. The horn’s place as the instrument that joins the woodwind to the brass was settled by practice long before anybody measured a bell’s cutoff — it is in the scoring from the middle of the eighteenth century onward, and the reason given has always been the softness of its tone. This rung offers a different quantity pointing the same way, and the honest status of that is a coincidence worth noticing rather than an explanation: one number out of a bore model agreeing with two centuries of orchestration is a reason to compute more of them, not a reason to believe this one.

Where this ladder goes next

Seven rungs. The ear builds objects and sometimes offers a choice; what makes two partials one note; a spectrum’s inharmonicity read as a perceptual count; the same count under the cue that needs time; the two arbitrated with the exchange rate swept; the onset cue that made the exchange rate stop mattering; and now the blown note, which does not put it back and which introduces a different missing number in its place.

What is owed after this is the criterion, and it is a measurement rather than an arithmetic. Every claim above is a function of one quantity — how far into its rise a partial has to be before the grouping mechanism treats it as having begun — and this collection can compute the consequences of every value it might take without being able to choose one. The experiment that would settle it is stated by the figures themselves: take a spectrum whose partials rise at rates in a known ratio, sweep the ratio, and find where a listener starts hearing the fast partial out. Nothing in this collection can do that. What it can do meanwhile is the compass — the same census at every note each instrument plays rather than at one — which would say whether the ordering of the five families survives being asked about a whole range, and the impedance sweeps for that are already computed and sitting unused.

Part 7 of 8

One essay in the series on auditory scene. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Air columnAttack transientAuditory scene analysisCommon fateFusionOnsetQuality factor