Perception and the listener

Read at two different heights

Ten placements of these figures, one value: the settling criterion is nine tenths in every one of them, and nothing is measured behind it. It is a multiplicative constant only inside the mechanism that has it — a bow's capture and an exciter's contact contain no criterion at all — so moving it rescales one of three clusters against two that stand still. The most-quoted number here, a factor of sixty-nine between the instrument's account and the listener's, is 5.9 at the criterion the listener's own measurements use, and the ordering an earlier essay was written about does not exist below a fifth.

Assumes: A note starts twice · What the tongue actually removes

A note takes a number of periods to speak opened this anchor with an identity and a warning attached to it. The identity is that a resonance driven from rest reaches a fraction pp of its steady amplitude in ln(1p)Q/πf-\ln(1-p)\cdot Q/\pi f seconds, so in periods the wait is the Q and nothing else. The warning is about the fraction:

to is the fraction of the steady amplitude counted as “speaking” and is the asserted part: 0.9 is the usual convention for a rise time and there is no measurement behind it. Every ratio reported in this anchor is independent of it, because it enters as one multiplicative constant.

The first sentence is true. The second is true of one third of the anchor, and the third of it that it is true of is the third the anchor spends least time on.

What the criterion is not in

A resonance building is one of three mechanisms this anchor is made of, and the other two do not contain a criterion at all.

A bowed string is at full amplitude the instant Helmholtz motion begins, and what takes time is the bow reaching a force inside Schelleng’s window — a threshold crossing, with no exponential approach to anything, and no fraction to choose. A plucked or struck string is at full amplitude at t=0t = 0 and its only duration is the exciter’s own contact, which is a contact time and not a criterion either.

So moving the criterion multiplies one cluster of the census and leaves two standing still.

The census with the criterion moved under it. Every instrument's speaking time in milliseconds, against the fraction of the steady amplitude counted as speaking. The criterion is in the wind instruments alone: a resonance takes −ln(1−p)·Q/(πf) to reach a fraction p, so those lines rise across the whole picture, while a bow's capture and an exciter's contact contain no criterion at all and are flat. Every earlier figure sits at 0.9, where the wind instruments are the slowest things in the collection by a factor of 20.5. At 0.05 they are the fastest: a violin's G3 string is the slowest at 15.5 milliseconds and a trumpet takes 1.4. The three clusters cross at a criterion between 0.18 and 0.39, which is inside the range the perceptual measurements work in — their three named criteria are 15 decibels below peak, 6 decibels below peak, and ninety per cent — and the settling figures have only ever used the third of them.
Fig. 1 Every instrument in the census against the criterion. The five wind instruments rise across the whole picture and the seven other rows are horizontal lines.

Every figure this anchor has drawn sits at the right-hand end of that picture, where the wind cluster is above everything else by a factor of four. At the left-hand end it is below everything else. The three clusters pass through one another somewhere in between, and where each crossing happens is arithmetic:

crosses the violin’s G string at
alto saxophone 0.20
clarinet 0.21
tenor trombone 0.28
F horn 0.39
trumpet 0.44

Below a criterion of a fifth, every wind instrument in this collection speaks faster than every string a bow can catch. Above two fifths, every one of them is slower. The band between is where the census’s own taxonomy stops sorting the instruments and starts interleaving them.

Three ways a note starts, on one pair of axes. Every instrument in this collection, with how long its note takes to speak measured twice: across, in periods of the note itself; up, in milliseconds. The three clusters are three different pieces of physics. A wind instrument accumulates energy in a resonance, so its wait is the resonance's Q — 2, 2, 2, 2, 3 periods here. A bowed string is at full amplitude the moment Helmholtz motion begins and what takes time is the bow reaching a force inside Schelleng's window, which is 3.0, 3.8, 5.1, 7.6 periods. A plucked or struck string is at full amplitude at once and the only duration in it is the exciter's own contact, under a tenth of a period. The two axes do not agree: the slowest instrument in milliseconds is a violin's G3 string at 15, and in periods it is a violin's E5 string at 8.
Fig. 2 The census at a criterion of about a sixth rather than nine tenths. The three clusters are still three clusters and they have changed places: the bowed strings are now the slow ones.

Nothing about the picture’s structure changes. The three mechanisms are still three mechanisms, they are still separated on both axes, and the argument that they are three different quantities reported in one unit survives intact. What changes is which of them is at the top, and that is what every sentence of this anchor’s prose has been about.

The statistic that is not even monotone

The third rung reports a number for how far the two readings disagree — Spearman’s footrule over the two orderings, normalised so that a complete reversal is one — and the value it quotes is 0.250 at the criterion it was computed at.

Sweeping the criterion gives 0.250 at a twentieth, 0.250 at a tenth, 0.361 at a sixth, 0.500 at three tenths, 0.389 at two fifths, 0.278 at a half, and 0.250 from two thirds all the way up to 0.99.

That is not a number drifting with a convention. It is a number that returns to exactly the same value at both ends and doubles in the middle, which means a reader cannot even tell from it which direction the convention should move to make the disagreement smaller. A statistic reported to three decimals that is non-monotone in an unmeasured parameter is a statistic that needs the parameter printed beside it.

The same instruments, ordered twice. On the left, slowest first in periods of the note being played; on the right, slowest first in milliseconds. 9 of 12 instruments change place, and the crossing lines are the whole argument: the number of periods a resonance takes to settle is its Q and has no pitch in it, while the number of milliseconds is that count divided by the frequency. A clarinet's low E is 2 periods and 13 milliseconds; a saxophone's written middle is 3 periods and 14. Everything about ensemble timing is in milliseconds and everything about how much of a transient a listener hears as part of the note is in periods, so the two orderings are both wanted and neither is the answer.
Fig. 3 The two orderings side by side at a criterion of a sixth. The wind instruments have moved from the top of both columns to the bottom of one.

The reason for the shape is worth having. At the high end the wind cluster is far above the others in milliseconds and far above them in periods too, so both orderings put the same five instruments on top and the footrule counts only the shuffling inside the clusters. At the low end it is far below in both, and the same thing happens with the order reversed. In between the clusters overlap, the two axes disagree about how they overlap, and the footrule peaks. The statistic is measuring cluster separation and not the thing it is named for, and that only becomes visible when the axis it depends on is moved.

The same instruments, ordered twice. On the left, slowest first in periods of the note being played; on the right, slowest first in milliseconds. 9 of 12 instruments change place, and the crossing lines are the whole argument: the number of periods a resonance takes to settle is its Q and has no pitch in it, while the number of milliseconds is that count divided by the frequency. A clarinet's low E is 19 periods and 148 milliseconds; a saxophone's written middle is 37 periods and 159. Everything about ensemble timing is in milliseconds and everything about how much of a transient a listener hears as part of the note is in periods, so the two orderings are both wanted and neither is the answer.
Fig. 4 The same picture at the criterion used all along. The five wind instruments now occupy the top of both columns and the crossing lines are all inside the clusters.

The factor of sixty-nine

The anchor’s most-quoted result is the fourth rung’s: the instrument’s settling time set against the listener’s placing delay on the five instruments both accounts hold, with the ratio between them spanning a factor of sixty-nine and sorting perfectly by mechanism.

The factor of sixty-nine, against the convention it was measured at. An earlier essay set the instrument's own settling time against the listener's placing delay on the five instruments both accounts hold, and reported that the ratio between them spans a factor of sixty-nine. This is that number against the settling criterion. At 0.9 it is 69.0, which is the published figure. At 0.178 — the fraction of the peak the perceptual detection criterion sits at — it is 5.9. At 0.05 it is 3.1. The listener's side of the comparison has a criterion too, and moving it cannot appear on this picture at all: it multiplies every one of the five ratios by one common factor and leaves their spread exactly where it was. The spread is a property of the settling criterion alone, and that criterion has no measurement behind it.
Fig. 5 That factor of sixty-nine, against the settling criterion it was measured at. It is 69.1 at nine tenths, 20.9 at a half, 5.9 at a sixth and 3.1 at a twentieth.

Sixty-nine at nine tenths. Five point nine at a sixth, and three point one at a twentieth. The number spans a factor of twenty-two across a parameter with nothing measured behind it, and it does so smoothly and monotonically, so unlike the footrule it at least says which way it is going.

One thing about that curve is not obvious and is the load-bearing part. The listener’s side of the comparison has a criterion too, and moving it cannot appear on this picture at all. A note is heard after it starts reads each instrument’s measured envelope against a level criterion, and its lag is that criterion’s own logarithm times a rise time — so changing it multiplies every one of the five lags by one common factor. A common factor leaves the ratios’ spread exactly where it was, and it leaves the ordering exactly where it was too.

So the spread is a property of the settling criterion alone. It cannot be repaired by choosing a better criterion on the listener’s side, and it cannot be blamed on one.

Which is the mismatch this rung was written to find

The two ladders have criteria and they are not the same criterion.

The perceptual side carries three, each with a source: fifteen decibels below peak for detection, six decibels below peak for the onset a listener places a note at, and ninety per cent of peak for a perceptual attack time. As fractions of the peak those are 0.178, 0.501 and 0.900.

The settling side carries one, and it is 0.900.

The fourth rung paired the instrument’s time to reach nine tenths of its steady amplitude against the listener’s time to reach half of its peak. Two readings of the same rising curve, taken at two different heights, and the factor between them was reported as a finding about instruments.

Read both at the same height and the disagreement collapses in a specific way. At the perceptual ladder’s detection criterion the five ratios are 6.63, 6.37, 2.18, 1.79 and 1.13 — a spread of 5.9 rather than 69, and every one of them above one, meaning the listener is later than the instrument on all five rather than later on three and earlier on two. At the highest of the three criteria they are 6.63, 6.37, 2.18, 0.15 and 0.10, and the two wind instruments have gone to the other side of the diagonal.

And the ordering, which is the part that vanishes

How many of the five change place, against the convention. How many of the five instruments occupy a different rank in the two orderings — the instrument's settling time and the listener's placing delay — as the settling criterion moves. Below about a fifth of the steady amplitude, 0 of the five change place: the two accounts agree completely, and the bowed violin is the slowest on both. Above about two fifths, 3 change place, which is the disagreement that earlier essay was written about. A bowed string reaches Helmholtz motion early and then goes on growing in level; a wind instrument is quiet the whole time it builds. So a low criterion catches the violin late and a high one lets the winds overtake it, and which of those is the right account depends on where a listener's own threshold sits — which is a question for the perceptual account and not for this one.
Fig. 6 How many of the five occupy a different rank in the two orderings. Below a fifth it is none of them; above two fifths it is three.

Below a criterion of about a fifth, none of the five instruments changes place between the instrument’s ordering and the listener’s. The two accounts agree completely, and the bowed violin is the slowest on both. The clarinet passes the violin at 0.184 and the trumpet at 0.383, and above that three of the five have moved.

The fourth rung’s headline — the violin is third slowest to settle and the last to be heard — is a statement about a criterion of nine tenths. At the criterion the perceptual ladder’s own detection threshold sits at, the violin is the slowest to settle and the last to be heard, and there is no disagreement to write about.

That is not a demolition of the rung. It is the rung’s own claim, made precise: the two accounts differ about the ordering only when the physical side is read near the top of the envelope, and there is a reason for that with real content in it. A bowed string reaches Helmholtz motion early and then keeps growing in level, because the amplitude is set by the bow speed and the bow is still accelerating; the fourth rung says so itself. A wind instrument is quiet the whole time it builds. So a criterion low on the envelope catches the violin genuinely late, and a criterion high on it lets the winds overtake — and the two accounts are not disagreeing about instruments, they are disagreeing about where on a rising curve a note counts as having started.

A second ladder found the same parameter and said so

This is not the only place in the collection where the criterion turned out to be carrying an argument, and the other case is worth setting beside it because it was found from a completely different direction.

A blown note does not start late asks whether a wind instrument’s partials arrive together enough to be heard as one note. Its answer depends on the same choice: a clarinet’s eight partials are spread over 5.5 milliseconds at a tenth of the steady amplitude, 36 at a half and 121 at nine tenths, against a twenty-millisecond threshold for hearing a partial out of its note. So the same instrument fuses easily, marginally or not at all according to a number nobody measured, and that essay ends by naming the criterion as the quantity it is now missing.

Its arithmetic is the cleaner statement of why. The spread at a criterion pp is ln(1p)-\ln(1-p) times the difference between the largest and smallest time constant — one quantity, the spread of rates, multiplied by a factor that grows without limit as pp approaches one. The multiplier is 0.105 at a tenth and 2.30 at nine tenths, twenty-two times larger.

Twenty-two is the same factor this rung found in the spread between the two accounts, and it is the same factor for the same reason: both quantities are a difference of time constants, and the criterion is a common multiplier on all of them. Two rungs on two ladders, asking unrelated questions, arriving at one convention and one number. That is the strongest evidence available that the constant is doing real work rather than sitting inertly in a formula, and it is why the right response is a rule about quoting rather than a new default.

Which of the three is right

The honest answer is that they measure different events and the collection needs all three.

Ninety per cent is the right convention for a rise time, which is an engineering quantity about when a system has settled, and it is what an instrument maker or a bore calculation wants. Nothing above argues that it is wrong; the first rung chose it correctly for what it was doing.

Six and fifteen decibels below peak are thresholds for hearing, and they are the right criteria for anything about when a note arrives to somebody. What the tongue actually removes is a rung about a player’s control over onset and it inherits the same problem, though more gently: the ramp’s cost is about half its own length at every criterion, so the milliseconds a tongue is worth are nearly criterion-independent even though the wait it is subtracted from is not.

What follows is a rule about which number to quote rather than about which convention to adopt. Any claim in this anchor that compares an instrument to a listener must state the criterion on both sides, and any claim that compares the build mechanism to either of the other two must state it at all. Everything internal to the build mechanism — the ordering of the five wind instruments, the ratio between periods and milliseconds, the compass sweeps — is genuinely independent of it, because there the constant does divide out.

Up a clarinet, the two clocks disagree. Every impedance peak of a clarinet, with the settling time each implies. The Q rises up the ladder and so does the frequency, and the settling time is their ratio — so in milliseconds the wait falls from 148 at the B2 to 23 at the E♭7, while in periods it rises from 19 to 56. Both curves are monotone and they point opposite ways, which means a player going up the instrument gets notes that arrive sooner and take longer in their own terms. Nothing here is a measurement of an instrument: it is the transmission-line solve of a clarinet-shaped bore, whose peak Qs are sensitive to how finely the sweep is sampled at about five per cent.
Fig. 7 A clarinet’s own peaks, which is a comparison entirely inside one mechanism. Here the criterion is a scale on the vertical axis and nothing on this page applies.

Which computation produced the numbers

The settling time is ln(1p)Q/(πf)-\ln(1-p)\,Q/(\pi f), so a census computed at one criterion is a census at every criterion under one multiplication: the ratio of ln(1/(1p))\ln(1/(1-p)) at the two values. That is what the sweep does, rather than re-solving five bores at each point, and the figures assert the rescaled answer against a recomputed one at a criterion in the middle of the range rather than trusting the shortcut.

The capture and impulse rows carry no criterion, so they are copied across unchanged, and that is exactly the property this rung is about.

The perceptual side is pCentreTable’s, unchanged. Its three criteria are the ones the perceptual ladder names, at the fractions of the peak given above.

The footrule is the third rung’s own: the summed rank displacement between the two orderings, over the largest value it can take.

Where the model stops

Three criteria is not a measurement of a criterion. The perceptual ladder’s three come from two published sources and they are what a listener was asked to do in two experiments, not a single number a listener has. The right criterion for a given musical question is somewhere among them and this collection cannot narrow it.

The envelope shapes are not the same object on both sides. The settling side computes an exponential approach to a steady amplitude from a Q; the perceptual side reads a measured ten-to-ninety rise time and assumes an exponential with that rise. Both are exponentials and the two are fitted to different things, so reading them at the same fraction is closer to a like-for-like comparison than reading them at different ones and is not the same as measuring both.

And the criterion is not the only unmeasured constant on this page. The bow’s force ramp, which the rung before this one turned into an argument about articulation, has never been measured either, and it scales the whole capture cluster the way the criterion scales the build cluster. The difference is that the capture cluster is four rows and the build cluster is five, and nobody has quoted a headline off the capture cluster yet.

What the picture cannot show

It cannot show a listener choosing. Everything above is about which convention makes two computations agree, and whether a real listener placing a real note behaves like a fixed fraction of the peak at all is the question the perceptual ladder’s own third rung is about and this one inherits.

It cannot show what the criterion does to the anchor’s other rungs. The higher note speaks sooner and takes longer compares peaks of one instrument and is safe; playing louder is playing earlier is on the perceptual side and inherits the perceptual criterion instead. Auditing which of this collection’s claims sit inside one mechanism and which cross between two is a piece of work this rung has done for five instruments and not for the collection.

Nor whether any of these differences is audible. A factor of sixty-nine between two accounts and a factor of six between the same two accounts read consistently are both large, and neither is a claim that anybody can hear the difference between two instruments arriving at their ninth tenth in different orders.

Where this ladder goes next

Six rungs. A wait is a Q in one unit and a Q over a frequency in another; up a brass instrument the two units disagree monotonically; on a bowed string the wait is a capture and the hardest place is the latest; the instrument’s number set against the listener’s; the articulation, which removes a ramp rather than adding an impulse; and now the criterion under all five, which turns out to carry the fourth rung’s headline entirely and to be a different number on each side of the comparison it was measured across.

What the anchor owes now is a listener, and it is the first debt on this ladder that arithmetic cannot pay. Everything above is a computation checked against another computation, and the question that decides between three criteria is what a person does when asked to place a note — which is an experiment with a stimulus, a task and a group of listeners, and this collection has none of those. What it can do meanwhile is smaller and worth doing: the perceptual ladder’s own lags are derived from published ten-to-ninety rise times for nine instruments, and the settling ladder computes rise times from first principles for five bores. Those two sets overlap on three instruments and have never been compared as rise times rather than as lags, which would say whether the two accounts even describe the same envelope before anybody argues about where to read it.

Part 6 of 6

One essay in the series on onset time. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientEnsemble asynchronyEnvelopeHelmholtz motionOnsetPerceptual-centreQuality factorTransient