Perception and the listener

Playing louder is playing earlier

An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.

Assumes: The players who have to be early · A hammer is not an impulse

The first two rungs of this ladder held one thing fixed without saying so. Every attack time in them is a single number per instrument — a bowed violin is ninety milliseconds, a piano is eight — and every figure was drawn for a note at one dynamic.

An instrument does not have one attack time. It has a range of three or four to one, and the thing that moves through that range is the dynamic.

How much earlier an accent is heard, by mechanismAn accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller.24.0 ms05101520051015202530decibels louder than the notes around itmilliseconds earlier the accent is heardcriterion at the note's own peakon an unchanging envelope: nothingthe same criterion, on theshorter rise a loud note hascriterion at a fixed levelboth, which is what aninstrument actually does
Fig. 1 An accented note on an instrument with a ninety-millisecond attack, against how much louder the accent is, with three mechanisms drawn apart. The flat line is a criterion tied to the note’s own peak on an envelope that does not change shape: it cannot move, by construction. The lower curve is the same criterion on the shorter rise a harder-driven instrument produces. The dashed curve is a criterion at a fixed level, crossed sooner by a bigger rise. The top curve is both together, which is what an instrument actually does — and at a twelve-decibel accent it is twenty-four milliseconds.

Twenty-four milliseconds is not a subtlety in this collection. It is very nearly the whole heard-moment spread of a piano trio, most of the way to the twenty milliseconds at which two events stop being one, and it happens because the player played louder.

The first mechanism is already on this site

The attack shortening with force is not an assumption. For one instrument it is computed here, and has been since the excitation ladder’s second rung.

A dynamic mark is an instruction about the spectrum. Six dynamic markings, given a hammer velocity each in a stated sequence of factors of two, with what the string then does. The level rises 35.1 decibels from pp to ff, which is the part everybody means. The contact time falls from 2.26 to 0.95 milliseconds, so the first null of the hammer's own pulse moves from partial 2.5 to partial 6.0 and the spectral centroid rises by 56 per cent. The partials between those two nulls are not quieter at pp; they are not there.
Fig. 2 A hammer is not an impulse: the felt is a nonlinear spring whose stiffness rises with compression, so a harder blow makes the contact shorter. Each dynamic marking is drawn with the contact time it produces and the partial at which the excitation pulse first has a null. The contact time goes as the force to the power −0.25, which is a factor of 1.7 across the marked range — and a contact time is an attack time.

The exponent is worth a sentence about where it comes from, because it is the one number in this essay that was derived rather than asserted. Felt compresses nonlinearly, so its effective stiffness rises with how hard it is squeezed; a stiffer spring against a fixed mass has a shorter half-period; and integrating that gives a contact time falling as the force to the power −0.25. Nothing about the piano was fitted to produce it and the brightness that comes with it is the same computation read as a spectrum.

A piano’s contact falls from about 2.3 milliseconds at a soft blow to 1.3 at a hard one. That instrument’s whole attack is small enough that the consequence for its heard moment is under a millisecond, so the piano is the case where this rung does almost nothing.

The instruments where it does a great deal are the slow ones, and there the exponent is asserted rather than computed. A bow driven harder reaches Helmholtz motion sooner — the site’s own attack model says so — but by how much, as a power of the driving force, is not a number this collection has measured. The figures use 0.15 and say so; at 0.25 the lower curve would be two-thirds higher and at zero it would be flat.

A constant bow force cannot start a note. Schelleng's minimum and maximum bow force, evaluated at the speed the bow has after one period rather than at the speed the note will be sustained at. Both bounds are proportional to bow speed and the speed after one period is the acceleration divided by the frequency, so both are proportional to acceleration and the region is a wedge through the origin. Its width as a ratio is 9.0 to 1 at every acceleration — the attack is not a narrower window than the sustain, it is the same window somewhere else. A fixed force, drawn here at 0.02, is inside it at one acceleration, about 0.18, which is why the force has to rise with the speed and why beginning a note is a trajectory through this plane rather than a point in it.
Fig. 3 The bowed attack the exponent is about: how many periods of scraping precede Helmholtz motion, against how hard the bow is pressed and how fast it accelerates. Harder and faster is fewer periods, which is a shorter attack — the direction is secure and comes out of the same model the note has to start somewhere is built on. What is not secure is the exponent, and the hero figure’s middle curve is the only quantity in this essay that depends on it.

The second mechanism is the ladder’s own fork

The first rung recorded that two families of experiment give two criteria for the heard moment and left the choice open. The choice has been irrelevant so far, because every claim was a difference between two instruments and a difference is the same slope times two attack times either way.

It stops being irrelevant here, because the two criteria say different things about the same note played at two dynamics.

Under one model an accent arrives early, under the other it does notThe heard moment of a note with a 90 millisecond attack, against how much louder it is than the notes around it. The flat line is a criterion taken as a fraction of the note's own peak: the criterion rises with the note, so the crossing does not move at all. The falling line is a criterion taken as a fixed level: a louder note crosses the same line sooner, so an accent of 12 decibels is heard 23 ms early with no change in when the key was pressed. Nothing but an accent distinguishes the two models, which is why this figure is the experiment rather than the illustration.23 ms early05101520051015202530decibels louder than the surrounding notesmilliseconds from onset to the heard momentcriterion as a fractionof this note's own peakcriterion as a fixed levelset by everything else
Fig. 4 The fork. A criterion taken as a fraction of a note’s own peak follows the note up and the crossing never moves — the flat line. A criterion taken as a fixed level, set by whatever else the listener is hearing, is crossed sooner by a louder note: twenty-three milliseconds sooner at twelve decibels on this attack. The two models agree about every ensemble computed earlier and disagree entirely about an accent.

This is the same structure as the fork the loudness ladder found in its own first rung, and it resolves the same way: the two models agree wherever the quantity being compared is a ratio at one level, and part company as soon as the level itself is the variable. The quietest thing audible is a criterion of exactly this kind — an absolute level below which nothing is heard — and it is the one criterion in this collection that is unambiguously not relative to anything.

The second is the more plausible of the two on its face. A listener is not measuring each note against its own eventual maximum, because the eventual maximum has not happened yet when the heard moment arrives; whatever the criterion is, it has to be available at the time. A level set by the preceding music is available; the note’s own peak is not.

But the first is what the published criteria are stated as — 6 decibels below maximum, 90 per cent of peak — and those statements come from experiments in which every stimulus was at the same level, where the two are indistinguishable.

So the model that is easier to state is the one that could not be causal, and the experiments that produced it could not have told.

What this predicts that could be checked

The two mechanisms have different signatures, and this is the useful part.

How much earlier an accent is heard, by mechanismAn accented note on an instrument with a 8 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 0.7 milliseconds at 12 decibels. A criterion at a fixed level gives 2.0. Both together give 2.2, and the rise at that dynamic is 6 milliseconds rather than 8. The rise-shortening exponent is stipulated at 0.25 rather than measured, and the two upper curves would separate further if it were smaller.2.2 ms0510152000.511.522.5decibels louder than the notes around itmilliseconds earlier the accent is heardcriterion at the note's own peakon an unchanging envelope: nothingthe same criterion, on theshorter rise a loud note hascriterion at a fixed levelboth, which is what aninstrument actually does
Fig. 5 The same figure for a piano, whose attack is eight milliseconds and whose rise-shortening exponent is the measured 0.25 rather than a stipulated 0.15. Every curve is the same shape and the whole vertical scale is a tenth of the previous figure’s. An accent on a piano moves its own heard moment by two milliseconds, which nothing can measure in a performance.

The prediction is that the size of the effect is proportional to the instrument’s attack time, and the attack times run from three milliseconds to a hundred and ten. So the same experiment run on a marimba and on a singer should give answers that differ by a factor of thirty, and either model predicts that.

What separates them is the ratio rather than the size. At twelve decibels on a ninety-millisecond attack, a relative criterion with a shortening rise gives 5 milliseconds and a fixed criterion gives 23 — a factor of four and a half. Both are large enough to measure on a slow instrument and neither is measurable on a fast one, so the experiment has to be done on a voice or a bow.

Every criterion makes the lag proportional to the rise. The heard moment against the attack time, for the three criteria the literature supports. All three are straight lines through the origin, because each is a fixed fraction of the same envelope — so the criterion decides the slope and nothing else. At a 180 millisecond attack the three give 16 ms, 57 ms, 189 ms, a spread of a factor of 12. Every claim here is a difference between two instruments, and a difference is the same multiple of the same slope whichever line is taken.
Fig. 6 The reason the experiment has to be done on a slow instrument: every criterion is a straight line through the origin, so at a short attack all of them are close together and any discrimination between them is a few milliseconds wide. The instruments where the models separate are the ones at the right of this axis.

The accent and the beat, which are now the same cue twice

There is a consequence for the rhythm ladders that is worse than a correction.

The beat is inferred from a list of onsets and their accents, and every model of that inference on this site treats the two as independent evidence: a note is at a position, and separately it is loud. Syncopation is a number computed from where the loud notes fall against where the metre says they should.

If a loud note is also heard early, the two are not independent. A model given the heard moments of a passage sees loud notes systematically ahead of the grid, on any instrument slow enough for it to matter — and a phase estimate fitted to those onsets is pulled early by exactly the amount the accents were. On a bowed line with twelve-decibel accents that is twenty-four milliseconds of bias in the inferred beat, which is about four per cent of a half-second beat and is the same order as the deviations the microtiming ladder is made of.

The same measurement, read from onsets and read from heard moments. Three published timing deviations, each drawn twice: the open mark is what an onset detector measured and the filled mark is where the note is heard, given what the two instruments are. The correction has a sign. A slow-attack soloist against a fast-attack timekeeper is heard further behind than the measurement says — 0 milliseconds becomes 28 — and two players on the same instrument get no correction at all, which is why a string section's internal timing needs none of this.
Fig. 7 A published timing deviation drawn twice: the open mark is what an onset detector measured and the filled mark is where the note is heard, given what the two instruments are. The correction has a sign.

A measured deviation is not a perceptual one, and the gap between them is the size of the effects the microtiming literature reports. Twenty-four milliseconds — the accent shift this essay computes — is the same order as the correction an instrument pairing introduces before any player does anything, which is why the two have to be separated before either can be interpreted.

None of the models on this site is wrong as a result, because all of them were run on quantised or on onset-measured material rather than on heard moments. What changes is what they are models of.

Which computation produced the numbers

Three lines, each a call to the same envelope inverse.

The unaccented lag is the criterion fraction inverted through a rise of 90 milliseconds: 28.5 milliseconds at the perceptual-onset criterion.

The rise-only shift replaces the rise with rise × amplitude^−0.15 and inverts the same fraction: an amplitude ratio of 3.98 — which is twelve decibels — gives a rise of 73 milliseconds and a lag of 23.2, so the shift is 5.3.

The level-only shift keeps the rise at 90 milliseconds and divides the criterion by the amplitude ratio, so the note has to reach 12.6 per cent of its peak rather than 50.1: the lag is 5.5 milliseconds and the shift is 23.0.

Both together give 24.0, which is very slightly less than the sum, because a shorter rise crossing a lower criterion reaches it a little later than the two effects computed separately would suggest. That the two are nearly additive is a property of the exponential and not a general one.

The criterion’s time constant, which turns out not to be a problem

The caveat below says the criterion has to be updated, that an instant update would remove the effect and a frozen one make it maximal, and that the truth — a second or two — is “in the worst possible place for a simple model”. Modelling it settles that, and it settles it the other way.

Let the criterion relax from the background level toward the note’s own with time constant τ. At τ = ∞ it is frozen and the fixed-level model applies; at τ = 0 it tracks the note and the relative model applies. Both limits come out exactly right, which is the check. Between them:

τ accent lag shift
instant (relative) 28.5 ms 0.0
20 ms 17.1 11.4
50 ms 8.3 20.2
200 ms 6.0 22.5
1 s 5.6 22.9
frozen 5.5 23.0

The transition is over by two hundred milliseconds, and half of it has happened by twenty. At the published time constant of one to two seconds the effect is at a hundred per cent of the frozen-criterion value, to the precision of the table.

So a lag of a second or two is not in an awkward place; it is twenty times past the point where the criterion is effectively frozen for the duration of an attack. The genuinely awkward region is τ around twenty milliseconds — a fifth of the rise time — and nothing in the loudness literature puts it anywhere near there. The one parameter the essay flags as a threat to the argument turns out to be the one that cannot threaten it, because the quantity it has to be compared with is not the note or the phrase but the ninety-millisecond rise.

And a loudness criterion costs much less than a third

The other caveat estimates that a criterion on loudness rather than amplitude would give shifts three times smaller. It gives shifts a quarter smaller.

Loudness in sones goes as amplitude to about 0.6, so a twelve-decibel accent is an effective ratio of 2.29 rather than 3.98 — seven decibels rather than twelve. The shift falls from 23.0 milliseconds to 18.4. Even a cube-root compression, which is more aggressive than any published loudness exponent, gives 12.9 rather than the 7.7 a factor of three would imply.

The reason is that the shift saturates. The lag is a logarithm of the criterion over the amplitude, so as the accent grows the lag runs down toward zero and the shift runs up toward the unaccented lag of 28.5 milliseconds and stops. A twelve-decibel accent is already eighty-one per cent of the way to that ceiling, and compressing the level scale moves a quantity that is nearly saturated very little.

That makes the conclusion considerably safer than the caveat suggested. Any monotone compression of the level scale leaves the effect within a factor of two, because the ceiling is set by the attack time rather than by the dynamic.

How much earlier an accent is heard, by mechanismAn accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller.24.0 ms05101520051015202530decibels louder than the notes around itmilliseconds earlier the accent is heardcriterion at the note's own peakon an unchanging envelope: nothingthe same criterion, on theshorter rise a loud note hascriterion at a fixed levelboth, which is what aninstrument actually does
Fig. 8 An accented note on an instrument with a ninety-millisecond attack, against how many decibels louder it is, with the three ways it can arrive early drawn separately: a shorter rise, a fixed-level criterion, and the two together.

The three are not the same size and they do not add to the same thing. A criterion tied to the note’s own peak on an unchanging envelope gives exactly nothing — so the whole effect depends on either the rise shortening or the criterion being absolute, and the essay’s number is a sum over two mechanisms rather than a measurement of one.

Where the model stops

The exponent is asserted for every instrument except the piano. It is the same weakness this collection has recorded twice already — a direction that is secure and a size that is not — and here the conclusion is not sensitive to it: the level mechanism dominates the rise mechanism at every exponent between 0 and 0.25, so the ordering of the curves is safe and only the gap between the lower two is at risk.

A crescendo is not an accent. All of this is about a note louder than its neighbours. A note in a passage that is uniformly loud has neighbours that are loud too, so a criterion set by the surrounding music moves with it and nothing happens at all. The effect is about contrast, which means it is largest at exactly the moments a composer marks — a sforzando into a quiet bar — and absent through a fortissimo tutti.

The criterion’s own time constant is now modelled, and the section above finds it harmless. An instant update removes the effect and a frozen criterion maximises it, as recorded — but the transition between them is complete by two hundred milliseconds, and the published lag of a second or two sits far past it. The quantity the time constant has to be compared with is the attack, not the note and not the phrase, and it is more than ten times longer than the longest attack here.

Whose music, and the notation that has been carrying this

The claim is about listeners, so it applies everywhere. Where it becomes a claim about a repertoire is in what has been written down.

European notation has two separate marks for what this rung says is partly one thing. A > is an instruction about loudness and an agogic accent — a note held slightly long, or arrived at slightly late — is an instruction about timing, and performers are taught them as different expressive resources. The arithmetic here says that executing the first produces some of the second for free, on any instrument with a slow attack, and none of it on a marimba.

That is the same shape of finding as the mark that is not a level, which found that a dynamic marking is not an instruction about level but about effort, and that the timbre comes along with it. This is the third thing that comes along with it: the timing.

And it explains, without needing a theory of expression, why performance research keeps finding that accented notes are played fractionally late. If an accent is heard 24 milliseconds early, a player who wants it on the beat has to put it down 24 milliseconds late — and a measurement made from onsets records a late accent and calls it expressive lengthening.

What the picture cannot show

Whether a listener’s criterion is a level at all. The two models here are both criteria on an amplitude envelope, which is the simplest thing that could be true. A criterion on loudness rather than amplitude was recorded here as giving shifts three times smaller; computed, it gives 18.4 milliseconds against 23.0 — a quarter smaller, because the shift saturates at the unaccented lag and a twelve-decibel accent is already most of the way there. A criterion on rate of change rather than on level would behave differently again, and that one is still not computed.

The loudness the criterion is set by is not the loudness of one note. A chord is not as loud as its notes, and what sets the reference for an accent in a texture is the whole texture — so an accent in a thin passage has further to rise above its neighbours than the same accent in a tutti, and is heard earlier for it. That is a second-order effect this model does not carry and the loudness ladder now has the machinery for.

And nothing here is measured. Every number is a model evaluated at published parameters. The one measurement that would settle it — the same instrument, the same notated rhythm, two dynamics, listeners asked to say which note is early — is a listening experiment and this collection has none.

What the two computations above buy is not accuracy but robustness, and it is worth being clear that those are different. Neither settles what a listener does; both establish that the answer is insensitive to a parameter the essay had flagged as a risk. A model whose conclusion survives its time constant varying by two orders of magnitude and its level scale being compressed to a cube root is a model whose remaining uncertainty is somewhere else — in whether the criterion is a level at all, and in the bowed exponent, which is the one number still asserted.

Where this ladder goes next

Three rungs. A note is heard after it starts, by a fraction of its own attack; an ensemble that mixes attack families carries a spread it cannot play its way out of; and the attack is not a constant of the instrument but a thing the dynamic moves, so an accent arrives early.

What the ladder still owes is the thing all three rungs have held at one value: the pitch. Every envelope here rises as fast as its instrument allows, and a low note cannot rise faster than its own period — a note at 40 hertz has 25 milliseconds in every cycle and cannot establish an amplitude inside one. So there is a floor on the attack time that rises as the pitch falls, it is computable from the period and the number of cycles a resolvable amplitude needs, and it would predict that a double bass is heard later than a violin playing the same written rhythm even at the same bow speed. That is arithmetic this site already has, in the ladder about how long a note has to be before it has a pitch at all.

Part 3 of 9

One essay in the series on Perceptual-centre. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientContact timeDynamicsMetrical weightOnsetPerceptual-centreTiming deviationTransient