Rhythm and metre

A note is heard after it starts

Every rhythm essay until now has treated a note's onset as the moment it happens. It is not: the instant a listener aligns a note with a beat is later than its physical start by an amount the note's own attack decides, and for a sung or bowed note that amount is about thirty milliseconds — the size of the whole quantity six essays on microtiming set out to measure.

Assumes: The first fifty milliseconds · The beat is inferred, and sometimes wrongly

A note starts at a definite instant. The first sample that is not silence is a fact about a recording, it can be found automatically, and every rhythm essay in this collection has used it: the beat is inferred from a list of onsets, a groove is a list of departures from a grid, two players correcting toward each other are two lists of onsets compared.

The instant a listener hears the note is a different instant, and it is later.

One envelope, and the three places a listener might be said to hear itThe amplitude envelope of a note with a 90 millisecond exponential attack, with the three criteria the literature offers drawn across it. The heard moment is 8.0 ms at the detection criterion, 28 ms at the perceptual-onset criterion and 94 ms at the perceptual-attack criterion. The physical onset is at zero on this axis and no criterion puts the heard moment there. The buttons play this attack against a two-millisecond one, started at the same instant.8.0 ms28 ms94 ms05010015020000.20.40.60.81milliseconds after the physical onsetamplitude, as a fraction of the note's own peak15 dB below peak6 dB below peak90% of peak90 ms attack
Fig. 1 The amplitude envelope of a note with a ninety-millisecond attack — a bowed string, an average one — with the three criteria the literature offers for where the heard moment is. The physical onset is at zero and none of the three puts the heard moment there. The rightmost is Gordon’s perceptual attack time, measured on listeners aligning instrument tones with a beat; the two on the left are Vos and Rasch’s perceptual onset, measured on listeners saying when a tone begins. The buttons play this note against a two-millisecond attack, started at the same instant, and then with the fast one delayed to match.

The gap is not small. On the middle criterion it is 28 milliseconds, which is very nearly the whole of the thirty-millisecond lag a jazz soloist sits behind the ride cymbal — a deviation that ladder treats as an expressive choice, measured from onsets.

Two notes, one onset, two moments

The claim is easiest to see with two notes rather than one, because the absolute lag depends on a criterion nobody agrees about and the difference between two lags does not.

Two notes started together, heard 32 ms apart. Two amplitude envelopes rising from the same instant: a piano with an 8 millisecond attack and a sung vowel with 110. The horizontal line is the criterion — 6 dB below peak, from Vos & Rasch 1981 — and the two dots are where each envelope crosses it. Nothing about the onsets differs; the heard moments differ by 32 ms, which is why the sung vowel has to start early to be heard on the beat. The buttons play the pair as written and then with the piano delayed by that amount.
Fig. 2 A piano note and a sung note, told to start at the same instant, drawn as amplitude envelopes. The piano is at full amplitude before the voice has reached a tenth of its own; the criterion line crosses the two envelopes 32 milliseconds apart. The first button plays them as written and they do not sound simultaneous; the second delays the piano by that 32 milliseconds and they do.

Started together, they are not heard together. The piano is heard 32 milliseconds before the voice, which is more than the twenty milliseconds at which two events stop being one and become a note and a late note.

This is not a defect of anybody’s playing. It is arithmetic on the two envelopes, and it says that an accompanist who wants to sound simultaneous with a singer has to play about a thirtieth of a second late.

What the envelope is, and why it is the whole of the mechanism

The site already has the object this runs on. The shape of a note is an attack, a decay, a sustain and a release, and the first of the four is the one under discussion.

A marimba’s attack is three milliseconds, a trumpet’s thirty, a bowed violin’s ninety, and the whole quantity this essay is about lives inside the first twentieth of a two-second note — which is why every other figure here is drawn at a quarter of a second. The reason so small a corner matters is that it is not the same corner for everyone playing at once.

Which ensembles have this problem and which do not. The width of the heard-moment spread built into 5 standard instrumentations, before any player does anything. An ensemble drawn from one attack family has a spread of zero — every note is heard the same distance after it is started, so a common onset is a common heard moment, and this is true of a string quartet and of a gamelan for the same reason and in the same amount. A mixed ensemble carries between 19 and 33 milliseconds of it.
Fig. 3 The width of the heard-moment spread built into five standard instrumentations, before any player does anything about it. An ensemble drawn from one attack family has a spread of zero — every note is heard the same distance after it is started, so a common onset is a common heard moment — and a mixed one does not.

The attack time quoted for an instrument is a 10-to-90 per cent rise: the seconds between the envelope reaching a tenth of its eventual peak and reaching nine tenths. It is a measurable number and it varies enormously — three milliseconds for a struck bar, ninety for a bow, a hundred and ten for a sung vowel.

The model of the heard moment is then one line. The note is heard when its envelope crosses a criterion. Everything else is a question about which criterion, and that question has two answers because two different experiments were done.

  • Vos and Rasch (1981) asked when a tone begins, and found the answer at 6 to 15 decibels below its eventual maximum.
  • Gordon (1987) asked when a tone is aligned with a beat, on sixteen instrument sounds, and found the answer near the top of the rise rather than near its foot.

These are not rival measurements of one quantity; they are measurements of two. A note can be audible well before it is a beat. Both are drawn on every figure here, and the honest position is that the truth is somewhere between them.

Every criterion makes the lag proportional to the rise. The heard moment against the attack time, for the three criteria the literature supports. All three are straight lines through the origin, because each is a fixed fraction of the same envelope — so the criterion decides the slope and nothing else. At a 180 millisecond attack the three give 16 ms, 57 ms, 189 ms, a spread of a factor of 12. Every claim here is a difference between two instruments, and a difference is the same multiple of the same slope whichever line is taken.
Fig. 4 The lag against the attack time, under all three criteria. Every one is a straight line through the origin, because each is a fixed fraction of the same rising envelope — so the choice of criterion sets the slope and nothing else. The spread between the outermost two is a factor of twelve, which is a very poor state for an absolute number and an irrelevant one for a comparison: a difference between two instruments is the same multiple of whichever slope is used.

That is the reason this ladder can proceed at all. No number here rests on knowing the criterion, because every claim below is a difference — one instrument against another, one dynamic against another — and a difference of two lags is the difference of two rise times times a slope that cancels.

The slope is a property of the envelope’s shape as well

That protection is exactly right and it is narrower than it sounds, because the slope is not a property of the criterion alone. It is a property of the criterion and the shape of the rise, and the model assumes one shape for all nine instruments.

Every envelope here is a resonator filling from rest — one minus an exponential — which is a good description of a bowed string or a blown pipe and a poor one of a struck bar, whose amplitude is not filling anything. Computing the same three criteria on four plausible rise shapes, each normalised to the same 10-to-90 time:

lag ÷ rise time detection perceptual onset perceptual attack
exponential (a resonator filling) 0.089 0.317 1.048
linear ramp 0.222 0.626 1.125
raised cosine 0.469 0.848 1.347
square root (a fast onset) 0.040 0.314 1.012

At a fixed criterion the shape moves the slope by a factor of 2.7. That is smaller than the factor of twelve between the criteria, and unlike the criterion it does not cancel in a difference — because two instruments with different rise shapes have different slopes, and a difference of two lags is then not a common factor times anything.

What survives, and what the ladder should carry

Taking the essay’s own extreme pair, a marimba at 3 milliseconds against a bowed violin at 90, and moving each source of uncertainty in turn:

  • Changing the criterion, with both instruments exponential, moves the gap from 7.8 ms to 27.5 to 91.2 — a factor of twelve, which is the criterion spread and is exactly what the figure above reports.
  • Changing the shapes, at the middle criterion, moves it from 25.7 ms to 75.4 — a factor of three.
  • Changing the fast instrument’s shape alone moves it hardly at all: 27.5, 26.6, 25.9, 27.5 across the four. Three milliseconds times any slope is a small number, so the shape of a struck attack is very nearly irrelevant.

So the argument is repaired rather than lost, and repaired into a sharper statement. What cancels in a difference is the criterion, provided the two instruments share a rise shape. What does not cancel is the shape of the slower one, and every gap this ladder quotes is proportional to it. The whole of the uncertainty sits in one number — how a bow or a voice actually gets loud — and nothing in the model measures it.

That also says where a measurement would be worth making. Attack times are published for every instrument in the table and rise shapes are not; a single measured envelope for a bowed note would fix a factor of three in every number this ladder produces, and would do more for it than settling the criterion argument would, because the criterion argument was never load-bearing for a comparison and this is.

It is worth adding that the exponential is not an arbitrary choice made for convenience — it is what a linear resonator driven from rest does, so it is right for the mechanism at the slow end, which is the end that matters. The square-root row is there as the opposite extreme rather than as a candidate. The two shapes that would genuinely change the answers are the raised cosine, if a bow’s rise turns out to be gradual at both ends, and anything with a plateau in it, which is what the noisy scrape before Helmholtz motion establishes would produce.

Nine instruments, in order

How late each instrument is heard. Nine measured attack times converted to a heard moment at the 6 dB below peak criterion. The bar is the lag for the family's typical attack and the line through it is the range a player can produce on that instrument — which for the bowed and sung rows is wider than the gap between several of the other rows, so the ordering is a claim about typical playing and not about any single note. The fastest here is the marimba at 0.9 ms and the slowest the sung vowel at 35 ms.
Fig. 5 Nine measured attack times converted to heard moments at the middle criterion. The bar is the family’s typical attack; the line through it is the range a player can produce on that instrument. The struck and plucked family is under three milliseconds throughout, the blown family is ten to twenty-five, and the bowed and sung family is around thirty — and the range on the bowed row is wider than the gap between several whole families, which is why this is a claim about typical playing rather than about any particular note.

The ordering is the useful product and it is not a surprise once stated: an instrument whose sound is a struck object is heard at once, an instrument whose sound is a resonator being filled is heard when the resonator is full.

The blown family sitting between the two is the interesting row, because it is the family whose attack time is a choice. A marimba bar has one attack and a bow has a range of about four to one; a trumpet has a range of four to one as well, and a player moves through it deliberately — a tongued attack at one end, a slurred entry at the other. So a trumpeter’s heard moment is under the player’s control in a way a marimba player’s is not, and the control is exercised by an articulation mark rather than by a timing decision. That is the first place in this collection where a notated articulation turns out to be a notated timing.

What is a surprise is the size. Thirty milliseconds is not a subtlety. It is:

Two notes started together, heard 28 ms apart. Two amplitude envelopes rising from the same instant: a marimba with a 3 millisecond attack and a bowed violin with 90. The horizontal line is the criterion — 6 dB below peak, from Vos & Rasch 1981 — and the two dots are where each envelope crosses it. Nothing about the onsets differs; the heard moments differ by 28 ms, which is why the bowed violin has to start early to be heard on the beat. The buttons play the pair as written and then with the marimba delayed by that amount.
Fig. 6 The extreme case inside an ordinary ensemble: a marimba and a bowed violin, started together. Twenty-seven milliseconds separate the heard moments, and every one of them is a consequence of the two instruments being what they are. Nothing about the players is in the picture.

The measurement that was already there

None of this is a new observation about ensembles. It is a new explanation of an old one.

Rasch measured onset asynchrony in small ensembles in 1979, by recording trios playing from the same score and finding when each part actually began each note. The answer was between thirty and fifty milliseconds of spread, consistently, in groups of professionals who sounded perfectly together. It was reported as a limit on human precision — the players were trying to be simultaneous and this was how close they got.

The table above says something is wrong with that reading. If a recorder and a viol are heard thirty milliseconds apart when their onsets coincide, then a recorded onset spread of thirty milliseconds is compatible with the players being heard exactly together, and the measurement that looks like imprecision may be the measurement of a correction. The two readings predict opposite things: under the first, the spread should be unstructured and roughly the same whatever the instruments are; under the second, it should have a fixed sign and a size that tracks the difference in attack times.

Which of the two it is is not decided here — it is the whole of the next rung — but the possibility is enough to make the reading of every published asynchrony figure conditional. A number measured from onsets is a number about onsets.

The lag is not only a property of the note’s shape. It is a property of the note’s shape and of what the criterion is a fraction of, and those two come apart the moment a note is louder than its neighbours.

Under one model an accent arrives early, under the other it does notThe heard moment of a note with a 90 millisecond attack, against how much louder it is than the notes around it. The flat line is a criterion taken as a fraction of the note's own peak: the criterion rises with the note, so the crossing does not move at all. The falling line is a criterion taken as a fixed level: a louder note crosses the same line sooner, so an accent of 12 decibels is heard 23 ms early with no change in when the key was pressed. Nothing but an accent distinguishes the two models, which is why this figure is the experiment rather than the illustration.23 ms early05101520051015202530decibels louder than the surrounding notesmilliseconds from onset to the heard momentcriterion as a fractionof this note's own peakcriterion as a fixed levelset by everything else
Fig. 7 The heard moment of a note with a ninety millisecond attack, against how much louder it is than the notes around it, under two readings of the criterion. Taken as a fraction of the note’s own peak the criterion rises with the note and the crossing does not move at all. Taken as a fixed level a louder note crosses it sooner, so an accent is heard early.

Nothing in this essay chooses between those, and the choice is not a detail: one of them says an accent is on time and the other says a twelve-decibel accent arrives measurably before the note it is written with. A criterion relative to the note is a claim that the ear normalises, and a criterion at a fixed level is a claim that it does not — which is an experiment rather than an argument, and the collection does not have it.

What it is not, which is a reflection

There is a second window on this site measured in the same units, and confusing the two would be easy.

How late a reflection has to be before it is an echo. What a single reflection does to the sound it follows, against its delay, on a logarithmic axis. Under a millisecond the two combine into one image that is pulled towards the earlier source. From there out to a few tens of milliseconds the reflection is not heard as a separate event at all and does not move the image — it only changes the timbre. Past the echo threshold it becomes a second sound, and the threshold is five times later for speech than for a click.
Fig. 8 The precedence window, with this essay’s quantity marked on it. A reflection arriving 28 milliseconds after a direct sound is inside the region where the first wavefront wins — it is not heard as a separate event and it does not move the image. The two numbers are the same size and they are about different things: that one is two arrivals of one sound and this one is one arrival of one sound, heard late because it took time to get loud.

The distinction matters because both effects are present in a real hall at once, and only one of them depends on the instrument. A reflection is late by the geometry of the room; a bowed note is late by the shape of its own attack. Move the player and the first changes; nothing moves the second.

Which computation produced the numbers

Two functions and one constant.

The envelope is a resonator filling from rest, so its amplitude is one minus an exponential and its time constant is the measured 10-to-90 rise divided by ln 9 = 2.197. That is the only conversion in the model, and it is there because the rise time is what the literature reports while the time constant is what the physics has.

The heard moment is then the inverse of that: the time at which the envelope first reaches a stated fraction of its peak. At the middle criterion — 6 decibels below peak, which is half the peak amplitude — the lag is 0.317 times the rise time. At the detection criterion it is 0.089 times, and at the perceptual-attack criterion it is 1.048 times, which is to say very slightly after the nominal end of the rise.

The three slopes are worth having in one place, because every later number in this ladder is one of them times a rise time:

  • detection, 15 decibels below peak: 0.089 × the attack time;
  • perceptual onset, 6 decibels below peak: 0.317 ×;
  • perceptual attack, 90 per cent of peak: 1.048 ×.

Nothing is fitted. The three fractions are the published criteria, the rise times are published attack measurements, and the arithmetic between them is two lines. What the model does not contain is any dependence on pitch, on loudness or on what else is sounding — and the third rung of this ladder is about the second of those.

The three slopes above are quoted for the exponential envelope, which is the one every figure here draws. The shape comparison uses the same criteria against three other rise shapes, each rescaled so that its own 10-to-90 time is the quantity being varied — which is the step that makes the four comparable at all, since a shape and a duration are otherwise entangled. The 10-to-90 point of a raised cosine is 0.590 of its total rise and of a linear ramp is 0.8 of it, both by direct solution rather than by fitting.

None of the four shapes is measured. They are chosen to bracket what a rise can plausibly look like, and the exponential is the only one with a mechanism behind it. What the comparison establishes is the size of the exposure, not which shape is right.

Where the model stops

The envelope is one number and a real attack is not. A bowed note’s amplitude rise is not smooth: the Helmholtz motion has to establish itself, and until it does the sound is not the note but a noisy scrape whose spectrum is nothing like the eventual one. A criterion on the amplitude envelope will fire during that scrape, and whether a listener’s heard moment does the same is not something this model knows.

The criterion is a band a factor of twelve wide. Every absolute number above is quoted at one end of it. Every relative number is not, which is why the essay is written in differences — and the section above adds the qualification that a difference is free of the criterion only when both instruments share a rise shape, which is assumed here and measured nowhere.

And the heard moment is not necessarily a moment. The experiments that produced these criteria asked listeners to align a tone with a click or a beat, and got a distribution rather than a point; the width of that distribution, on slow attacks, is comparable to the lag itself. What the model computes is where the middle of it sits.

Whose music, and where it shows

The claim is about listeners and applies to any repertoire, but the consequences are about ensembles, and they show wherever instruments of very different attack are asked to play together.

The clearest case in the European tradition is the one that has been written about since the seventeenth century without this explanation: an organ and a choir. Both are slow-attack sources, so neither is heard when it starts, and an organist accompanying plainchant is famously told to play slightly ahead. A harpsichord continuo under a viol consort is the opposite pairing and gets the opposite advice.

The gamelan is the case where the effect nearly vanishes, and for a reason the table above makes obvious: almost every instrument in the ensemble is a struck bar or a struck gong, so the lags are all under three milliseconds and all equal. An ensemble of one attack family has no synchronisation problem of this kind at all — which is worth stating as a fact about the instruments rather than as a claim about the players.

Where the beats actually fall. Measured timing deviations from a strict grid, in milliseconds, for three published profiles. The right-hand column converts each deviation into the note value it would have to be written as, at three tempi — and because a fixed number of milliseconds is a different fraction of the beat at every tempo, no single notated rhythm describes any of these.
Fig. 9 The measured deviations all this now has to be read against: a Viennese waltz’s early second beat, a jazz soloist’s constant thirty milliseconds behind the ride, and a sequencer’s zero. Every number in this figure was measured from physical onsets. An earlier essay asks how much of the middle row is the drummer’s ride cymbal and the soloist’s saxophone being different instruments.

What the picture cannot show

Whether a listener uses the envelope at all. The model is a criterion on amplitude because that is what was measured, but a listener plausibly uses spectral change, or the arrival of a stable pitch, or the loudness rather than the amplitude — and on a slow attack those three do not coincide. Two of them would give a later heard moment than this model does and one would give an earlier one.

The pitch. Every figure here is drawn at one frequency. A low note’s envelope cannot rise faster than its own period, so a 40-hertz note has a floor on its attack of about 25 milliseconds that no player can get under, and the table above does not contain that floor.

And the number the ladder is really about is a difference of differences. What a conductor cares about is not that a violin is heard 28 milliseconds after it starts, but that a violin and a flute are heard 9 milliseconds apart, and that the gap changes when either of them plays louder.

Where this ladder goes next

One rung. A note is heard after it starts, the lag is a fixed fraction of its own attack time, and for the slow-attack instruments the fraction times the attack is about thirty milliseconds — the same size as everything the timing ladders measure.

The rung after it is the one the table makes unavoidable. If two instruments are heard at different moments when started together, then an ensemble that sounds together is not playing together, and the leads that have been measured in ensembles for fifty years — an organist ahead of a choir, a melody instrument ahead of its accompaniment — have been read as behaviour when part of them is arithmetic. How much is a computation, because the attack times are published and the leads are published, and nobody has put the two columns beside each other.

Part 1 of 9

One essay in the series on Perceptual-centre. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 15.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientEnsemble asynchronyEnvelopeIntegration timeInter-onset intervalOnsetPerceptual-centreSimultaneityTiming deviation