A note is heard after it starts
Assumes: The first fifty milliseconds · The beat is inferred, and sometimes wrongly
A note starts at a definite instant. The first sample that is not silence is a fact about a recording, it can be found automatically, and every rhythm essay in this collection has used it: the beat is inferred from a list of onsets, a groove is a list of departures from a grid, two players correcting toward each other are two lists of onsets compared.
The instant a listener hears the note is a different instant, and it is later.
The gap is not small. On the middle criterion it is 28 milliseconds, which is very nearly the whole of the thirty-millisecond lag a jazz soloist sits behind the ride cymbal — a deviation that ladder treats as an expressive choice, measured from onsets.
Two notes, one onset, two moments
The claim is easiest to see with two notes rather than one, because the absolute lag depends on a criterion nobody agrees about and the difference between two lags does not.
Started together, they are not heard together. The piano is heard 32 milliseconds before the voice, which is more than the twenty milliseconds at which two events stop being one and become a note and a late note.
This is not a defect of anybody’s playing. It is arithmetic on the two envelopes, and it says that an accompanist who wants to sound simultaneous with a singer has to play about a thirtieth of a second late.
What the envelope is, and why it is the whole of the mechanism
The site already has the object this runs on. The shape of a note is an attack, a decay, a sustain and a release, and the first of the four is the one under discussion.
A marimba’s attack is three milliseconds, a trumpet’s thirty, a bowed violin’s ninety, and the whole quantity this essay is about lives inside the first twentieth of a two-second note — which is why every other figure here is drawn at a quarter of a second. The reason so small a corner matters is that it is not the same corner for everyone playing at once.
The attack time quoted for an instrument is a 10-to-90 per cent rise: the seconds between the envelope reaching a tenth of its eventual peak and reaching nine tenths. It is a measurable number and it varies enormously — three milliseconds for a struck bar, ninety for a bow, a hundred and ten for a sung vowel.
The model of the heard moment is then one line. The note is heard when its envelope crosses a criterion. Everything else is a question about which criterion, and that question has two answers because two different experiments were done.
- Vos and Rasch (1981) asked when a tone begins, and found the answer at 6 to 15 decibels below its eventual maximum.
- Gordon (1987) asked when a tone is aligned with a beat, on sixteen instrument sounds, and found the answer near the top of the rise rather than near its foot.
These are not rival measurements of one quantity; they are measurements of two. A note can be audible well before it is a beat. Both are drawn on every figure here, and the honest position is that the truth is somewhere between them.
That is the reason this ladder can proceed at all. No number here rests on knowing the criterion, because every claim below is a difference — one instrument against another, one dynamic against another — and a difference of two lags is the difference of two rise times times a slope that cancels.
The slope is a property of the envelope’s shape as well
That protection is exactly right and it is narrower than it sounds, because the slope is not a property of the criterion alone. It is a property of the criterion and the shape of the rise, and the model assumes one shape for all nine instruments.
Every envelope here is a resonator filling from rest — one minus an exponential — which is a good description of a bowed string or a blown pipe and a poor one of a struck bar, whose amplitude is not filling anything. Computing the same three criteria on four plausible rise shapes, each normalised to the same 10-to-90 time:
| lag ÷ rise time | detection | perceptual onset | perceptual attack |
|---|---|---|---|
| exponential (a resonator filling) | 0.089 | 0.317 | 1.048 |
| linear ramp | 0.222 | 0.626 | 1.125 |
| raised cosine | 0.469 | 0.848 | 1.347 |
| square root (a fast onset) | 0.040 | 0.314 | 1.012 |
At a fixed criterion the shape moves the slope by a factor of 2.7. That is smaller than the factor of twelve between the criteria, and unlike the criterion it does not cancel in a difference — because two instruments with different rise shapes have different slopes, and a difference of two lags is then not a common factor times anything.
What survives, and what the ladder should carry
Taking the essay’s own extreme pair, a marimba at 3 milliseconds against a bowed violin at 90, and moving each source of uncertainty in turn:
- Changing the criterion, with both instruments exponential, moves the gap from 7.8 ms to 27.5 to 91.2 — a factor of twelve, which is the criterion spread and is exactly what the figure above reports.
- Changing the shapes, at the middle criterion, moves it from 25.7 ms to 75.4 — a factor of three.
- Changing the fast instrument’s shape alone moves it hardly at all: 27.5, 26.6, 25.9, 27.5 across the four. Three milliseconds times any slope is a small number, so the shape of a struck attack is very nearly irrelevant.
So the argument is repaired rather than lost, and repaired into a sharper statement. What cancels in a difference is the criterion, provided the two instruments share a rise shape. What does not cancel is the shape of the slower one, and every gap this ladder quotes is proportional to it. The whole of the uncertainty sits in one number — how a bow or a voice actually gets loud — and nothing in the model measures it.
That also says where a measurement would be worth making. Attack times are published for every instrument in the table and rise shapes are not; a single measured envelope for a bowed note would fix a factor of three in every number this ladder produces, and would do more for it than settling the criterion argument would, because the criterion argument was never load-bearing for a comparison and this is.
It is worth adding that the exponential is not an arbitrary choice made for convenience — it is what a linear resonator driven from rest does, so it is right for the mechanism at the slow end, which is the end that matters. The square-root row is there as the opposite extreme rather than as a candidate. The two shapes that would genuinely change the answers are the raised cosine, if a bow’s rise turns out to be gradual at both ends, and anything with a plateau in it, which is what the noisy scrape before Helmholtz motion establishes would produce.
Nine instruments, in order
The ordering is the useful product and it is not a surprise once stated: an instrument whose sound is a struck object is heard at once, an instrument whose sound is a resonator being filled is heard when the resonator is full.
The blown family sitting between the two is the interesting row, because it is the family whose attack time is a choice. A marimba bar has one attack and a bow has a range of about four to one; a trumpet has a range of four to one as well, and a player moves through it deliberately — a tongued attack at one end, a slurred entry at the other. So a trumpeter’s heard moment is under the player’s control in a way a marimba player’s is not, and the control is exercised by an articulation mark rather than by a timing decision. That is the first place in this collection where a notated articulation turns out to be a notated timing.
What is a surprise is the size. Thirty milliseconds is not a subtlety. It is:
- larger than the twenty milliseconds inside which two onsets fuse into one event;
- the same order as the whole systematic deviation in every timing profile the microtiming ladder measured;
- about a twentieth of the half-second beat that the preferred tempo sits at, so a bar of four carries four of them.
The measurement that was already there
None of this is a new observation about ensembles. It is a new explanation of an old one.
Rasch measured onset asynchrony in small ensembles in 1979, by recording trios playing from the same score and finding when each part actually began each note. The answer was between thirty and fifty milliseconds of spread, consistently, in groups of professionals who sounded perfectly together. It was reported as a limit on human precision — the players were trying to be simultaneous and this was how close they got.
The table above says something is wrong with that reading. If a recorder and a viol are heard thirty milliseconds apart when their onsets coincide, then a recorded onset spread of thirty milliseconds is compatible with the players being heard exactly together, and the measurement that looks like imprecision may be the measurement of a correction. The two readings predict opposite things: under the first, the spread should be unstructured and roughly the same whatever the instruments are; under the second, it should have a fixed sign and a size that tracks the difference in attack times.
Which of the two it is is not decided here — it is the whole of the next rung — but the possibility is enough to make the reading of every published asynchrony figure conditional. A number measured from onsets is a number about onsets.
The lag is not only a property of the note’s shape. It is a property of the note’s shape and of what the criterion is a fraction of, and those two come apart the moment a note is louder than its neighbours.
Nothing in this essay chooses between those, and the choice is not a detail: one of them says an accent is on time and the other says a twelve-decibel accent arrives measurably before the note it is written with. A criterion relative to the note is a claim that the ear normalises, and a criterion at a fixed level is a claim that it does not — which is an experiment rather than an argument, and the collection does not have it.
What it is not, which is a reflection
There is a second window on this site measured in the same units, and confusing the two would be easy.
The distinction matters because both effects are present in a real hall at once, and only one of them depends on the instrument. A reflection is late by the geometry of the room; a bowed note is late by the shape of its own attack. Move the player and the first changes; nothing moves the second.
Which computation produced the numbers
Two functions and one constant.
The envelope is a resonator filling from rest, so its amplitude is one minus an exponential and its time constant is the measured 10-to-90 rise divided by ln 9 = 2.197. That is the only conversion in the model, and it is there because the rise time is what the literature reports while the time constant is what the physics has.
The heard moment is then the inverse of that: the time at which the envelope first reaches a stated fraction of its peak. At the middle criterion — 6 decibels below peak, which is half the peak amplitude — the lag is 0.317 times the rise time. At the detection criterion it is 0.089 times, and at the perceptual-attack criterion it is 1.048 times, which is to say very slightly after the nominal end of the rise.
The three slopes are worth having in one place, because every later number in this ladder is one of them times a rise time:
- detection, 15 decibels below peak: 0.089 × the attack time;
- perceptual onset, 6 decibels below peak: 0.317 ×;
- perceptual attack, 90 per cent of peak: 1.048 ×.
Nothing is fitted. The three fractions are the published criteria, the rise times are published attack measurements, and the arithmetic between them is two lines. What the model does not contain is any dependence on pitch, on loudness or on what else is sounding — and the third rung of this ladder is about the second of those.
The three slopes above are quoted for the exponential envelope, which is the one every figure here draws. The shape comparison uses the same criteria against three other rise shapes, each rescaled so that its own 10-to-90 time is the quantity being varied — which is the step that makes the four comparable at all, since a shape and a duration are otherwise entangled. The 10-to-90 point of a raised cosine is 0.590 of its total rise and of a linear ramp is 0.8 of it, both by direct solution rather than by fitting.
None of the four shapes is measured. They are chosen to bracket what a rise can plausibly look like, and the exponential is the only one with a mechanism behind it. What the comparison establishes is the size of the exposure, not which shape is right.
Where the model stops
The envelope is one number and a real attack is not. A bowed note’s amplitude rise is not smooth: the Helmholtz motion has to establish itself, and until it does the sound is not the note but a noisy scrape whose spectrum is nothing like the eventual one. A criterion on the amplitude envelope will fire during that scrape, and whether a listener’s heard moment does the same is not something this model knows.
The criterion is a band a factor of twelve wide. Every absolute number above is quoted at one end of it. Every relative number is not, which is why the essay is written in differences — and the section above adds the qualification that a difference is free of the criterion only when both instruments share a rise shape, which is assumed here and measured nowhere.
And the heard moment is not necessarily a moment. The experiments that produced these criteria asked listeners to align a tone with a click or a beat, and got a distribution rather than a point; the width of that distribution, on slow attacks, is comparable to the lag itself. What the model computes is where the middle of it sits.
Whose music, and where it shows
The claim is about listeners and applies to any repertoire, but the consequences are about ensembles, and they show wherever instruments of very different attack are asked to play together.
The clearest case in the European tradition is the one that has been written about since the seventeenth century without this explanation: an organ and a choir. Both are slow-attack sources, so neither is heard when it starts, and an organist accompanying plainchant is famously told to play slightly ahead. A harpsichord continuo under a viol consort is the opposite pairing and gets the opposite advice.
The gamelan is the case where the effect nearly vanishes, and for a reason the table above makes obvious: almost every instrument in the ensemble is a struck bar or a struck gong, so the lags are all under three milliseconds and all equal. An ensemble of one attack family has no synchronisation problem of this kind at all — which is worth stating as a fact about the instruments rather than as a claim about the players.
What the picture cannot show
Whether a listener uses the envelope at all. The model is a criterion on amplitude because that is what was measured, but a listener plausibly uses spectral change, or the arrival of a stable pitch, or the loudness rather than the amplitude — and on a slow attack those three do not coincide. Two of them would give a later heard moment than this model does and one would give an earlier one.
The pitch. Every figure here is drawn at one frequency. A low note’s envelope cannot rise faster than its own period, so a 40-hertz note has a floor on its attack of about 25 milliseconds that no player can get under, and the table above does not contain that floor.
And the number the ladder is really about is a difference of differences. What a conductor cares about is not that a violin is heard 28 milliseconds after it starts, but that a violin and a flute are heard 9 milliseconds apart, and that the gap changes when either of them plays louder.
Where this ladder goes next
One rung. A note is heard after it starts, the lag is a fixed fraction of its own attack time, and for the slow-attack instruments the fraction times the attack is about thirty milliseconds — the same size as everything the timing ladders measure.
The rung after it is the one the table makes unavoidable. If two instruments are heard at different moments when started together, then an ensemble that sounds together is not playing together, and the leads that have been measured in ensembles for fifty years — an organist ahead of a choir, a melody instrument ahead of its accompaniment — have been read as behaviour when part of them is arithmetic. How much is a computation, because the attack times are published and the leads are published, and nobody has put the two columns beside each other.
Part 1 of 9
One essay in the series on Perceptual-centre. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 15.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Attack transientEnsemble asynchronyEnvelopeIntegration timeInter-onset intervalOnsetPerceptual-centreSimultaneityTiming deviation
- Which notes have to be played early attack transient, ensemble asynchrony, onset, perceptual-centre, timing deviation
- Playing louder is playing earlier attack transient, onset, perceptual-centre, timing deviation
- The blend arrives before the note does attack transient, envelope, onset, perceptual-centre
- What the tongue actually removes attack transient, ensemble asynchrony, envelope, onset
- A blown note does not start late, it starts slowly attack transient, onset
- A hammer is not an impulse attack transient, envelope