The players who have to be early
Assumes: A note is heard after it starts · Two players and no clock
A note is heard after it starts, by a fixed fraction of its own attack time. The consequence for one instrument is a curiosity. The consequence for several instruments at once is a number every ensemble carries around with it, and it can be computed from the instrument list alone.
The list is short and the two ends of it are the whole finding.
A string quartet and a gamelan get the same answer for opposite reasons. Four bowed strings are all heard about twenty-eight milliseconds after they are started; four struck bars are all heard about one millisecond after. In both cases the number is the same for every player, so a common onset is a common heard moment and the spread is exactly zero. What matters is not whether an ensemble is fast or slow but whether it is uniform.
The lead, per player
The three-instrument case is the one where a musician’s intuition and the arithmetic can be compared, because a piano trio has a documented practice and it is not this one.
The number for a piano trio is 26 milliseconds and the pianist is the one who has to move. That is the opposite of the folk advice, which puts the strings in charge of the ensemble and the piano in support of it; the arithmetic says the strings are the ones who cannot help being late and the piano is the one with the room to compensate.
A lutenist plucking with the string already under tension has an attack of about five milliseconds and a singer’s vowel takes a hundred and ten to arrive; the difference between the two is more than the twenty milliseconds at which two events stop being one.
Two players is the case where nothing can be blamed on coordination. A singer and a lutenist looking at each other, with no conductor and no ensemble to keep together, still have thirty-three milliseconds between the moment each of their notes is heard if both start at the same instant.
One family, and the ensemble that is one instrument
The zero rows are worth more attention than the large ones, because a quantity that is exactly zero for a whole class of ensembles is usually saying something about how the class was chosen.
A choir is the extreme case. Every singer has the same attack family, so the spread is zero however many of them there are — and what a choir does that a soloist cannot is already a result about many nominally identical sources being heard as one. The two results are the same shape: a section of one instrument type has no internal problem of this kind, in pitch or in time, because whatever error each member has, they all have the same one.
The moment the choir is joined by anything else the zero disappears. A choir with an organ carries eleven milliseconds; a choir with a lute or a harp carries thirty-three; a choir with a string orchestra carries about six, because a bowed attack and a sung attack are nearly the same length. That last number is small enough to be within the range a single player produces, which is a way of saying that voices and bowed strings are the same instrument for this purpose — a claim the vowel and the body filter would make on spectral grounds and which arrives here from timing.
An orchestra is every family at once, so its spread is the full nineteen to twenty-eight milliseconds, and the section that is always heard last is the strings.
The measurement this changes
Rasch recorded small ensembles in 1979 and measured when each part actually began each note. The spread was thirty to fifty milliseconds, consistently, in professionals who sounded entirely together, and it has been read ever since as the precision limit of ensemble playing.
The reading has a problem, and the problem has a sign.
The correction is not a constant offset that could be subtracted once. It is the difference of two lags, so it is zero within a section and largest between families — which means a measurement made on a string quartet needs none of this and the identical measurement made on a wind quintet needs nineteen milliseconds of it.
That is a bad property for a published figure to have, because the figure is usually quoted without the instrumentation.
Where the correction goes the wrong way
The most-quoted deviation in this collection is the jazz soloist thirty milliseconds behind the ride cymbal, and it is the case where this rung makes the problem worse rather than better.
A ride cymbal is a struck metal object and is heard within a millisecond. A saxophone is a blown reed with an attack of forty milliseconds or so and is heard about thirteen milliseconds after it starts. So a soloist whose onsets are thirty milliseconds behind the cymbal’s is heard forty-three milliseconds behind it.
The laid-back feel is not an artefact of onset measurement. It is thirty per cent larger than the measurement said.
What the ensemble is actually correcting for
The other half of this rung is that ensembles obviously do correct, and there is already a model of how on this site.
A correction loop with a systematic offset in it does not remove the offset — it removes the variance around it. Two players nudging toward each other’s onsets will converge on aligned onsets, which is exactly the state in which their heard moments are thirty milliseconds apart. Two players nudging toward each other’s heard moments will converge on the offset onsets and sound together.
Which of the two an ensemble does is an experiment nobody here can run, and it is the difference between an ensemble that sounds tight and one that merely is.
There is one piece of indirect evidence, and it is in the direction of the second. A correction loop driven by onsets would produce the same asynchrony whatever the instruments were, because it converges on zero onset difference by construction. What Rasch actually reported was that the spread varied with the ensemble, and was smallest in the group whose instruments were most alike. That is what a loop correcting toward heard moments would produce and not what a loop correcting toward onsets would.
It is weak evidence — the groups differed in more than their attack times — but it is the right shape of evidence, and it is the kind this rung makes it possible to look for.
Which computation produced the numbers
Each instrument’s heard-moment lag is its attack time times the criterion slope from the previous rung — 0.317 at the middle criterion — and the ensemble’s spread is the largest of those minus the smallest. That is the whole calculation, and the reason it is worth doing is that nobody has: the attack times have been in the timbre literature since the nineteen-seventies and the asynchrony figures have been in the performance literature since 1979, and the two columns have not been put side by side.
The leads are then the individual differences from the latest-heard member, which is why the latest member always shows zero: it is the ensemble’s own reference and cannot be early relative to itself.
Every number is quoted at one criterion and every number would change together if a different one were taken. A spread is a difference of two lags, both of which are the same slope times an attack time, so the ratio of any two ensembles’ spreads is fixed and the absolute scale is not. The piano trio’s twenty-six milliseconds against the wind group’s nineteen is a ratio this model is confident about; the twenty-six itself is a factor of three uncertain.
Where the model stops, and the thing that is the same size
The room is the same size as the effect and is not in the model. Sound covers a metre in 2.9 milliseconds. A symphony orchestra is about twelve metres deep, so the physical flight time from the back desks to the conductor is thirty-five milliseconds — larger than every spread in the hero figure. A player at the back who plays to be heard in time at their own position is late everywhere else in the hall, and no amount of attack arithmetic touches that.
The two effects are separable in principle and they behave differently: the flight time changes when the player moves and the attack lag does not, and the flight time is the same for every instrument at a given position while the attack lag is not. In a small ensemble sitting close together the flight time is a few milliseconds and the attack spread is tens, so this rung is the larger term; in an orchestra it is the smaller one.
And in an orchestra they very nearly cancel. Putting the two together over a standard seating — strings at two to six metres, wind at eight, brass at nine to twelve, timpani and percussion at thirteen and fourteen — the attack lag spans 27.5 milliseconds on its own, the flight time spans 33.5 on its own, and the two together span 11.7. Less than either, because they run in opposite directions:
| section | attack lag | flight | heard at |
|---|---|---|---|
| first violins | 28.5 | 7.3 | 35.8 ms |
| clarinets | 14.2 | 23.3 | 37.6 |
| horns | 9.5 | 27.7 | 37.2 |
| timpani | 0.9 | 37.9 | 38.9 |
| percussion | 0.9 | 40.8 | 41.8 |
The correlation between a section’s distance from the podium and its attack lag is −0.96. The instruments heard latest sit closest and the instruments heard soonest sit furthest away, almost exactly, so most of what one quantity gives away the other takes back. Reversing the seating order — percussion at the front, violins at the back — takes the combined spread to 61 milliseconds, five times worse than the arrangement everybody uses.
The invariance is worth one line because it is not obvious. Moving the listener back along the centre line adds the same flight time to every section, so the spread does not change at all: the 11.7 is what a conductor hears and also what row Q hears. It compresses only for a listener well off to the side, where the depth differences foreshorten.
None of this is a claim that anybody sat the orchestra down this way for this reason. The layout has obvious explanations — balance, sightlines, and putting the loudest instruments furthest from the audience — and it long predates any of the measurements. What the arithmetic says is that the layout those reasons produced happens to undo three quarters of a quantity nobody knew was there, and that the ensemble usually held up as the hardest to keep together is, on this measure, the best-arranged one in the hero figure. A string quartet gets zero by uniformity; an orchestra gets close to it by geometry.
Two consequences follow that are worth having. The first is that the residual is not evenly distributed: the double basses come out latest of anything on the platform, at 47 milliseconds, because they are the one slow-attack section that sits a long way back, and the first violins earliest at 36. Whatever an orchestra is fighting on this measure, it is fighting it between the front and the back of the string section rather than between the strings and the brass. The second is that the cancellation is a property of a deep platform. Halve every depth — a chamber orchestra on a shallow stage — and the flight spread halves while the attack spread does not, so the two stop matching and the combined figure rises rather than falls. The arrangement is tuned, accidentally, to the size of platform the repertoire was written for.
That difference is worth stating carefully, because both are called asynchrony and only one of them is fixable by playing differently. The first wavefront wins governs what a listener does with two arrivals of one sound; this rung governs one arrival of one sound. A hall’s first eighty milliseconds contains both, and the two are not separable in a recording made at one microphone position.
The attack time is not a constant of an instrument. The table’s ranges say so — a bowed attack runs from forty to a hundred and eighty milliseconds — and a player choosing an articulation is choosing a heard moment. That makes the ensemble spread a variable under partial control rather than a fixed property, and it makes the fixed claim above true only of typical playing.
And the conductor is not in the model at all. A beat pattern is a visual signal with its own latency, and the relationship between seeing a beat and putting a note down is a different literature with different numbers.
Whose music, and the design consequence
The claim is about instruments and applies wherever instruments are combined, but the interesting observation is about which ensembles were assembled and which were not.
The three ensembles in the hero figure with a spread of zero are the string quartet, the gamelan and — not drawn, because it is trivial — any single-instrument section. Each of the three is also, in its own tradition, the ensemble held up as the model of perfect ensemble playing. The European chamber-music ideal is four instruments of one family; the Javanese ideal is a set of instruments built together, struck, and tuned as one object.
The gamelan case is the stronger of the two, because the instruments were not merely chosen from one family but built as one set, tuned to each other rather than to a standard, and — as the tuning that is not a table of cents records — not transferable between ensembles. An instrument set assembled as a single object has uniform attack for the same reason it has uniform tuning: it was made that way.
That is not evidence that anybody knew about perceptual centres. It is evidence that whatever the ear rewards, these two traditions found it independently, and one of the things it rewards is uniformity of attack.
Against that, the ensembles with the largest spreads are the ones whose repertoire is most explicitly about conversation between unlike voices — a voice with a lute, a keyboard with strings — and in every one of them there is a documented performance tradition of one part leading. The tradition is usually explained as a matter of hierarchy, and the arithmetic here offers a duller explanation that predicts which part leads and by how much.
What the picture cannot show
Whether players compensate. Everything above says what the instruments do; nothing says what is done about it. If ensembles fully compensate, the spread is invisible in every recording and the effect is real and entirely absorbed; if they compensate partially, it shows up as an instrument-dependent bias in exactly the measurements that have been read as noise.
The onset is not always findable either. An onset detector reports the first energy, and a bowed note begins with a scrape that is not the note. Where the detector puts the onset of a slow entry is itself a choice of criterion, which means the published asynchrony figures already contain one criterion and this rung adds a second.
And the picture is of steady notes. The lags here are all for a note starting from silence. A note in the middle of a phrase begins from the previous note rather than from nothing, and what a legato line does to an attack is to remove most of it — so the spread is largest at an entry and smallest inside a line, which is the opposite of where the measurement literature has looked.
Where this ladder goes next
Two rungs. A note is heard late by a fraction of its own attack, and an ensemble that mixes attack families carries a heard-moment spread of ten to thirty-three milliseconds that is a property of the instrument list rather than of the players.
The rung after it is the one the range in the table points at. Every number above uses a single attack time per instrument, and the ranges are three or four to one — so the heard moment is not a fixed property of an instrument at all but something a player varies while playing. The obvious variable is dynamics, because a louder note is a faster rise on almost every instrument here. That is a computation, and it is the one that also decides between this ladder’s two models: a criterion taken as a fraction of a note’s own peak cannot move with level, and a criterion taken as a fixed level must.
Part 2 of 9
One essay in the series on Perceptual-centre. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
- A low note cannot start on time
- Which notes have to be played early
- An attack time is not an attack
- How many bars an ensemble needs
- The blend arrives before the note does
- Twelve violins are more punctual than one
- A note takes a number of periods to speak
- The higher note speaks sooner and takes longer
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Attack transientEnsemble asynchronyEntrainmentExpressive timingMotor delayOnsetPerceptual-centreTiming deviation
- Playing louder is playing earlier attack transient, onset, perceptual-centre, timing deviation
- Read at two different heights attack transient, ensemble asynchrony, onset, perceptual-centre
- A note starts twice attack transient, onset, perceptual-centre
- The passage that separates two players attack transient, onset, perceptual-centre
- What the tongue actually removes attack transient, ensemble asynchrony, onset
- A blown note does not start late, it starts slowly attack transient, onset