Instruments and their design

The players who have to be early

Ensembles have been measured for fifty years and found to be about forty milliseconds out of alignment, which has always been reported as the limit of human precision. Part of it is not: an ensemble that mixes attack families carries a heard-moment spread of ten to thirty-three milliseconds before anybody plays a note, and an ensemble drawn from one family carries none at all — which is true of a string quartet and of a gamelan for the same reason.

Assumes: A note is heard after it starts · Two players and no clock

A note is heard after it starts, by a fixed fraction of its own attack time. The consequence for one instrument is a curiosity. The consequence for several instruments at once is a number every ensemble carries around with it, and it can be computed from the instrument list alone.

Which ensembles have this problem and which do not. The width of the heard-moment spread built into 7 standard instrumentations, before any player does anything. An ensemble drawn from one attack family has a spread of zero — every note is heard the same distance after it is started, so a common onset is a common heard moment, and this is true of a string quartet and of a gamelan for the same reason and in the same amount. A mixed ensemble carries between 11 and 33 milliseconds of it.
Fig. 1 Seven standard instrumentations, drawn against the width of the heard-moment spread built into each of them. The two at the top are zero, and they are a string quartet and a gamelan — the slowest attacks in the collection and the fastest. Everything in between mixes families and carries between eleven and thirty-three milliseconds of spread before any player has done anything at all.

The list is short and the two ends of it are the whole finding.

A string quartet and a gamelan get the same answer for opposite reasons. Four bowed strings are all heard about twenty-eight milliseconds after they are started; four struck bars are all heard about one millisecond after. In both cases the number is the same for every player, so a common onset is a common heard moment and the spread is exactly zero. What matters is not whether an ensemble is fast or slow but whether it is uniform.

The lead, per player

What each player in piano trio has to do to be heard together. Each instrument of piano trio, drawn against the instant it has to put its note down in order for every note to be heard at the same moment. Zero is the common heard moment's own reference, which belongs to the latest-heard instrument in the group; everybody else starts earlier by the difference in their attack times. The whole width is 26.0 milliseconds, and it is a property of the instrumentation rather than of any performance.
Fig. 2 A piano trio, drawn against the instant each player has to put the note down for all three to be heard together. The two string players are the reference — they are heard last — and the pianist has to be twenty-six milliseconds early. That is about a thirty-eighth of a second, which at a moderate tempo is a twentieth of a beat, and it is not a preference or a style.

The three-instrument case is the one where a musician’s intuition and the arithmetic can be compared, because a piano trio has a documented practice and it is not this one.

The number for a piano trio is 26 milliseconds and the pianist is the one who has to move. That is the opposite of the folk advice, which puts the strings in charge of the ensemble and the piano in support of it; the arithmetic says the strings are the ones who cannot help being late and the piano is the one with the room to compensate.

What each player in voice and lute has to do to be heard together. Each instrument of voice and lute, drawn against the instant it has to put its note down in order for every note to be heard at the same moment. Zero is the common heard moment's own reference, which belongs to the latest-heard instrument in the group; everybody else starts earlier by the difference in their attack times. The whole width is 33.2 milliseconds, and it is a property of the instrumentation rather than of any performance.
Fig. 3 The largest spread in the list, and the smallest ensemble: a singer and a lute. Thirty-three milliseconds, because a plucked string is the fastest attack on the table and a sung vowel is the slowest. Two people, one part each, and a third of a tenth of a second built into the pairing.

A lutenist plucking with the string already under tension has an attack of about five milliseconds and a singer’s vowel takes a hundred and ten to arrive; the difference between the two is more than the twenty milliseconds at which two events stop being one.

Two players is the case where nothing can be blamed on coordination. A singer and a lutenist looking at each other, with no conductor and no ensemble to keep together, still have thirty-three milliseconds between the moment each of their notes is heard if both start at the same instant.

One family, and the ensemble that is one instrument

The zero rows are worth more attention than the large ones, because a quantity that is exactly zero for a whole class of ensembles is usually saying something about how the class was chosen.

A choir is the extreme case. Every singer has the same attack family, so the spread is zero however many of them there are — and what a choir does that a soloist cannot is already a result about many nominally identical sources being heard as one. The two results are the same shape: a section of one instrument type has no internal problem of this kind, in pitch or in time, because whatever error each member has, they all have the same one.

The moment the choir is joined by anything else the zero disappears. A choir with an organ carries eleven milliseconds; a choir with a lute or a harp carries thirty-three; a choir with a string orchestra carries about six, because a bowed attack and a sung attack are nearly the same length. That last number is small enough to be within the range a single player produces, which is a way of saying that voices and bowed strings are the same instrument for this purpose — a claim the vowel and the body filter would make on spectral grounds and which arrives here from timing.

An orchestra is every family at once, so its spread is the full nineteen to twenty-eight milliseconds, and the section that is always heard last is the strings.

How late each instrument is heard. Nine measured attack times converted to a heard moment at the 6 dB below peak criterion. The bar is the lag for the family's typical attack and the line through it is the range a player can produce on that instrument — which for the bowed and sung rows is wider than the gap between several of the other rows, so the ordering is a claim about typical playing and not about any single note. The fastest here is the marimba at 0.9 ms and the slowest the sung vowel at 35 ms.
Fig. 4 The nine attack families the ensemble arithmetic is built from, with the range each instrument can produce. Reading an ensemble’s spread off this figure is subtraction: the top of the group minus the bottom. A choir and an organ are the two slowest rows and eleven milliseconds apart; a lute and a voice are the top and the bottom of the whole table.

The measurement this changes

Rasch recorded small ensembles in 1979 and measured when each part actually began each note. The spread was thirty to fifty milliseconds, consistently, in professionals who sounded entirely together, and it has been read ever since as the precision limit of ensemble playing.

The reading has a problem, and the problem has a sign.

The same measurement, read from onsets and read from heard moments. Three published timing deviations, each drawn twice: the open mark is what an onset detector measured and the filled mark is where the note is heard, given what the two instruments are. The correction has a sign. A slow-attack soloist against a fast-attack timekeeper is heard further behind than the measurement says — 30 milliseconds becomes 43 — and two players on the same instrument get no correction at all, which is why a string section's internal timing needs none of this.
Fig. 5 Three published timing deviations, each drawn twice: open where an onset detector measured it, filled where the note is heard. The correction depends entirely on which two instruments are involved. The soloist behind the timekeeper is heard further behind than the measurement says; the two violins get no correction at all, because the correction is a difference and theirs is zero; and the organ, playing exactly with the choir, is heard eleven milliseconds ahead of it.

The correction is not a constant offset that could be subtracted once. It is the difference of two lags, so it is zero within a section and largest between families — which means a measurement made on a string quartet needs none of this and the identical measurement made on a wind quintet needs nineteen milliseconds of it.

That is a bad property for a published figure to have, because the figure is usually quoted without the instrumentation.

Where the correction goes the wrong way

The most-quoted deviation in this collection is the jazz soloist thirty milliseconds behind the ride cymbal, and it is the case where this rung makes the problem worse rather than better.

A ride cymbal is a struck metal object and is heard within a millisecond. A saxophone is a blown reed with an attack of forty milliseconds or so and is heard about thirteen milliseconds after it starts. So a soloist whose onsets are thirty milliseconds behind the cymbal’s is heard forty-three milliseconds behind it.

The laid-back feel is not an artefact of onset measurement. It is thirty per cent larger than the measurement said.

Two notes started together, heard 22 ms apart. Two amplitude envelopes rising from the same instant: a plucked string with a 5 millisecond attack and an organ flue pipe with 75. The horizontal line is the criterion — 6 dB below peak, from Vos & Rasch 1981 — and the two dots are where each envelope crosses it. Nothing about the onsets differs; the heard moments differ by 22 ms, which is why the organ flue pipe has to start early to be heard on the beat. The buttons play the pair as written and then with the plucked string delayed by that amount.
Fig. 6 The pairing that produces the largest correction available in an ordinary church: a plucked string and an organ flue pipe. Twenty-two milliseconds separate the heard moments at a common onset. An organist accompanying a congregation is compensating for this, for the building’s reverberation, and for the fact that a congregation is not listening to either — and only the first of the three is computable from the instruments.
Two kinds of systematic timing, which share a word. Each measured profile split into a constant offset from the grid and a pattern that varies by position in the bar. Viennese waltz, second beat is −10.0 ms of offset and 33.4 ms of pattern; jazz soloist against the ride is 28.8 ms of offset and 1.9 ms of pattern; quantised is 0.0 ms of offset and 0.0 ms of pattern. A motor deviation anywhere in the 8 to 20 ms range published for skilled performers leaves 74–95% of Viennese waltz, second beat's variation systematic, 1–5% of jazz soloist against the ride's variation systematic. The two quantities are independent and no single deviation figure distinguishes them.
Fig. 7 The two systematic quantities a timing profile carries, separated: a constant offset from the grid, and a pattern that varies with position in the bar. The correction this essay supplies is entirely in the first column and not at all in the second — an attack time does not know which beat of the bar it is on — so a profile that is nearly all pattern, like the Viennese waltz, is untouched by any of this, and one that is nearly all offset, like the jazz soloist’s, is the case where it matters most.

What the ensemble is actually correcting for

The other half of this rung is that ensembles obviously do correct, and there is already a model of how on this site.

Two players, and the correction that keeps them together. The spread of the asynchrony between two players, in milliseconds, against beat number, for 3 correction gains, averaged over 120 seeded runs each. It reaches 120 ms after 64 beats at a gain of 0, 27 ms after 64 beats at a gain of 0.1, 20 ms after 64 beats at a gain of 0.3. With no correction at all the asynchrony is a random walk and grows without bound; with any correction it settles at a fixed spread within a few beats and stays there. Two people cannot share a timekeeper, so the fact that ensembles do not come apart is itself the evidence that they are correcting.
Fig. 8 The two-player correction model from two players and no clock: each player nudges toward the other by a gain, and the pair’s asynchrony settles at a width the two gains set. The defect those essays recorded is that the model cannot say who is following, because only the sum of the gains is observable. This one adds a second thing the model cannot say: some of the asynchrony it is correcting is not an error.

A correction loop with a systematic offset in it does not remove the offset — it removes the variance around it. Two players nudging toward each other’s onsets will converge on aligned onsets, which is exactly the state in which their heard moments are thirty milliseconds apart. Two players nudging toward each other’s heard moments will converge on the offset onsets and sound together.

Which of the two an ensemble does is an experiment nobody here can run, and it is the difference between an ensemble that sounds tight and one that merely is.

There is one piece of indirect evidence, and it is in the direction of the second. A correction loop driven by onsets would produce the same asynchrony whatever the instruments were, because it converges on zero onset difference by construction. What Rasch actually reported was that the spread varied with the ensemble, and was smallest in the group whose instruments were most alike. That is what a loop correcting toward heard moments would produce and not what a loop correcting toward onsets would.

It is weak evidence — the groups differed in more than their attack times — but it is the right shape of evidence, and it is the kind this rung makes it possible to look for.

Which computation produced the numbers

Each instrument’s heard-moment lag is its attack time times the criterion slope from the previous rung — 0.317 at the middle criterion — and the ensemble’s spread is the largest of those minus the smallest. That is the whole calculation, and the reason it is worth doing is that nobody has: the attack times have been in the timbre literature since the nineteen-seventies and the asynchrony figures have been in the performance literature since 1979, and the two columns have not been put side by side.

The leads are then the individual differences from the latest-heard member, which is why the latest member always shows zero: it is the ensemble’s own reference and cannot be early relative to itself.

Every number is quoted at one criterion and every number would change together if a different one were taken. A spread is a difference of two lags, both of which are the same slope times an attack time, so the ratio of any two ensembles’ spreads is fixed and the absolute scale is not. The piano trio’s twenty-six milliseconds against the wind group’s nineteen is a ratio this model is confident about; the twenty-six itself is a factor of three uncertain.

Where the model stops, and the thing that is the same size

The room is the same size as the effect and is not in the model. Sound covers a metre in 2.9 milliseconds. A symphony orchestra is about twelve metres deep, so the physical flight time from the back desks to the conductor is thirty-five milliseconds — larger than every spread in the hero figure. A player at the back who plays to be heard in time at their own position is late everywhere else in the hall, and no amount of attack arithmetic touches that.

The two effects are separable in principle and they behave differently: the flight time changes when the player moves and the attack lag does not, and the flight time is the same for every instrument at a given position while the attack lag is not. In a small ensemble sitting close together the flight time is a few milliseconds and the attack spread is tens, so this rung is the larger term; in an orchestra it is the smaller one.

And in an orchestra they very nearly cancel. Putting the two together over a standard seating — strings at two to six metres, wind at eight, brass at nine to twelve, timpani and percussion at thirteen and fourteen — the attack lag spans 27.5 milliseconds on its own, the flight time spans 33.5 on its own, and the two together span 11.7. Less than either, because they run in opposite directions:

section attack lag flight heard at
first violins 28.5 7.3 35.8 ms
clarinets 14.2 23.3 37.6
horns 9.5 27.7 37.2
timpani 0.9 37.9 38.9
percussion 0.9 40.8 41.8

The correlation between a section’s distance from the podium and its attack lag is −0.96. The instruments heard latest sit closest and the instruments heard soonest sit furthest away, almost exactly, so most of what one quantity gives away the other takes back. Reversing the seating order — percussion at the front, violins at the back — takes the combined spread to 61 milliseconds, five times worse than the arrangement everybody uses.

The invariance is worth one line because it is not obvious. Moving the listener back along the centre line adds the same flight time to every section, so the spread does not change at all: the 11.7 is what a conductor hears and also what row Q hears. It compresses only for a listener well off to the side, where the depth differences foreshorten.

None of this is a claim that anybody sat the orchestra down this way for this reason. The layout has obvious explanations — balance, sightlines, and putting the loudest instruments furthest from the audience — and it long predates any of the measurements. What the arithmetic says is that the layout those reasons produced happens to undo three quarters of a quantity nobody knew was there, and that the ensemble usually held up as the hardest to keep together is, on this measure, the best-arranged one in the hero figure. A string quartet gets zero by uniformity; an orchestra gets close to it by geometry.

Two consequences follow that are worth having. The first is that the residual is not evenly distributed: the double basses come out latest of anything on the platform, at 47 milliseconds, because they are the one slow-attack section that sits a long way back, and the first violins earliest at 36. Whatever an orchestra is fighting on this measure, it is fighting it between the front and the back of the string section rather than between the strings and the brass. The second is that the cancellation is a property of a deep platform. Halve every depth — a chamber orchestra on a shallow stage — and the flight spread halves while the attack spread does not, so the two stop matching and the combined figure rises rather than falls. The arrangement is tuned, accidentally, to the size of platform the repertoire was written for.

That difference is worth stating carefully, because both are called asynchrony and only one of them is fixable by playing differently. The first wavefront wins governs what a listener does with two arrivals of one sound; this rung governs one arrival of one sound. A hall’s first eighty milliseconds contains both, and the two are not separable in a recording made at one microphone position.

The attack time is not a constant of an instrument. The table’s ranges say so — a bowed attack runs from forty to a hundred and eighty milliseconds — and a player choosing an articulation is choosing a heard moment. That makes the ensemble spread a variable under partial control rather than a fixed property, and it makes the fixed claim above true only of typical playing.

And the conductor is not in the model at all. A beat pattern is a visual signal with its own latency, and the relationship between seeing a beat and putting a note down is a different literature with different numbers.

Whose music, and the design consequence

The claim is about instruments and applies wherever instruments are combined, but the interesting observation is about which ensembles were assembled and which were not.

The three ensembles in the hero figure with a spread of zero are the string quartet, the gamelan and — not drawn, because it is trivial — any single-instrument section. Each of the three is also, in its own tradition, the ensemble held up as the model of perfect ensemble playing. The European chamber-music ideal is four instruments of one family; the Javanese ideal is a set of instruments built together, struck, and tuned as one object.

The gamelan case is the stronger of the two, because the instruments were not merely chosen from one family but built as one set, tuned to each other rather than to a standard, and — as the tuning that is not a table of cents records — not transferable between ensembles. An instrument set assembled as a single object has uniform attack for the same reason it has uniform tuning: it was made that way.

That is not evidence that anybody knew about perceptual centres. It is evidence that whatever the ear rewards, these two traditions found it independently, and one of the things it rewards is uniformity of attack.

Against that, the ensembles with the largest spreads are the ones whose repertoire is most explicitly about conversation between unlike voices — a voice with a lute, a keyboard with strings — and in every one of them there is a documented performance tradition of one part leading. The tradition is usually explained as a matter of hierarchy, and the arithmetic here offers a duller explanation that predicts which part leads and by how much.

What the picture cannot show

Whether players compensate. Everything above says what the instruments do; nothing says what is done about it. If ensembles fully compensate, the spread is invisible in every recording and the effect is real and entirely absorbed; if they compensate partially, it shows up as an instrument-dependent bias in exactly the measurements that have been read as noise.

The onset is not always findable either. An onset detector reports the first energy, and a bowed note begins with a scrape that is not the note. Where the detector puts the onset of a slow entry is itself a choice of criterion, which means the published asynchrony figures already contain one criterion and this rung adds a second.

And the picture is of steady notes. The lags here are all for a note starting from silence. A note in the middle of a phrase begins from the previous note rather than from nothing, and what a legato line does to an attack is to remove most of it — so the spread is largest at an entry and smallest inside a line, which is the opposite of where the measurement literature has looked.

Where this ladder goes next

Two rungs. A note is heard late by a fraction of its own attack, and an ensemble that mixes attack families carries a heard-moment spread of ten to thirty-three milliseconds that is a property of the instrument list rather than of the players.

The rung after it is the one the range in the table points at. Every number above uses a single attack time per instrument, and the ranges are three or four to one — so the heard moment is not a fixed property of an instrument at all but something a player varies while playing. The obvious variable is dynamics, because a louder note is a faster rise on almost every instrument here. That is a computation, and it is the one that also decides between this ladder’s two models: a criterion taken as a fraction of a note’s own peak cannot move with level, and a criterion taken as a fixed level must.

Part 2 of 9

One essay in the series on Perceptual-centre. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientEnsemble asynchronyEntrainmentExpressive timingMotor delayOnsetPerceptual-centreTiming deviation