Rhythm and metre

How many bars an ensemble needs

The map of required leads is something nobody tells the players, because the leads are a property of the instruments' attacks. So an ensemble has to find them, and the mechanism is the one the microtiming essays describe: each player hears sounds rather than onsets and moves toward the others. It converges on the map to within a fifth of a millisecond, in five beats, and there is a best correction gain.

Assumes: Which notes have to be played early · Two players and no clock

Which notes have to be played early produced a map: for a scored chord, which parts have to start before the beat and by how much, from the attack family, the pitch and the dynamic of each. Its last paragraph named the thing nobody had done with it.

The map above is exactly the input that model wants: a set of required leads that the players do not know and have to discover by listening.

Nobody tells the players those numbers. They are a property of the instruments’ attacks, not of the music, and they change with the register and the dynamic of every chord. So an ensemble has to find them — and this collection has a model of exactly the mechanism by which it would.

An ensemble finding an asynchrony nobody told it about. An earlier essay produced a map of required leads — which notes of a scoring have to be played early, and by how much — and nothing tells the players those numbers, because they are a property of the instruments' attacks rather than of the music. So an ensemble has to find them, and the mechanism is already here: each player hears sounds rather than onsets and moves their next onset toward the mean of the others'. The spread of arrival times starts at 20 milliseconds and settles at 4, crossing 5 milliseconds after 5 beats — about 1.3 bars of four. The leads it converges on match that map to within 0.2 milliseconds, which is what makes this a convergence rather than a coincidence: the fixed point of players listening to each other is every player leading by their own attack.
Fig. 1 Four parts starting together and listening. The spread of their sounds falls from twenty milliseconds to under five in five beats, and the leads they arrive at are the earlier map to within a fifth of a millisecond.

The mechanism, and why its fixed point is the map

Two players and no clock is the model: each player, on each beat, moves their next onset toward what they heard, with a gain. There is no conductor and no metronome; the ensemble’s timing is the fixed point of everybody correcting.

The thing that makes it produce this map is a distinction the fifth rung established and this rung uses.

Players hear sounds, not onsets. A note’s physical onset and the moment a listener places it differ by the perceptual-centre lag, which for a bowed violin at a moderate dynamic is 28 milliseconds and for a forte trumpet is 9. So when four players listen to each other, what each of them hears is everybody’s onset plus everybody’s lag.

If each player moves toward the mean of what they heard, the only arrangement nobody wants to move away from is the one where all four sounds coincide. That requires onsets separated by exactly the differences of the lags — which is the fifth rung’s map.

So the map is not something the ensemble has to be told. It is the fixed point of listening, and the model’s job is to say how long it takes to get there.

Five beats, and the leads come out right

Running it with a correction gain of 0.25, four parts starting together and no other information:

The spread of arrival times starts at 19.9 milliseconds — the map’s own spread, because everybody is starting on the beat and their lags are what they are.

It falls to under 5 milliseconds after 5 beats. At a moderate tempo that is a bar and a quarter.

It settles at about 3.7 milliseconds, which is the residue of the players chasing each other’s motor noise.

And the leads the ensemble arrives at are the fifth rung’s map: 19.9 for the violin, 13.2 for the flute, 12.0 for the piano and 0 for the trumpet, matched to within 0.22 milliseconds.

That last number is the check rather than the finding. The convergence recovering the map is what makes this a joining of two models rather than a new one, and if it had converged somewhere else one of the two would have been wrong.

What the fifth rung got backwards, and why it matters here

Running the two models together found an error in the fifth rung, and it is recorded here because this rung is where it showed.

The fifth rung reported its leads with the sign inverted. A part whose sound arrives late — a bowed violin, with the slowest attack of the four — must start early, and the published figure gave the violin a lead of zero and the trumpet, whose attack is the quickest, a lead of twenty milliseconds.

The physics is not in doubt: a note heard 28 milliseconds after it starts and a note heard 9 milliseconds after it starts coincide only if the first begins 19 milliseconds before the second. The convergence model has no choice about this — its fixed point is whatever makes the sounds line up — so running it was what surfaced the error.

The fifth rung’s essay is corrected. Its headline changes and its finding survives in a better form: the piano at E1 needs a lead of twelve milliseconds despite having the quickest family attack of the four, because its pitch floor has taken that attack away from it. The surprise is not that a blown instrument is earlier than a struck one; it is that a struck instrument in the bass is nearly as early as a bowed one.

Which notes of a scored chord have to be played early. Four parts of one chord, each with its own instrument, its own pitch and its own dynamic, and the perceptual centre that comes out of all three. piano, sforzando on E1: an attack family of 8 milliseconds against a pitch floor of 97, so the pitch is what limits it, shortened by the dynamic to 65, heard 20.5 after it starts and needing to be played 12.0 early; flute, quiet on A5: an attack family of 60 milliseconds against a pitch floor of 5, so the instrument is, shortened by the dynamic to 69, heard 21.8 after it starts and needing to be played 13.2 early; violin, mezzo forte on E4: an attack family of 90 milliseconds against a pitch floor of 12, so the instrument is, shortened by the dynamic to 90, heard 28.5 after it starts and needing to be played 19.9 early; trumpet, forte on A3: an attack family of 30 milliseconds against a pitch floor of 18, so the instrument is, shortened by the dynamic to 27, heard 8.6 after it starts and needing to be played 0.0 early. The spread is 19.9 milliseconds, which is well above the two or three a listener resolves, so a conductor asking for these four to sound together is asking for four different physical onsets.
Fig. 2 The corrected map: which parts have to be early, and by how much, for a scored chord. The violin is earliest because its sound arrives latest, and the piano is nearly as early for a completely different reason.

Two spreads, and only one is a convergence

There is a modelling decision behind the figure that is worth exposing, because getting it wrong makes a converged ensemble look like a stalled one.

The spread that a listener would measure is the spread of what actually arrives, and it includes every player’s motor noise on that beat. That quantity floors at whatever the noise gives — around ten milliseconds for four players with four milliseconds of noise each — and it never gets below it, however perfectly the ensemble has learnt.

The spread that is converging is the spread of the players’ intentions: where each of them is aiming, which is what a rehearsal changes and what a player could report. That falls to zero, or rather to a residue set by how hard the players are correcting.

Plotting the first as though it were the second says an ensemble stops improving after two beats, which is false. Plotting only the second says the ensemble becomes perfect, which is also false. The figure plots the intentions and reports the noise floor beside them, and the honest summary is that the ensemble’s aim converges and its execution does not.

That distinction is the same one the deviations are not noise draws on the microtiming ladder: a performance’s timing has a systematic part and a random part, and every measurement of the first has to get past the second.

Where the beats actually fall. Measured timing deviations from a strict grid, in milliseconds, for three published profiles. The right-hand column converts each deviation into the note value it would have to be written as, at three tempi — and because a fixed number of milliseconds is a different fraction of the beat at every tempo, no single notated rhythm describes any of these.
Fig. 3 The other account’s version of the same separation: what is systematic in a performance’s timing and what is not. The convergence in this essay is a claim about the first and the floor it settles on is the second.

There is a best gain, and it is where players are

The correction gain is the one quantity in this model that belongs to the players rather than to the instruments, and sweeping it produces a genuine optimum.

gain beats to agree within 5 ms settled spread
0.05 21 1.5 ms
0.10 11 2.2
0.15 7 2.8
0.20 6 3.2
0.30 4 4.1
0.40 4 4.9
0.50 never 5.8
0.70 never 7.6

Two things move in opposite directions. A stronger correction converges faster and settles worse, because a player who moves a long way in response to each observation is also moving a long way in response to each observation’s noise.

Past a gain of about 0.4 the settled spread exceeds five milliseconds and the ensemble never gets inside it at all: the players are chasing each other’s tremble faster than they are converging on anything.

The best compromise is between 0.15 and 0.4, and the gains measured for real duos on the microtiming ladder are between 0.2 and 0.5. That is not a prediction that has been tested — the measured range comes from a different experimental setup and a different task — but the overlap is close enough to be worth stating.

How many beats it takes, against how hard the players correct. The correction gain is the one quantity in this model that belongs to the players rather than to the instruments, and an earlier essay on microtiming measures it at between a fifth and a half for real duos. Over that range the ensemble finds its asynchrony in 6 beats at worst and 4 at best, which is inside the first phrase. A weak corrector takes much longer and a very strong one is not much faster, because the limit at high gain is the motor noise rather than the correction: past about 0.30 the players are chasing each other's tremble.
Fig. 4 How many beats it takes against how hard the players correct. Weak correction is slow; strong correction never settles, because past a certain gain the players are following each other’s motor noise.

What a bar and a quarter means for a rehearsal

The headline number is small and it is worth saying what it does and does not describe.

Five beats is fast. An ensemble meeting a new chord finds its asynchrony inside a bar and a half, without discussion, from listening alone. That is consistent with the fact that ensembles do not rehearse asynchrony explicitly and generally do not talk about it.

It is also per chord. The map is a property of the scoring, so a passage in which every chord has a different instrumentation, register and dynamic has a different fixed point at every chord — and five beats is longer than most chords last. An ensemble playing a fast passage never converges on anything; it is always chasing.

That is the useful consequence and it points somewhere specific — and “never converges on anything” is a sentence the same model can be made to check, by moving the target and measuring what is left.

Give the four players a second chord: the same instruments at different registers and dynamics, which is a map of 4.6, 17.1, 24.0 and 10.2 milliseconds against the first chord’s 20.5, 21.8, 28.5 and 8.6. Alternate the two every N beats and let the ensemble chase.

chord lasts spread of what is heard of the available correction
never changes 3.68 ms 100%
16 beats 5.48 89%
8 beats 7.34 77%
4 beats 9.26 65%
2 beats 10.34 58%
1 beat 10.70 56%
no correction at all 19.67 0%

Chasing is worth more than half of catching. At one chord a beat — the fastest the target can possibly move — the ensemble still ends up at 10.7 milliseconds against the 19.7 it would have with nobody listening, which is fifty-six per cent of the whole available improvement. A chord lasting a bar gets sixty-five per cent.

The curve also saturates in the wrong place for the earlier claim. Between four beats and one it moves by only 1.4 milliseconds, so the penalty for a fast passage is nearly all paid by the time a chord is down to a bar, and shortening it further costs almost nothing.

The reason is that the correction is a low-pass filter with a time constant of several beats, so an ensemble tracking a moving target settles near the average of the targets rather than nowhere. And the two maps share most of their shape — the violin is latest in both and the trumpet earliest in both — so an average of them is a good answer to each.

So the honest version is narrower than the one above. The convergence model works for a sustained texture and degrades gracefully for a moving one, and what a fast passage loses is the last forty per cent rather than the whole of it.

Which suggests that what an ensemble actually learns in rehearsal is not a set of leads. It is a map from scoring to leads — an internalised version of the fifth rung’s figure — applied on the fly, with the listening correction handling the residual. That is a claim about expertise rather than about timing and it is not something this model can test.

What the spread would be without listening

A useful control is the ensemble that does not correct at all — four players each starting on the beat, as written.

That ensemble’s spread is the map’s own: 19.9 milliseconds, permanently. A note is heard after it starts puts a listener’s resolution for an asynchrony between notes of similar attack at a few milliseconds, and rather more between notes of different attack — but not twenty. A twenty-millisecond spread is comfortably audible and is heard as a chord that does not quite arrive together.

So the correction is doing something a listener would notice. An ensemble of four playing exactly on the beat sounds ragged, and the raggedness is entirely a property of the instruments rather than of the players.

There is a version of that observation which is much older than any of this. Orchestral players routinely describe the experience of playing “with” a section and being told they are late or early, and the instruction to a brass or a wind section to anticipate is common enough to be a cliché. What the model adds is that the amount is computable and that nobody has to be told it — the ensemble finds it by listening, if the texture stays still long enough.

Two notes started together, heard 32 ms apart. Two amplitude envelopes rising from the same instant: a piano with an 8 millisecond attack and a sung vowel with 110. The horizontal line is the criterion — 6 dB below peak, from Vos & Rasch 1981 — and the two dots are where each envelope crosses it. Nothing about the onsets differs; the heard moments differ by 32 ms, which is why the sung vowel has to start early to be heard on the beat. The buttons play the pair as written and then with the piano delayed by that amount.
Fig. 5 The earlier figure: the spread an ensemble carries from its attack families alone, with every player at one pitch and one dynamic. That is the quantity the correction has to remove, and it is a property of the instrumentation.

Which computation produced the numbers

The lags come from the fifth rung’s scoring model: each part’s rise time from its attack family, floored by four cycles of its own pitch and shortened by its dynamic, converted to a perceptual-centre lag by the relative criterion the whole ladder uses.

The dynamics are the microtiming ladder’s: each player’s next onset is their current one minus a gain times the difference between their own heard sound and the mean of the others’, with Gaussian motor noise added to each sound.

Two spreads are available and only one of them converges. The spread of what actually arrives includes each player’s motor noise on that beat and therefore floors at whatever the noise gives, however well the ensemble has learnt. The spread of the players’ intentions converges to zero and is what is plotted, with the noisy spread reported beside it.

Every curve is a mean over eighty runs with different noise seeds.

Where the model stops

Nobody is leading. Every player here corrects toward the mean of the others with the same gain, which is a democratic ensemble. Real ensembles have a leader, a conductor, or a part everybody follows, and two players and no clock is explicit that its measure cannot say which player is following.

The motor noise is Gaussian and independent. A real player’s timing errors are correlated across beats and correlated with the music, and a quantity resting against a wall is the rung about how badly a boundary distorts what is recovered from them.

There is no tempo. The model has beats and no duration between them, so nothing here says whether the convergence is faster at a fast tempo. It probably is in beats and is not in seconds — and the beat has a preferred rate is the collection’s evidence that a beat is not an arbitrary unit.

And the lags are computed at one criterion. A perceptual centre is the moment an envelope reaches a stated fraction of its own peak, and how late is a different note is about how much the choice of fraction moves everything downstream.

And the correction is symmetric. Each player can move earlier or later without limit, which the walled-duet rung shows is not true when a player is already at a physical limit.

What the picture cannot show

It cannot show that any ensemble does this. The convergence is what the model does; whether players’ asynchronies converge on the perceptual-centre map in rehearsal is a measurement on performances, and it is one that could be made.

Nor can it show the listener. Five milliseconds is well inside what a listener resolves for notes of similar attack and well outside it for notes of different attack, and which case applies here is precisely the thing being converged on.

And it cannot show the score. The map is one chord’s. A real passage’s leads change from chord to chord and the model has no way to represent a player anticipating the change rather than reacting to it — which is what a good ensemble is doing.

How much earlier an accent is heard, by mechanismAn accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller.24.0 ms05101520051015202530decibels louder than the notes around itmilliseconds earlier the accent is heardcriterion at the note's own peakon an unchanging envelope: nothingthe same criterion, on theshorter rise a loud note hascriterion at a fixed levelboth, which is what aninstrument actually does
Fig. 6 One of the three terms behind every lag: what a change of dynamic does to a note’s rise and therefore to when it is heard. A passage in which the dynamics move is a passage in which the fixed point moves, which is the case this model cannot follow.

Whose ensembles, and when

The scoring is a mixed one — piano, flute, violin and trumpet — chosen because its attack families span the range. That mixture is a chamber-music or orchestral texture of the nineteenth century onwards, and the second rung of this ladder argues that a heterogeneous ensemble has a spread a homogeneous one does not.

The correction gains come from experimental work on duos tapping or playing simple material, mostly since the 1990s. Applying them to a four-part chord is an extrapolation in two directions at once — more players, and real music — and the model’s own arithmetic says the number of players matters: correcting toward the mean of three others is a different observation from correcting toward one.

The one historical claim worth making is the one the second rung already makes. A homogeneous ensemble — a string quartet, a consort of viols, a choir — has almost no spread to converge on, because its instruments share an attack family. The mixed ensemble that needs this correction is a specific development, and it acquired a conductor at about the point it became standard.

Where this ladder goes next

Six rungs. A note is heard after it starts; an ensemble that mixes attack families carries a spread; the dynamic moves each attack; the pitch puts a floor under all of it; the three added, for a scoring, as a map; and now the map found by an ensemble that was never told it, in five beats.

What is owed after this is the moving target, and the section above has taken the first step: chasing recovers more than half of what catching would, so the model degrades rather than failing. What is left is the other forty per cent, and the fix for it is not a faster gain but a different kind of player. An ensemble that has internalised the map applies it rather than discovering it, which is a feedforward correction where this is a feedback one, and the two are distinguishable in a performance: a feedback correction lags the change of scoring by several beats and a feedforward one does not. That is a measurement on a recording rather than an arithmetic, and it is the first thing this anchor has wanted that a corpus would answer.

Part 6 of 9

One essay in the series on Perceptual-centre. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AsynchronyAttack transientCorrection gainEnsemble asynchronyEntrainmentMotor noisePerceptual-centreRehearsal