Instruments and their design

Two players and no clock

Two people cannot share a timekeeper, and two independent ones drift a hundred and twenty milliseconds apart inside a minute. Ensembles do not, so something is correcting — and the measurement everybody reaches for recovers the pair's total responsiveness exactly and cannot tell which of the two is doing it. Four tenths from one player and two tenths each give the identical number.

Assumes: Swing is a ratio, and it is not two to one

One player producing three against two turned out to be running one timekeeper rather than two, and the argument was that two independent clocks drift apart without bound while one clock cannot drift at all.

Two players cannot use that solution. There is no shared timekeeper to be had; each has their own, in their own head, and the drift argument applies to them in full.

Two players, and the correction that keeps them together. The spread of the asynchrony between two players, in milliseconds, against beat number, for 3 correction gains, averaged over 120 seeded runs each. It reaches 120 ms after 64 beats at a gain of 0, 27 ms after 64 beats at a gain of 0.1, 20 ms after 64 beats at a gain of 0.3. With no correction at all the asynchrony is a random walk and grows without bound; with any correction it settles at a fixed spread within a few beats and stays there. Two people cannot share a timekeeper, so the fact that ensembles do not come apart is itself the evidence that they are correcting.
Fig. 1 The spread of the asynchrony between two players over sixty-four beats, for three amounts of mutual correction. With none at all it reaches 120 milliseconds and is still growing. With a gain of 0.1 it settles at 27 and stays there; with 0.3, at 20.

So the fact that ensembles do not come apart is itself the evidence. Something is pulling the two players together, beat by beat, and the something has a size.

The correction, and what it does to the asynchrony

The standard model is one line per player. Each hears the asynchrony between their own last onset and the other’s, and removes a fraction of it from their own next interval.

Call those fractions α for the first player and β for the second. Both zero is two independent clocks. Anything positive is correction, and correction turns a random walk into a stationary process — the asynchrony still varies, and it no longer accumulates.

That is a qualitative change rather than an improvement. With no correction the spread after sixty-four beats is 120 milliseconds; with a gain of 0.1 it is 26 after four beats and 27 after sixty-four. The bound arrives almost immediately and then nothing further happens.

Two players, and the correction that keeps them together. The spread of the asynchrony between two players, in milliseconds, against beat number, for 4 correction gains, averaged over 120 seeded runs each. It reaches 120 ms after 64 beats at a gain of 0, 36 ms after 64 beats at a gain of 0.05, 22 ms after 64 beats at a gain of 0.2, 20 ms after 64 beats at a gain of 0.5. With no correction at all the asynchrony is a random walk and grows without bound; with any correction it settles at a fixed spread within a few beats and stays there. Two people cannot share a timekeeper, so the fact that ensembles do not come apart is itself the evidence that they are correcting.
Fig. 2 Four gains including a very small one. Even a gain of 0.05 — removing a twentieth of the error each beat — is enough to bound the spread, at around 40 milliseconds rather than 120 and rising. The difference between correcting and not correcting is much larger than the difference between correcting a little and correcting a lot: from 0.1 to 0.5, a fivefold change in gain, the spread moves from 27 to 20.

Which explains something about ensemble playing that is otherwise puzzling: why it is possible at all with quite poor listening. A player who barely responds still holds together with one who does, because the gains add.

It also explains why an ensemble can hold a tempo that no member is holding. Correction bounds the difference between the players and says nothing about where the pair as a whole ends up, so a duo drifts in tempo together while staying together — which is the ensemble version of the drift one performer’s two hands cannot have, and is why a long unaccompanied passage ends at a different tempo from the one it started at.

The measurement, and what it recovers

The quantity a study can measure is the asynchrony series itself, and the standard thing to compute from it is its lag-one autocorrelation.

There is an exact identity available. In the absence of motor noise, that autocorrelation is

1 − α − β.

So a measured autocorrelation of 0.6 says the two gains add to 0.4, which is a genuine and useful recovery of something invisible: how responsive the pair is, as a pair, in a unit that can be compared across duos.

Five ways of dividing one total, and one number for all of them. The lag-one autocorrelation of the asynchrony between two players, for five ways of splitting a total correction gain of 0.4 between them — one player doing all of it, through to an even share. The values are 0.578, 0.578, 0.578, 0.578, 0.578, a spread of 0.000. The autocorrelation is 1 minus the sum of the two gains, so it recovers the pair's total responsiveness exactly and says nothing whatever about who is following whom.
Fig. 3 Five ways of splitting a total gain of 0.4 between two players — one doing all of it, through to an even share. Every one of them gives an autocorrelation of 0.578, a spread of zero. The measurement recovers the sum and is blind to the division.

And there the recovery stops. The autocorrelation is a function of α + β and of nothing else, so it cannot distinguish a duo in which one player is doing all the correcting from one in which both are doing half — which is exactly the question the word leader names, and exactly the question anybody measuring a duo wants answered.

The limitation is worse than a limitation of one statistic, and the reason is one line of algebra. Writing e for the asynchrony, the model says

e next equals e now times (1 − α − β), plus noise.

The two gains appear only as their sum, in the dynamics of the asynchrony itself. So it is not that the autocorrelation happens to miss the split: no statistic of the asynchrony series can recover it, because the series is generated by a process in which the split does not appear. The mean is the same, the spread is the same, every lag of the autocorrelation is the same, and running the simulation confirms it — a 0.4-and-0 duo and a 0.2-and-0.2 duo produce an asynchrony with the same mean to two decimal places and the same spread to one.

Why that matters more than it sounds

A leader is not a metaphor in ensemble playing. It is a specific claim: that one part is being followed and the other is being tracked, that the responsibility is unequal, and that swapping the roles changes the result.

The autocorrelation gives a single number for the pair. Two duos with the same number can be organised completely differently — one a soloist and an accompanist, the other two equals — and no amount of care taken over the measurement will separate them, because the quantity does not contain the information.

Recovering the split needs a different measurement: the cross-correlation between each player’s own interval changes and the asynchrony that preceded them. A player who is correcting shows that dependence and a player who is not does not, and the two can be computed separately. That is more data and more analysis, and it is what the studies that do report leadership actually do.

Two players, and the correction that keeps them together. The spread of the asynchrony between two players, in milliseconds, against beat number, for 2 correction gains, averaged over 120 seeded runs each. It reaches 161 ms after 96 beats at a gain of 0, 24 ms after 96 beats at a gain of 0.3. With no correction at all the asynchrony is a random walk and grows without bound; with any correction it settles at a fixed spread within a few beats and stays there. Two people cannot share a timekeeper, so the fact that ensembles do not come apart is itself the evidence that they are correcting.
Fig. 4 The same comparison run half as long again. The uncorrected spread has reached 161 milliseconds by beat ninety-six and the corrected one is within a few milliseconds of where it was at beat four. What separates them is the growth, which is a property of the series over time — and it is the one property a single-lag autocorrelation is not designed to see.

The thing this rung expected to find, and did not

The slate for this essay said that the asynchrony’s lag-one autocorrelation is negative under mutual correction, and that a shared-clock model predicts zero — so the sign would be the signature.

It is not. At the correction gains people are actually measured at, roughly 0.1 to 0.4 each, the autocorrelation is comfortably positive: 0.58 for a total of 0.4, 0.44 for a total of 0.4 with motor noise added, 0.22 for a total of 0.6. It goes negative only when the two gains together exceed 1, which is over-correction — a duo that removes more than the whole error each beat, and oscillates.

This is the second time in this phase that a lag-one autocorrelation has been reached for as a sign test and has failed to be one. The single-performer case was the first, where both competing models turned out to give negative values for the same reason and the sign said nothing about the question.

The pattern is worth naming. An autocorrelation is an estimate of a parameter, not a test of a structure. Used as a parameter estimate — α + β from the asynchrony, the motor share from the intervals — it is exact and valuable. Used as a yes-or-no signature of a mechanism, it has now been wrong twice.

What would recover it

The split is not unrecoverable in principle; it is unrecoverable from the asynchrony alone. What carries it is each player’s own intervals, which is a different series.

A player who corrects has intervals that depend on the asynchrony that preceded them, and one who does not has intervals that do not. So the cross-correlation between player A’s interval changes and the previous asynchrony estimates α directly, and the same computation on B estimates β. Two numbers, from two series, and the sum of them can then be checked against the asynchrony’s autocorrelation as a consistency test.

Two players, and the correction that keeps them together. The spread of the asynchrony between two players, in milliseconds, against beat number, for 2 correction gains, averaged over 120 seeded runs each. It reaches 36 ms after 64 beats at a gain of 0.05, 20 ms after 64 beats at a gain of 0.3. With no correction at all the asynchrony is a random walk and grows without bound; with any correction it settles at a fixed spread within a few beats and stays there. Two people cannot share a timekeeper, so the fact that ensembles do not come apart is itself the evidence that they are correcting.
Fig. 5 Two duos with very different total responsiveness, which the asynchrony does distinguish: 0.05 each settles at around 40 milliseconds and 0.3 each at 20. What the asynchrony cannot say about either of them is which player is providing the gain.

That is more work than computing one autocorrelation, and it needs each part’s onsets separately rather than a mixed recording. Which is presumably why the single number gets quoted, and it is worth knowing what it is a number for.

What is left when correction is subtracted

Correction bounds the asynchrony; it does not remove it. The bound in the figures sits between 20 and 34 milliseconds depending on the gain, which is the region a real duo occupies.

That residual is not error in any useful sense. It is the sum of the motor noise the players cannot suppress and the correction lag they cannot beat, and it sets a floor under any offset one player carries relative to the other. An intended offset of thirty milliseconds and a residual asynchrony of twenty are the same size, and separating them in a recording requires averaging over many beats.

Two kinds of systematic timing, which share a word. Each measured profile split into a constant offset from the grid and a pattern that varies by position in the bar. Viennese waltz, second beat is −10.0 ms of offset and 33.4 ms of pattern; jazz soloist against the ride is 28.8 ms of offset and 1.9 ms of pattern; quantised is 0.0 ms of offset and 0.0 ms of pattern. A motor deviation anywhere in the 8 to 20 ms range published for skilled performers leaves 74–95% of Viennese waltz, second beat's variation systematic, 1–5% of jazz soloist against the ride's variation systematic. The two quantities are independent and no single deviation figure distinguishes them.
Fig. 6 The two measured profiles again, split into offset and pattern, with the motor band at the foot. A duo’s residual asynchrony sits in the same range as that band and as the offsets themselves, which is why an offset is a statistical claim about many bars rather than an observation about one.

The room is the size of the effect

The omission that most obviously matters is that the model has the two players hearing each other instantly.

Sound travels about 343 metres a second, so ten metres of separation is 29 milliseconds each way. That is the same size as the residual asynchrony, the same size as a measured jazz offset, and larger than the motor floor.

Distance is a term in the model too, and it arrives long before a room’s reverberant field overtakes the direct sound. At ten metres apart two players hear each other 29 milliseconds late — a whole intended offset’s worth of lag, introduced by geometry alone and by nothing either of them is doing.

The consequence is that a correction gain measured in a laboratory with headphones is not the gain operating in a hall, and that ensembles across a large stage are correcting toward a version of each other that is a beat’s fraction out of date. Orchestral players know this and deal with it by watching rather than listening, which is a solution the model has no term for at all.

There is a second acoustic term that is larger still. In a reverberant room an onset does not arrive at an instant; it arrives smeared, and the early reflections are within the window in which the ear is deciding when the note began. A model whose input is an onset time is assuming that question is settled.

The gain is not free either

A large correction gain looks like an unmixed good in the figures — 0.5 holds tighter than 0.1 — and it is not.

A player correcting hard is a player whose intervals are being set by somebody else, which means their own timing intentions are being overwritten. An intended offset or a bar-position pattern is exactly the kind of thing a high gain removes: the correction cannot tell an expressive deviation from an error, so it pulls both back toward the other part.

That predicts a trade with a shape. A duo with high gains is tight and flat; one with low gains is loose and expressive; and the middle is where playing together is supposed to happen. Whether measured gains actually sit in a middle, and whether they drop in passages where expression is wanted, is the obvious study and is not something this site can settle.

Two players, and the correction that keeps them together. The spread of the asynchrony between two players, in milliseconds, against beat number, for 4 correction gains, averaged over 120 seeded runs each. It reaches 27 ms after 64 beats at a gain of 0.1, 20 ms after 64 beats at a gain of 0.3, 21 ms after 64 beats at a gain of 0.6, 39 ms after 64 beats at a gain of 0.9. With no correction at all the asynchrony is a random walk and grows without bound; with any correction it settles at a fixed spread within a few beats and stays there. Two people cannot share a timekeeper, so the fact that ensembles do not come apart is itself the evidence that they are correcting.
Fig. 7 Four gains, up to a very high one. Past about 0.5 the spread stops improving and begins to worsen: a pair that removes nearly the whole error each beat overshoots, and each correction produces the error the next one corrects. Tight playing has an optimum and it is not at the top.

Whose ensembles, and what the model leaves out

Every gain in this essay is from tapping and duo-performance studies, mostly of Western trained musicians, mostly finger-tapping or piano duets, mostly since 1990. The model is the standard one in that literature and it is a model of two people.

Three things it does not have are worth naming because they are what a real ensemble is made of.

It has no ears. The model corrects toward the other player’s onset as though it were known exactly and instantly. In a real room the sound takes time to arrive — a few metres of separation is several milliseconds — and the asynchrony a player perceives is not the asynchrony a microphone at the centre records.

It has no score. Both players are producing an isochronous beat. Real parts have rests, different densities and different registers, and a player who is not playing cannot correct. Nor does it have a metre: where the beat is at all is assumed settled and shared, when in a difficult passage it is exactly what the two players may disagree about.

And it has no more than two players. An orchestra is not a duo scaled up. Nor is a duo playing three against two between them the same problem as one player doing it, and this model is the right one for the first while the single-timekeeper account is the right one for the second. With many players correcting toward a common something — a conductor, a section leader, the loudest part — the dynamics are different, and the question of who is following whom becomes a network rather than a pair.

What the picture cannot show

It cannot show a leader. That is the substance of this rung and it is worth repeating as a limitation as well as a finding: every figure here is symmetric or is drawn as though the asymmetry were invisible, because to the measurement it is.

It cannot show what the players are doing about it. The offset and the pattern a performer produces are the deliberate half of ensemble timing and the correction is the involuntary half, and every figure here has only the second. A duo in which one player intends to sit behind and the other intends to be exactly on the beat is a duo with a target asynchrony, which this model does not have a term for either.

It cannot show intention. A player who deliberately sits behind the beat and a player who is correcting slowly produce similar traces. The first is an offset and the second is a low gain, and telling them apart requires the offset’s mean to be distinguishable from zero over many bars.

It cannot show that the gain is constant. α is a fixed number in every figure. Real correction gains change with tempo, with familiarity, with how loud the other part is, and with whether the passage is difficult. A single number for a whole performance is an average over something that varies.

It cannot show the room. Sound arriving late is the largest omission in the model and it has a size that is easy to compute and hard to include: at ten metres apart, thirty milliseconds each way, which is the same size as everything else in this essay.

It cannot show what happens when one player stops. A part with rests cannot correct and cannot be corrected toward, so the effective gains drop to zero for that stretch and the asynchrony resumes its random walk. That is a computable consequence of this model and it takes one line: with no correction, two players’ timings accumulate independent noise, so the asynchrony after n beats has a standard deviation of σ√(2n).

Set σ at the motor band’s own 4 to 8 milliseconds and ask how long before the asynchrony reaches the twenty this essay’s own bound sits at:

motor noise reaches 20 ms reaches 30 ms
4 ms 12.5 beats 28 beats
6 ms 5.6 12.5
8 ms 3.1 7.0

A duo with the tightest plausible motor floor has about three bars of common time before its silence costs it as much as its correction was buying, and one at the loose end has under one. At a hundred beats a minute those are seven seconds and two.

Which is short enough to be a fact about repertoire rather than a curiosity. It is the reason a passage where one player rests for four bars is a passage that needs a cue, and the reason continuo parts are written to keep sounding rather than to keep quiet.

And the correction is only about phase. Both players hold the same period throughout. A real duo also corrects its tempo toward the other, which is the second gain a beat tracker has, and a model with both has behaviours neither has alone.

Three players is not two players and a half

Every figure here has two lines in it, and an ensemble does not.

With three players each correcting toward the other two, the asynchronies are no longer one series but three, and they are not independent — the triangle closes, so the third is determined by the other two. The system is stable for a wide range of gains and its behaviour is not a scaled version of the duo’s.

The interesting case is the one every tradition has arrived at independently: a common reference that everybody corrects toward and that corrects toward nobody. A conductor, a timekeeper, a bell, a click. That reduces an n-player network to n independent one-way corrections, each of which is the tracker problem rather than the duet problem, and it is why a bell pattern is played by one person and never varies.

The trade is visible in the arithmetic. Mutual correction keeps a pair together and lets the pair’s tempo wander, because nothing anchors the average. A one-way reference anchors the tempo absolutely and gives up the ability to breathe with the players — which is exactly the complaint about playing to a click, and it is a structural consequence rather than a matter of taste.

A beat kept through the bars that do not mark it. A 16-step pattern played for 6 bars, with every onset removed from the beat itself in the middle two, and seeded timing noise on every onset. The line is how far the tracker's beat sits from where it started, in milliseconds of a 560 ms beat. It coasts through beats 9 to 16, keeps its period, and comes out at most 23 ms from true — 4 per cent of a beat. Handed the silenced bars on their own, the preference rules move the downbeat from phase 2 to phase 0.
Fig. 8 The one-way case: a tracker following a reference that does not respond to it. This is what playing to a click is, and it is the model from the essays on metre rather than the one in this essay — one gain instead of two, and no possibility of the reference moving.

The ladder from here

The residual asynchrony is twenty to thirty milliseconds and an intended offset is thirty. Neither is heard as a wrong note, and something must be making the difference between a note that is late and a note that is a different note.

There is an answer and it is the same shape as the answer in the pitch domain: the deviation lives inside a category, and the category has an edge. The edges in time turn out to be arithmetic, and the swing ratio walks across two of them.

There is a third kind of systematic timing and it is the one the next rung is about: the swing ratio, which falls with tempo because the short note holds a roughly constant hundred milliseconds. Everything in this essay has been measured in milliseconds and a ratio is dimensionless, and that difference turns out to decide whether a deviation stays inside a category or leaves it.

The other thing left open here is the network. Two is the case the literature has solved and the case every figure on this page draws; three is where the interesting structure starts, and an orchestra is neither.

Part 4 of 9

One essay in the series on microtiming. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 11.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

EntrainmentMicrotimingMotor delayRandom walkTimekeeperTiming deviation