Rhythm and metre

The passage that separates two players

An ensemble that has learnt where the asynchronies are applies them; one that is listening discovers them. In steady state the two are identical, which is why nobody has separated them. Change the scoring mid-phrase and they are not: one ensemble is wrong by twelve milliseconds for three beats and the other is not wrong at all — and twelve milliseconds is under what a listener notices and far above what a microphone resolves.

Assumes: How many bars an ensemble needs · Which notes have to be played early

How many bars an ensemble needs put the fifth rung’s map of required leads into a model of players adjusting toward each other, and found that an ensemble which is told nothing converges on it in about five beats. Its last paragraph names the alternative that model cannot be:

An ensemble that has internalised the map applies it rather than discovering it, which is a feedforward correction where this is a feedback one, and the two are distinguishable in a performance.

In steady state they are not distinguishable at all. Both end up with every player leading by their own attack; both produce the same arrival times; both look exactly like an ensemble playing well together. That is why the distinction has never been made — there is nothing to see.

Change the scoring and there is.

The passage that separates them, and a listener cannot hear it. A scoring changes at the halfway bar, and the two maps of required leads differ by 18.1 milliseconds at their widest. An ensemble that has internalised the map applies the new one on the first note of it and its spread never leaves zero. An ensemble that is listening to each other has to re-converge: its spread jumps to 11.8 milliseconds and takes 3 beats to get back under 5. The dashed line is twenty milliseconds, which is what a listener notices — and the disagreement never reaches it. So the two accounts are separable on a recording and very nearly not separable by ear, which is why nobody has noticed the distinction and why the measurement is worth making.
Fig. 1 A scoring changes at the halfway bar. One ensemble applies the new map on the first note of it; the other has to find it again.

The construction

Take two scorings whose maps of required leads differ. A marimba, a trumpet and a violin is one; a sforzando piano in the bass, a quiet flute at the top and a voice in the middle is another. Their maps differ by 18.1 milliseconds at the widest, which is the fifth rung’s arithmetic and not a choice.

Write a phrase that plays twelve beats of the first and then switches, without warning, to the second.

A feedforward ensemble — players who know both maps — applies the second one on the first note of the second scoring. Its spread of arrival times is zero throughout, before and after.

A feedback ensemble — the sixth rung’s, each player moving toward the mean of what they hear — cannot. The first note of the new scoring is played with the old map, so it arrives with a spread of 11.8 milliseconds, and the ensemble then re-converges over 11.8, 8.8, 6.6, 5.0, 3.7, 2.8 and so on, reaching five milliseconds after three beats.

So the discriminating statistic is the spread in the beats immediately after the change, and it exists for three or four beats and then closes.

Which a listener cannot hear

The number that decides whether this is a measurement or a phenomenon is 11.8 against 20.

A note is heard after it starts puts the just-noticeable asynchrony for a listener at about twenty milliseconds. The peak disagreement between the two ensembles is under twelve, and it is under twelve for one beat and half of it by the third.

So the two kinds of ensemble sound the same. That is not a disappointment; it is the explanation for why the distinction has never been drawn. Nobody has noticed two kinds of ensemble because there is nothing to notice, and the reason there is nothing to notice is that the map change is smaller than a threshold.

A recording is a different matter. Onset times are extracted from an audio file to a millisecond or better as a matter of routine, and eleven milliseconds is ten times that. The measurement is easy and the perception is impossible, which is an unusual and useful place for a question to sit.

What the two hypotheses actually are

It is worth stating both carefully, because “feedback or feedforward” is a control-theory pair and the musical versions are not quite it.

The feedback hypothesis is that an ensemble’s togetherness is maintained continuously, note by note, by each player adjusting toward what they hear. It has no memory: an ensemble that has played a passage a hundred times behaves exactly as it did on the first read, because the mechanism is a servo and a servo does not learn.

The feedforward hypothesis is that the leads are learnt — in rehearsal, or over a career, or from knowing the instrument — and applied. It has no error signal: the players are not correcting toward each other at all, they are each producing a note early by an amount they know.

Every real ensemble is some of both, and the interesting question is the proportion. What the constructed passage measures is how much of the togetherness survives when the learnt component is made wrong: a purely feedforward ensemble whose map has just become obsolete is worse than a feedback one, for exactly as long as it takes anybody to notice.

That third case is the one the model does not draw and the recording would show. A feedforward ensemble handed a scoring it has not learnt applies the old map and stays wrong — its spread jumps to eleven milliseconds and stays there — which is neither of the two curves above and is the most likely thing to actually happen.

So the measurement has three outcomes rather than two, and they are cleanly separated: a flat zero is a learnt map that transferred, a decay is listening, and a step that does not come down is a learnt map that did not.

What would have to be recorded

The passage has to be constructed rather than found, and the requirements are specific enough to write down.

The scoring must change and the notes must not. If the melody changes at the same moment then the players have a new phrase to read and every timing statistic is contaminated. The cleanest form is the same line handed to different instruments — an orchestral device that exists anyway.

The change must be unrehearsed at least once. A feedback ensemble that has played the passage twenty times is a feedforward ensemble: it has learnt the map. So the measurement is a sight-reading one, and a second take is a different experiment rather than a repetition.

And the ensemble must be small enough to have a spread. With three or four parts the map has a range of tens of milliseconds; with sixty players in sections the spread is dominated by within-section variation, and the effect this looks for is buried — which is what many nominally identical sources do to any statistic and is the same argument the choir rung makes about loudness.

That is a chamber group, sight-reading, with a scoring change written into a phrase, recorded with a close microphone on each part. None of it is exotic and, as far as this collection can tell, none of it has been done.

An ensemble finding an asynchrony nobody told it about. An earlier essay produced a map of required leads — which notes of a scoring have to be played early, and by how much — and nothing tells the players those numbers, because they are a property of the instruments' attacks rather than of the music. So an ensemble has to find them, and the mechanism is already here: each player hears sounds rather than onsets and moves their next onset toward the mean of the others'. The spread of arrival times starts at 20 milliseconds and settles at 4, crossing 5 milliseconds after 5 beats — about 1.3 bars of four. The leads it converges on match that map to within 0.2 milliseconds, which is what makes this a convergence rather than a coincidence: the fixed point of players listening to each other is every player leading by their own attack.
Fig. 2 The earlier convergence, which is the left-hand half of the figure above with nothing before it. The passage constructed here is that experiment run twice in one phrase.

The gain decides how long the window is and not how wide

The one quantity players are known to differ in is how strongly they correct toward each other, and the model has it as a gain.

How long the window lasts, against how fast the players correct. The number of beats a feedback ensemble takes to get its spread back under 5 milliseconds after a change of scoring, against the correction gain — the one quantity players are known to differ in. At a gain of 0.05 it does not get there inside 24 beats; at 0.25, the value chosen earlier, it takes 3 beats; at 0.8 it takes 1. The peak disagreement is the same 11.8 milliseconds at every gain, because it is the map change and the first note after it is always wrong by the whole of it. So the gain decides how long the evidence lasts and not how large it is, and a slow ensemble is easier to catch than a quick one.
Fig. 3 How many beats a feedback ensemble takes to recover, against how fast its players correct.

At a gain of 0.05 the ensemble does not get back inside tolerance within the twenty-four beats of the passage. At 0.25 — the sixth rung’s own value, taken from the microtiming ladder’s measurements of two players adjusting — it takes three beats. At 0.8 it takes one.

The peak disagreement is 11.8 milliseconds at every gain, and it has to be: the first note after the change is played with the wrong map whatever the gain, so the error on that note is the whole map change and nothing about the correction has happened yet.

So a slow ensemble is easier to catch than a quick one, and the evidence for feedback is a decay rather than a level. There is a nice corollary in that: the ensembles hardest to distinguish are the very good ones, because a high correction gain closes the window in a beat — and the ensembles it is easiest to run the experiment on are the ones whose result matters least. That is a better statistic than the peak, because a peak of eleven milliseconds could be a mistake and a clean exponential decay over four beats could not.

One envelope, and the three places a listener might be said to hear itThe amplitude envelope of a note with a 90 millisecond exponential attack, with the three criteria the literature offers drawn across it. The heard moment is 8.0 ms at the detection criterion, 28 ms at the perceptual-onset criterion and 94 ms at the perceptual-attack criterion. The physical onset is at zero on this axis and no criterion puts the heard moment there. The buttons play this attack against a two-millisecond one, started at the same instant.8.0 ms28 ms94 ms05010015020000.20.40.60.81milliseconds after the physical onsetamplitude, as a fraction of the note's own peak15 dB below peak6 dB below peak90% of peak90 ms attack
Fig. 4 One envelope and the three places a listener might be said to hear it, established earlier. Every lead in either map is one of these read off a rise time, and the passage changes which rise times are in play.

And it makes the sixth rung’s convergence a prediction rather than a result

There is a reading of the sixth rung this rung sharpens.

That rung showed that an ensemble can converge on the fifth rung’s map by listening, and treated that as an account of how ensembles come to be together. This one points out that the account has never been tested against the obvious alternative, and that in every steady-state passage the two are identical.

Everything the sixth rung demonstrated is consistent with an ensemble that learnt the map in the first rehearsal and has been applying it since. The convergence model would then be a description of what happened on the first read-through and nothing at all about the performance.

Which of those is true is exactly what the constructed passage would say, and it is the sort of question this ladder is for: a distinction that is invisible in the ordinary case, visible in a designed one, and settled by a measurement anybody could make.

It is also the first thing this anchor has wanted that is not an arithmetic. Six rungs have been computations over published attack times, and this one is a request — for one recording of one constructed passage, sight-read once, with a microphone on each part. That is a small enough ask to be worth stating precisely, and it is stated here as precisely as the model can make it: what to write, how large the map change has to be, how many beats the evidence lasts, and what resolution it needs.

Which notes of a scored chord have to be played early. Four parts of one chord, each with its own instrument, its own pitch and its own dynamic, and the perceptual centre that comes out of all three. piano, sforzando on E1: an attack family of 8 milliseconds against a pitch floor of 97, so the pitch is what limits it, shortened by the dynamic to 65, heard 20.5 after it starts and needing to be played 12.0 early; flute, quiet on A5: an attack family of 60 milliseconds against a pitch floor of 5, so the instrument is, shortened by the dynamic to 69, heard 21.8 after it starts and needing to be played 13.2 early; violin, mezzo forte on E4: an attack family of 90 milliseconds against a pitch floor of 12, so the instrument is, shortened by the dynamic to 90, heard 28.5 after it starts and needing to be played 19.9 early; trumpet, forte on A3: an attack family of 30 milliseconds against a pitch floor of 18, so the instrument is, shortened by the dynamic to 27, heard 8.6 after it starts and needing to be played 0.0 early. The spread is 19.9 milliseconds, which is well above the two or three a listener resolves, so a conductor asking for these four to sound together is asking for four different physical onsets.
Fig. 5 The map itself, drawn earlier: which notes of a scoring have to be played early, and by how much. Two of these, one after the other, is the whole passage.
The three terms do not add. The spread of perceptual centres across one scoring, with each term flattened in turn. All three together give 19.9 milliseconds. no dynamic spread gives 21.2, all at one pitch gives 26.6, all one instrument gives 17.4. Removing the dynamic spread makes the total spread LARGER, which is the finding: on this scoring the sforzando is partly cancelling the pitch floor, so a term that is real on its own is subtracting from the sum. Three curves each computed with the others held fixed cannot show that, and every earlier figure is such a curve.
Fig. 6 Why the two maps differ at all, computed earlier: the three terms that set a part’s lead do not add, so changing an instrument, a register and a dynamic at once moves the map in a way none of them does alone.

Why the map change is as large as it is

The eighteen milliseconds between the two maps is worth a paragraph, because it is what makes the whole measurement possible and it is not obvious that it should be that big.

The three terms that set a part’s lead are the instrument family’s attack, the floor its pitch imposes, and the shortening its dynamic applies. Between the two scorings here all three move on every part: a marimba becomes a low sforzando piano, a trumpet becomes a quiet flute at the top of its range, a violin becomes a voice.

That is a deliberately extreme change and a real orchestral one — it is what happens when a phrase is handed from a bright group to a dark one. A milder change moves the map less and shrinks the effect proportionally, and a change that swaps two instruments of the same family moves it hardly at all.

So the experiment has a design parameter and it is the map change, which the fifth rung’s arithmetic computes for any pair of scorings before anybody plays anything. Choosing the pair is choosing the size of the effect, and the pair here is near the top of what an orchestra does.

How many beats it takes, against how hard the players correct. The correction gain is the one quantity in this model that belongs to the players rather than to the instruments, and an earlier essay on microtiming measures it at between a fifth and a half for real duos. Over that range the ensemble finds its asynchrony in 6 beats at worst and 4 at best, which is inside the first phrase. A weak corrector takes much longer and a very strong one is not much faster, because the limit at high gain is the motor noise rather than the correction: past about 0.30 the players are chasing each other's tremble.
Fig. 7 The earlier gain sweep, for comparison: how many beats a cold start takes at each correction gain. The passage here is the same curve entered halfway through a phrase instead of at a downbeat.

Which computation produced the numbers

The maps are the fifth rung’s pCentreScoring, unchanged: for each part, the larger of the instrument family’s measured attack time and the floor its pitch imposes, shortened by the dynamic, read through a level criterion into a lag.

The feedback ensemble is the sixth rung’s, unchanged: each player’s intention moves toward the mean of what is heard, by the gain, once a beat. The spread reported is of the arrival times, which is what a microphone would record.

The feedforward ensemble applies whichever map is current on the beat, so its spread is zero by construction. That is an idealisation — a real ensemble that knows the map still has motor noise — and the noise floor is the same for both, so it cancels out of the comparison.

The switch is at beat twelve of twenty-four, and the tolerance of five milliseconds is the sixth rung’s own.

Where the model stops

The switch is instantaneous and a phrase is not. Real players see the change coming on the page even when it is not announced, and reading ahead is a span of about four notes — which is a beat or two at most tempi and is exactly the window the effect lives in.

A feedforward ensemble is not perfect. Real players who know a passage still have a few milliseconds of motor noise, which puts a floor under both curves and reduces the contrast. The floor is common to both and does not move the decay.

There is no rehearsal in between. The model has two regimes and one instantaneous switch; a real ensemble that has rehearsed the passage is somewhere between them, with a partly learnt map, and this figure has no term for partly.

The gain is one number for everybody. The microtiming ladder’s own measurements have players differing in it, and an ensemble of mixed gains would converge on something the mean does not describe.

Twenty milliseconds is one threshold of several. The asynchrony a listener notices depends on the material — it is smaller for clicks than for music, and smaller for a duo than for an orchestra — so the claim that eleven milliseconds is inaudible is safe for an ensemble and not for a metronome. The threshold this collection uses is the musical one.

And the maps are computed. Every lead here comes from the fifth rung’s arithmetic rather than from a measurement of an ensemble, which is the standing caveat on this whole anchor.

What the picture cannot show

It cannot show a rehearsal letter. A change of scoring in real music is usually at a structural boundary a player can see coming from three bars away, and anticipation is precisely the thing the two hypotheses differ about. The passage has to hide the change, which is a compositional constraint as much as an experimental one — and which notes have to be played early is the map that has to change without the page announcing it.

It cannot show a low part. A low note cannot start on time puts a floor under a bass instrument’s lead that no amount of learning removes, so a scoring change that moves a part across the register changes what is learnable as well as what is learnt.

It cannot show a conductor. A beat given in advance is a feedforward signal from outside the ensemble, and it is the obvious confound: a conducted ensemble might show no lag at all for reasons that have nothing to do with what the players know.

Nor can it show the third outcome’s own decay. A feedforward ensemble stuck on the wrong map will eventually notice and correct, by some mixture of listening and being told, and how long that takes is a rehearsal timescale rather than a beat one.

Nor can it show which players are wrong. The spread is a summary; the interesting version of the measurement is per part, because a feedback ensemble’s error should be concentrated in whichever instrument’s lead changed most.

It cannot show the listener’s side. Twelve milliseconds is below the threshold for noticing an asynchrony, and whether it is below the threshold for noticing that a passage is slightly less together is a different experiment with a different threshold.

It cannot show the players’ own gains changing. How hard players correct is measured on pairs tapping, and an ensemble that knows it is about to change scoring might correct harder for a few beats — which would look exactly like feedforward and is a third hypothesis this passage does not separate.

And it cannot show a real scoring change. The two scorings here differ in all three parts at once, which maximises the effect; most orchestral changes of colour move one line, and the map change — and therefore the whole measurement — shrinks with them.

Whose playing, and when

The attack times are twentieth-century laboratory measurements of ordinary orchestral playing; the correction gain is from studies of two players tapping together. Neither is a measurement of an ensemble performing.

The device the passage is built around — the same line handed from one instrument to another mid-phrase — is a nineteenth-century orchestral commonplace and is much rarer before it. So the material for this measurement exists in quantity in the repertoire, and it has the wrong property: it is all thoroughly rehearsed. The experiment needs a sight-reading, which is the one condition the recorded repertoire does not supply.

Where this ladder goes next

Seven rungs. A note is heard after it starts; an ensemble that mixes attack families carries a spread; the dynamic moves each attack; the pitch puts a floor under all of it; the three added, for a scoring, as a map; the map found by an ensemble that was never told it; and now the passage that would say whether an ensemble finds it or already has it.

What is owed after this is the section. Every rung of this anchor treats a part as one player, and an orchestral part is a dozen of them playing the same line — so a section has an internal spread of its own, from a dozen instruments with the same attack and different players, and the map’s leads are of the same size as it. Whether a section’s own scatter swallows the lead it is supposed to be taking is an arithmetic this collection has both halves of: the fifth rung’s map, and the choir rung’s account of what happens when many nominally identical sources are added. If it does, the whole map applies to soloists and to nobody else.

Part 7 of 9

One essay in the series on Perceptual-centre. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientEnsemble timingInferenceMicrotimingOnsetOrchestrationPerceptual-centre