Rhythm and metre

Twelve violins are more punctual than one

Every essay until now treats a part as one player, and an orchestral part is a dozen. Sectioning does two things at once and only one of them was expected: it pulls the part's heard moment forward, by four milliseconds against a map spanning twenty-six, and it makes the part's arrival more accurate by very nearly the root of the number of players. So the map of required leads applies to an orchestra better than it applies to a quartet, and the case where it fails is three trumpets rather than fourteen violins.

Assumes: The passage that separates two players · What a choir does that a soloist cannot

The passage that separates two players ended by naming what every rung of this anchor has quietly assumed:

Every rung of this anchor treats a part as one player, and an orchestral part is a dozen of them playing the same line — so a section has an internal spread of its own, and the map’s leads are of the same size as it.

The worry in that sentence is precise and it is easy to state. The map of required leads asks a violin to start about nineteen milliseconds ahead of a trumpet. Fourteen violinists do not start at the same instant as each other, and if their disagreement is of the same size as the lead, then telling the section to be nineteen milliseconds early is telling it to do something finer than it can do. The map would then be a description of a chamber group and of nothing an orchestra has ever played.

The arithmetic for this is already here in two halves. This ladder has the criterion; the choir rung has what happens when many nominally identical sources are added together.

14 players summed, against the one at their average onset. Each thin line is one player's rising envelope, started at its own moment, with a spread of 30 milliseconds about the beat and a 90-millisecond attack. The heavy line is the section: nominally identical sources add incoherently, so their powers add and the sum is the root of the mean of their squares, drawn here as a fraction of the section's own peak. The dashed line is the single player who started at the section's average onset. The section reaches the 6 dB below peak criterion at 18.1 milliseconds and that player at 27.6, a difference of 9.5. The section is early because the players who started first are already sounding while the average one is still building, and nothing a late player does can make the sum quieter.
Fig. 1 Fourteen violinists on one note, each starting at their own moment, with a spread of thirty milliseconds about the beat and the family’s ninety-millisecond attack. The heavy line is what the section sounds like: nominally identical sources add incoherently, so their powers add and the sum is the root of the mean of their squares. The dashed line is the single player who started at the section’s average onset. The section crosses the criterion first, because the players who started early are already sounding while the average one is still building.

A section is a sum and not an average

The step that decides everything below is what “the section” means as a signal, and there are two candidates that give different answers.

If a section were an average of its players, its onset would be the mean onset and the whole question would be about how accurately a dozen people can agree on a mean. That is the intuition behind the debt, and it is wrong for the same reason the choir rung’s intuition about smoothness was wrong: nothing is averaging. The listener receives a sum.

Fourteen violins playing one note are fourteen tones within a few tens of cents of each other. Their phases are unrelated, so they add incoherently, which means their powers add rather than their amplitudes — twenty players are twelve decibels above one and not twenty-six. The section’s envelope is therefore the root of the sum of its players’ squared envelopes, and the section’s peak is the root of the count times one player’s.

Now put this ladder’s criterion on that. A perceptual centre is placed where the envelope reaches a fixed fraction of its own peak, and the section’s own peak is n\sqrt{n} times a player’s. So the section is heard when the root-mean-square of its fourteen amplitudes reaches the criterion — and that is not when the average player reaches it, because the players who started early have already got there and the ones who started late contribute nothing yet and cannot take anything away.

The consequence is one-directional and it is worth stating before any number. Scatter can only bring a section forward. Adding a late player adds energy at every instant or none; it can never make the sum quieter than it would have been. There is no arrangement of onsets that makes a section later than one player of the same attack starting at the section’s own mean.

That is the first half of the answer, and by itself it says the debt was right to worry: the section’s heard moment moves, and it moves in a direction that changes the map.

How far it moves, and how little the count matters

The size is the question, and the size turns out to depend on almost nothing anybody would have guessed.

A section of 14 is heard 4.1 ms early and arrives 3.5 times as accurately. A part played by 14 instruments instead of one, against how far apart the players start. The solid curve is when the section is heard, relative to its own mean onset; the horizontal rule at 28.5 milliseconds is where one player of the same 90-millisecond attack would be heard. The dashed curve is how far the section's arrival moves from note to note. At 30 milliseconds of scatter — the asynchrony at which the fusion model used here stops hearing several sources as one — the section is heard 4.1 milliseconds early and its arrival moves by 8.5 against a single player's 30. The bias is what a section's scatter costs the map of required leads and the scatter is what it buys, and over the range a section can actually have, the second is much the larger.
Fig. 2 A section of fourteen violins against how far apart its players start. The solid curve is when the section is heard, relative to its own mean onset; the rule at 28.5 milliseconds is where one violinist of the same ninety-millisecond attack would be heard. The dashed curve is how far the section’s arrival moves from note to note. The vertical rule is thirty milliseconds, which is the asynchrony at which the model of fusion used here stops hearing several sources as one thing.

At four milliseconds of scatter the section is heard where one player would be, to a hundredth of a millisecond. At twelve it has moved by 0.2. At twenty, by 1.3. At thirty — which is as scattered as a section can be and still be a section, for a reason the next section gives — it has moved by 4.1 milliseconds, against a map that spans twenty-six.

The curve is flat and then it is not, and the reason is a ratio rather than a threshold. The bias becomes appreciable when the scatter becomes comparable with the lag itself, and the lag here is 28.5 milliseconds. A section whose players agree to a tenth of their own attack time is a section whose heard moment is exactly one player’s.

Now the part that is genuinely surprising, and it is the reason this rung exists rather than merely confirming the map.

The scatter falls with the players and the bias does not. Two quantities against the number of instruments on one part, at 30 milliseconds of scatter between them and a 90-millisecond attack. The dashed curve is how far the section's arrival moves from note to note: it falls as the root of the count, from 31.1 milliseconds at 1 to 5.4 at 32, and the open marks behind it are the root law drawn from the single player's figure. The solid curve is how much earlier the section is heard than one player of the same attack, and it is flat: 3.7 milliseconds at 2 players and 4.3 at 32. Adding players to a part buys accuracy and does not change where the part is heard, which is why the map of required leads survives being scored for an orchestra.
Fig. 3 The two quantities against the number of instruments on one part, at thirty milliseconds of scatter and a ninety-millisecond attack. The dashed curve is how far the section’s arrival moves from note to note, and the open marks behind it are one player’s figure divided by the root of the count. The solid curve is how much earlier the section is heard than one player, and it is flat.

The bias does not depend on the number of players. Two violins are pulled forward by 3.7 milliseconds and thirty-two by 4.4. That is a null, and it has a reason: the crossing time is a property of the distribution the onsets are drawn from rather than of the sample, and two draws locate that distribution about as well as thirty-two do for this particular statistic. A section and a duet are pulled forward by the same amount.

The scatter does depend on it, and it falls as the root of the count. One player’s arrival moves by 29.5 milliseconds from note to note; a section of twelve moves by 8.8, and one of thirty-two by 5.6. The ratios are 3.4 and 5.4 against roots of 3.46 and 5.66. That agreement is a result and not an identity — nothing in the model is averaging anything, and the crossing of a summed envelope had no obligation to behave like a mean.

The scatter has a ceiling, and this collection set it

Everything above is a function of one number nobody has measured for an orchestral section, so the honest thing is to bound it rather than to assert it, and the bound is already in the collection.

A section is a section because a listener hears it as one thing. What makes two partials one note puts the onset asynchrony at which sources stop fusing at about thirty milliseconds: past that, the ear separates what it is hearing into more than one event. A string section scattered by more than thirty milliseconds is not a section with a smeared attack — it is fourteen violinists a listener can hear are not together, which is a thing every conductor stops the rehearsal for.

So thirty milliseconds is not an estimate of a section’s scatter. It is a ceiling on it, imposed by the same ear the whole ladder is about, and the bias at the ceiling is 4.1 milliseconds.

At the other end, the model this ladder already uses for players adjusting to one another carries four milliseconds of motor noise per player — the figure the convergence rung runs on — and at four milliseconds the bias is nothing at all.

The real number is somewhere between, and the whole range is small against a map spanning twenty-six milliseconds. That is the answer to the debt, and it is a stronger answer than a measurement would have been at this stage, because it does not depend on which value inside the range is right.

The map of required leads, for sections rather than soloists. Each part of an orchestral scoring, with the lead the earlier map asks of it as a soloist (the open mark) and the lead it needs when the part is played by the number of instruments beside its name (the filled one), at 4 milliseconds of scatter inside each section. The bar behind each is how far that section's own arrival moves from note to note. The soloist map spans 25.9 milliseconds and the section map 25.9. The order does not change, and no part moves by more than 0.0 milliseconds.
Fig. 4 The map of required leads for an orchestral tutti, with every part played by the number of instruments beside its name, at the four milliseconds of scatter the ensemble model used here gives a player. The open mark is the lead the part needs as a soloist and the filled one is the lead it needs as a section; the bar through each is how far that section’s own arrival moves from note to note. The two maps are the same map.

Which inverts the debt

Put the two halves together and the conclusion runs the other way from the question.

A soloist asked to play nineteen milliseconds early is asked to hit a target whose own note-to-note scatter, at any plausible motor noise, is a good fraction of the lead. A section of fourteen asked the same thing is asked to hit it with an arrival that moves by a third as much. The lead is more reliably deliverable by a section than by a player, and the part with the largest lead in an orchestral scoring — a bowed string, whose attack is the slowest — is precisely the part that is played by the largest section.

That is a coincidence of the instrument rather than a design, and it is worth naming as one. The families with the longest attacks are the bowed and sung ones; the families with the most players on a part are the strings and the chorus. The orchestra’s most demanding leads are asked of its most accurate timekeepers, and nobody arranged that.

It also puts the sixth and seventh rungs on firmer ground rather than weaker. The convergence model has each player moving toward the mean of what they hear; if a section were a smear, the thing being converged on would be ill-defined and the model would describe nothing. It is not a smear. A section is a source with a better defined arrival than any of its members, which is exactly the object that model needs.

Where it does break, and it is not where anybody looked

There is one case in which sectioning genuinely changes the map, and it is the opposite of the case the debt named.

Scoring the map for sections changes which part is heard first. Each part of an orchestral scoring, with the lead the earlier map asks of it as a soloist (the open mark) and the lead it needs when the part is played by the number of instruments beside its name (the filled one), at 30 milliseconds of scatter inside each section. The bar behind each is how far that section's own arrival moves from note to note. The soloist map spans 25.9 milliseconds and the section map 29.4. The order changes: a small section with a large scatter is dominated by whichever of its players was earliest, so trumpets moves to the front.
Fig. 5 The same tutti at the ceiling: thirty milliseconds of scatter inside every section, which is as far apart as players can be and still be heard as one thing. The order changes. Three trumpets are pulled forward by 12.1 milliseconds and two flutes by 6.8, while fourteen violins move by 4.3 — so the small sections overtake, and the part heard first is no longer the percussion.

The pull is largest for the smallest section above one. Three trumpets at the ceiling are heard twelve milliseconds early and fourteen violins four, and the reason is the criterion rather than the players.

The criterion is a fraction of the section’s own peak, and three players’ peak is 3\sqrt{3} times one player’s — about 1.73. Half of that is 0.87, which one trumpeter alone reaches while still nearly alone. So a section of three crosses its own criterion at roughly the moment its earliest member has spoken, and the earliest of three is well ahead of the average of three. A section of fourteen has a peak of 14\sqrt{14}, and half of that is 1.87 — more than any one player can supply, so the crossing waits for several and lands much nearer the middle of the distribution.

A larger section is therefore pulled forward less, not more. That is the reverse of every intuition about smearing, and it falls directly out of a relative criterion: a bigger section sets itself a higher bar.

The consequence for an orchestra is specific. The parts at risk are the ones with two, three or four players — the wind desks — and they are also the parts with the shortest attacks and the smallest leads to begin with. At the ceiling the map’s order changes and the trumpets arrive first. Whether that ever happens is a question about how ragged a wind section is, and it is a question a recording could answer in an afternoon.

Which notes of a scored chord have to be played early. Four parts of one chord, each with its own instrument, its own pitch and its own dynamic, and the perceptual centre that comes out of all three. percussion on C4: an attack family of 3 milliseconds against a pitch floor of 15, so the pitch is what limits it, shortened by the dynamic to 15, heard 4.8 after it starts and needing to be played 0.0 early; trumpets on A4: an attack family of 30 milliseconds against a pitch floor of 9, so the instrument is, shortened by the dynamic to 30, heard 9.5 after it starts and needing to be played 4.7 early; flutes on A5: an attack family of 60 milliseconds against a pitch floor of 5, so the instrument is, shortened by the dynamic to 60, heard 19.0 after it starts and needing to be played 14.2 early; first violins on E5: an attack family of 90 milliseconds against a pitch floor of 6, so the instrument is, shortened by the dynamic to 90, heard 28.5 after it starts and needing to be played 23.6 early; double basses on E1: an attack family of 90 milliseconds against a pitch floor of 97, so the pitch is what limits it, shortened by the dynamic to 97, heard 30.7 after it starts and needing to be played 25.9 early. The spread is 25.9 milliseconds, which is well above the two or three a listener resolves, so a conductor asking for these four to sound together is asking for four different physical onsets.
Fig. 6 The same tutti as it would have been computed earlier, with every part a soloist. Three terms set each lead: the family’s attack, the floor the note’s own period imposes, and the dynamic. The double basses are the latest not because a bass is bowed more slowly but because a forty-hertz note cannot establish an amplitude in less than four of its own cycles.
Twenty singers are not twenty times one. The level of n sources against n, on the two assumptions. Incoherent sources add in power and gain 3 dB per doubling, so 32 voices are 15 dB above one — about 6 times the pressure and rather less than that in loudness. Sources in phase would add in amplitude and reach 30 dB, which no choir does and no ensemble has ever needed to. The gap between the two curves is the whole reason a section of twenty exists rather than a soloist told to sing louder.
Fig. 7 The addition this whole essay rests on, from the essays on choirs. Doubling the players on a part adds three decibels and not six, because near-identical sources arrive in unrelated phases and their powers add. That is what makes a section’s peak the root of its count, and the root is where the criterion — and therefore the whole of the bias above — comes from.

Which computation produced the numbers

Each player is the same envelope this ladder has used since its first rung: an amplitude approaching its peak with a stated ten-to-ninety per cent attack time, started at that player’s own onset. The onsets are drawn from a normal distribution about the beat with the stated standard deviation, and the attack times may be given a spread of their own.

The section is the root of the sum of the players’ squared envelopes, which is incoherent addition, and the criterion is six decibels below the section’s own peak — the same fraction, from the same source, that every figure on this anchor uses.

Two quantities are reported. The bias is the section’s crossing relative to the mean of its own players’ onsets, averaged over several hundred independent draws, and it is compared with the closed-form lag of a single player. The arrival scatter is the standard deviation of the section’s crossing measured from the beat, across the same draws, which is what a microphone on the section would record.

The map is the fifth rung’s pCentreScoring, unchanged, with each part’s soloist lag shifted by the bias its own section size and scatter give it, and the leads re-referred to whichever part is then heard first.

Where the model stops

The scatter is a parameter and not a measurement. Nothing here is a published figure for how far apart an orchestral section starts. What the rung supplies is a ceiling from the fusion threshold and a floor from the ensemble model’s own motor noise, and the finding is that the answer is small across the whole range between them.

The players are independent. A real section watches a leader and each other, so its onsets are correlated rather than drawn afresh each note, and correlation would reduce the effective scatter and therefore the bias. The model’s number is the pessimistic one.

Incoherent addition is an idealisation. Fourteen violins in a hall are not fourteen uncorrelated sources at a listener’s ear; they are on a stage, at different distances, and the room adds its own arrivals. Coherent addition would raise the section’s peak faster and pull the crossing earlier still.

The envelope is one shape. Every rise here approaches its peak the way a resonator driven from rest does, and that is an assumption this anchor has never examined.

A section plays one dynamic. Each player has the same attack time here, or the same distribution of them. In a real section the outside desks play differently from the inside ones, which is a systematic spread rather than a random one — and a louder attack is a shorter one, so a dynamic gradient across a section is an attack gradient too.

And nothing decays. These envelopes rise and hold. A struck section — a percussion group, or pizzicato strings — is a set of notes that rise and immediately fall, and a sum of decaying envelopes started at different moments has a peak that depends on the scatter, which this arithmetic does not have. Partials that do not die together is the nearest thing the collection has to that case and it is about one instrument.

What the picture cannot show

It cannot show a divided section. Half the violins on one line and half on another is two sections of seven, and every number above changes with the count.

It cannot show the leader. A section’s first desk is louder and earlier by convention, and a section with one player weighted above the rest is nearer a soloist than the sum of equals drawn here.

Nor can it show the conductor. A beat given in advance is a reference the whole section shares, and it is the obvious way a section’s scatter is kept below the ceiling in the first place.

It cannot show the hall. The scatter a listener receives is the players’ scatter plus the spread of their distances to the seat, which in a large orchestra is several metres and therefore ten milliseconds or more of pure geometry.

Nor can it show the bowing. Fourteen players changing bow together is a coordinated event with its own timing, and a section’s raggedness is not a random variable so much as a rehearsed one.

And it cannot show whether any of this is heard. Four milliseconds is a fifth of what a listener needs to notice an asynchrony, so the bias computed here is a correction to a map rather than an audible event — which is the same situation the seventh rung found itself in, and for the same reason.

Whose playing, and when

The attack times are laboratory measurements of ordinary orchestral playing, and the fusion threshold is a psychoacoustic measurement made on synthetic tones. Neither is a measurement of a section.

The section sizes are the nineteenth-century orchestra’s: fourteen first violins, eight basses, three trumpets, two flutes. That distribution is a historical object rather than an acoustic one — it was arrived at by balancing loudness in halls of a certain size — and it happens to put the most players on the parts with the longest attacks, which is why the map survives being scored for it. An ensemble balanced differently would not automatically get the same answer, and a baroque orchestra with three players a part is nearer the case where the map changes than the case where it does not.

Where this ladder goes next

Eight rungs. A note is heard after it starts; an ensemble that mixes attack families carries a spread; the dynamic moves each attack; the pitch puts a floor under all of it; the three added for a scoring, as a map; the map found by an ensemble that was never told it; the passage that would say whether an ensemble finds it or already has it; and now the map scored for sections, which survives, and which survives better for an orchestra than for a quartet.

What the ladder owes now is the shape of the rise. Every envelope on this page and on the seven before it approaches its peak the way a resonator driven by a step does, and nothing anywhere in the anchor says so or offers an alternative — the source table records one shape for all nine of its families, and the map’s own arithmetic does not carry the parameter at all. That is the last variable in the model nobody has varied, and it can be varied with no new measurement: the published attack times fix a ten-to-ninety time and say nothing about the curve between. Whether the ladder’s milliseconds are properties of the instruments or of that curve is one afternoon’s arithmetic, and it is owed before any of the leads above are quoted again.

Part 8 of 9

One essay in the series on Perceptual-centre. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientEnsemble timingIncoherent additionOnsetOrchestrationPerceptual-centreUnison