Which notes have to be played early
Assumes: A low note cannot start on time · The players who have to be early
A note is heard some milliseconds after it starts, and this ladder has now found three separate things that decide how many.
The players who have to be early is the attack family: a marimba is heard three milliseconds after it starts and a sung vowel a hundred and ten, so an ensemble mixing the two carries a spread built into its instrumentation. Playing louder is playing earlier is the dynamic: a louder note has a shorter rise, so an accent is heard earlier as well as louder. A low note cannot start on time is the period: a note cannot establish an amplitude inside a cycle, so four cycles at forty hertz is a hundred milliseconds and the attack has a floor that rises as the pitch falls.
Every figure in all three varies one and holds the others. A real passage varies all three at once, and the fourth rung ended by naming the case where that matters: a sforzando low note on a piano against a quiet high one on a flute has the dynamic effect and the pitch effect pulling in opposite directions on one of the two instruments.
How the three terms compose
They are not three additive corrections and it matters what order they go in.
The pitch imposes a floor. A note’s amplitude cannot be established faster than a few of its own cycles, whatever the instrument. That is a hard bound and not a contribution: an instrument whose attack is faster than the floor simply gets the floor.
The instrument supplies a rise, which is what the attack family is, and it competes with the floor rather than adding to it. Whichever is larger is the one the note actually has.
The dynamic shortens whichever survived. A louder blow, bow or breath reaches its regime sooner, and the exponent differs by family — a quarter for a struck note, where the felt’s nonlinearity is measured, and an asserted fifteen hundredths for bowed and blown.
So the composition is max(family, floor) and then a level term on the result. Written that way it is obvious why the terms can cancel, and it is not obvious at all from three curves each drawn with the others held still.
The case the fourth rung named
The piano’s bottom E is 41 hertz. Four cycles of it is ninety-seven milliseconds, against the piano’s own attack of eight — so the note is pitch-limited by a factor of twelve, and everything the excitation ladder knows about how fast a hammer leaves a string is irrelevant to when the note is heard.
Play it sforzando and the felt’s nonlinearity shortens the rise to sixty-five milliseconds. The pitch pushed it to ninety-seven and the dynamic pulled it back to sixty-five, and what is left is a note that needs to be played twelve milliseconds early rather than the twenty-two the pitch alone would have demanded — which is to say the accent has moved the piano from being the earliest part of the chord to the third of four.
Meanwhile the quiet high flute has almost no floor — five milliseconds at A5 — so its attack is entirely its instrument’s sixty, and playing it quietly lengthens that to sixty-nine. The dynamic and the pitch are pulling the same way on the flute and opposite ways on the piano, in one chord.
Flattening the dynamics makes it worse
The check that the three terms interact rather than add is to flatten one of them and see which way the total moves.
Set every part to the same dynamic and the spread of perceptual centres goes from twenty milliseconds to twenty-one. Removing a source of variation made the total variation larger, which is not possible if the terms add and is exactly what happens if one of them is cancelling another.
Put every part at the same pitch and the spread goes to twenty-seven, because the pitch floor had been helping: the instrument with the longest attack happened to be at a pitch where the floor was irrelevant, and the one with the shortest was at a pitch where the floor was enormous. Put every part on the same instrument and the spread falls to seventeen, which is the only one of the three that behaves the way an additive model predicts.
The asserted exponent, swept
One of the three terms carries a number with no measurement behind it: the level exponent is a quarter for a struck note, where the felt’s nonlinearity has been measured, and an asserted fifteen hundredths for bowed and blown. Turning it from zero to four tenths says how much of this rung depends on it, and the answer separates cleanly into two halves.
| exponent | chord A’s spread | order, earliest first | chord B’s spread |
|---|---|---|---|
| 0.00 | 19.0 ms | violin, piano, flute, trumpet | 31.5 ms |
| 0.15, as asserted | 19.0 | violin, piano, flute, trumpet | 31.5 |
| 0.20 | 19.0 | violin, flute, piano, trumpet | 31.5 |
| 0.30 | 19.0 | violin, flute, piano, trumpet | 31.5 |
| 0.40 | 20.6 | flute, violin, piano, trumpet | 31.5 |
The spread does not move at all across the plausible range, and the reason is structural rather than lucky: the earliest and latest parts of this chord — the violin and the trumpet — are both at the reference dynamic, so the level term multiplies their rises by one whatever the exponent is. A parameter that only acts on the interior of a range cannot change the range.
What it does change is the interior, and that is where this rung’s most quoted observation lives. The piano and the flute are early by within a millisecond of each other, and which of them is earlier flips at an exponent of 0.20 — a third above the asserted value, which is well inside the uncertainty of a number nobody has measured. So the sentence about two parts arriving at the same lead for unrelated reasons is robust and the sentence about the order of those two is not.
The cancellation survives everything, and the reason is worth stating because it moves the result onto firmer ground. Flattening the dynamics makes chord A’s spread larger at every exponent tried, including at zero — where the bowed and blown parts have no level term at all. The cancellation is therefore being produced entirely by the piano’s struck exponent, which is the one with measurements behind it, and the asserted parameter is not carrying the finding.
That is a better outcome than the sweep going the other way, and it is worth recording as such: of the three terms, the pitch floor is arithmetic, the family attacks are published, the struck exponent is measured, and the one number that is asserted turns out to be responsible for nothing in this rung except the order of two middle parts.
What a conductor is being asked for
Twenty milliseconds is a long time. A listener resolves an asynchrony of two or three between two notes of similar attack, so a chord whose parts have a twenty-millisecond spread of perceptual centres is a chord that will not sound together unless somebody is early.
The map says who, and by how much. The violin is earliest by twenty milliseconds, which is the answer an intuition about attack families would give — a bowed string has the slowest onset of the four and its sound therefore arrives last unless it starts first. The trumpet is on the beat, because at A3 its own thirty-millisecond attack is well above its floor and the forte shortens it to twenty-seven, which is the quickest onset in the chord.
The surprise is in the middle of the list. The piano is twelve milliseconds early — three quarters of the way to the violin — and everybody would name it as the instrument that plays last, because it is a struck instrument with an eight-millisecond attack. Its pitch has taken that attack away from it. The quiet flute is thirteen milliseconds early for a completely different reason: its family attack is long and the quiet dynamic lengthens it further.
That is the practical content of the rung. Which player has to be early is a property of the scoring rather than of the instrumentation, and it changes when the dynamics change even if the notes do not: two of these four parts are early by within a millisecond of each other, and the two reasons have nothing in common.
A second scoring, where nothing cancels
The cancellation in the first figure is a property of that chord, and the honest way to show it is to change the chord.
Score the same four instruments so that the low note is the quiet one and the high note is the accented one — which is a perfectly ordinary thing to write — and the two terms now push the same way on both. The pitch floor is long at the bottom and the quiet dynamic lengthens it further; the floor is nothing at the top and the accent shortens the instrument’s own attack.
The spread goes from twenty milliseconds to thirty-one, and flattening the dynamics now behaves the way an additive model predicts. So the terms interact in a direction the scoring decides, and neither the cancelling case nor the reinforcing one is more typical than the other.
Why this is a map and not a curve
The word is worth defending, because the ladder’s four earlier rungs all produced curves and this produces a table.
A curve is what a model gives when one variable moves and the rest are fixed. Four of those exist here and they are all correct. A map is what it gives when the variables move together in the pattern a real object has — and the pattern is supplied by the scoring, which is a piece of music rather than a parameter.
That is the same distinction the fingerboard map draws for a violin: two curves that had each been drawn alone, multiplied over the space a player actually works in, with a worst place on the result. Both cases have the same moral, which is that the interesting structure is in the product and neither factor’s own figure can contain it.
Which computation produced the numbers
The attack families are PC_SOURCES, the second rung’s own table of rise times, each a published mid-range figure with a range attached — a marimba at three milliseconds, a sung vowel at a hundred and ten.
The floor is attackFloorMs: four cycles of the note’s own fundamental. The number of cycles is asserted and is the fourth rung’s, and it is the parameter that rung says its whole result rests on.
The dynamic term is RISE_LEVEL_EXPONENT: the rise goes as amplitude to a negative power, one quarter for struck and fifteen hundredths for bowed and blown. The struck exponent is measured — it is Hall and Askenfelt’s felt-compression figure, the same one the excitation ladder uses — and the other is asserted, which the third rung records at length.
The lag itself is pCentre with the relative criterion: the moment the envelope reaches a stated fraction of its own peak. The criterion is the ladder’s own default and every figure in it uses the same one, so the numbers here are comparable with the numbers there.
What twenty milliseconds is next to
It is worth putting the number beside the other times this collection works in, because a millisecond figure is easy to quote and hard to feel.
Twenty milliseconds is about a sixtieth of a crotchet at a moderate tempo. It is well above the two or three at which an asynchrony between two similar notes becomes audible, and well above the three to six milliseconds of standard deviation that separate one player’s groove from another’s. It is the same order as the deliberate microtiming that makes a style.
Which means the spread is not a subtlety to be tidied away. It is as large as the expressive timing being laid on top of it, so a performance’s microtiming and its perceptual-centre corrections are quantities of one size, and any measurement of the first that does not model the second is measuring their sum.
That is a real methodological consequence for the microtiming ladder, which measures deviations from a notated grid and attributes them to style. A mixed ensemble’s deviations contain a component that is not style at all: it is four players finding the asynchrony that makes them sound together.
Where the model stops
The envelope is one exponential. A real attack has a noise burst, an overshoot and a settling, and the ladder’s first rung says so. What is modelled is a rise to a peak, and everything downstream is a criterion on that rise.
Four cycles is an assertion. The whole pitch term scales with it. Three cycles or six would move every floor proportionally and would move which parts are pitch-limited — the piano would still be, the trumpet at A3 would be borderline.
Every part is one note and no part is a line. A player who has been sounding does not have an attack at all, so the whole apparatus applies to entries and to changes of note, and a sustained part passing through a chord contributes nothing to its spread. The map is of a simultaneous attack, which is the case a conductor’s downbeat is about and is not most of what an ensemble does.
And the parts are independent. A real ensemble’s players are listening to each other, which is what two players and no clock is about: they correct toward one another with a gain, so the spread that survives is not the spread the physics imposes. The map is the problem, not the solution.
What the picture cannot show
It cannot show that anyone plays this way. The leads are computed from a model of when a note is heard. Whether players actually place their onsets there is a measurement on performances, and the ladder’s second rung is careful to say that the published evidence is about ensembles that mix attack families rather than about individual corrections.
Nor can it show the room. Every part arrives at the listener at a different time because the players are in different places, which for an orchestra is tens of milliseconds — the same order as everything computed here and from an entirely different cause.
And it cannot show what happens to a passage. One chord is one instant. A passage has every part changing pitch and dynamic continuously, so each player’s required lead is a function of time, and playing it would mean a tempo that differs by part.
Whose ensembles, and when
The four parts chosen are a modern orchestral mixture and the numbers are for it. The claim generalises in a way that has a historical edge: an ensemble of one family — a consort of viols, a gamelan, a string quartet — has almost no spread from the family term, so its remaining spread is entirely the pitch and the dynamic, which are both under the players’ control.
A mixed ensemble does not have that option. The attack families are what they are, and the only way to bring the parts together is for somebody to be early. That is a real difference between a homogeneous ensemble and a heterogeneous one, and it is a plausible part of why the former needs no conductor and the latter acquired one at about the moment the mixture became standard.
Where this ladder goes next
Five rungs. A note is heard after it starts; an ensemble that mixes attack families carries a spread; the dynamic moves each attack through a range; the pitch puts a floor under all of it; and now the three added, for a scoring, as a map.
What is owed after this is the correction. This ladder’s own fourth rung on the microtiming side — two players and no clock — has a model of players adjusting toward one another with a gain, and the map above is exactly the input that model wants: a set of required leads that the players do not know and have to discover by listening. Running one into the other would say how long an ensemble takes to converge on the right asynchrony, in bars, which is a quantity every rehearsal is spending and nobody has computed.
Part 5 of 9
One essay in the series on Perceptual-centre. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 9.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Attack transientDynamicsEnsemble asynchronyOnsetOrchestrationPerceptual-centreRegisterTiming deviation
- A note is heard after it starts attack transient, ensemble asynchrony, onset, perceptual-centre, timing deviation
- Read at two different heights attack transient, ensemble asynchrony, onset, perceptual-centre
- A part entering is not a change of level dynamics, onset, orchestration
- The blend arrives before the note does attack transient, onset, perceptual-centre
- What the tongue actually removes attack transient, ensemble asynchrony, onset
- A blown note does not start late, it starts slowly attack transient, onset