Series

Perceptual-centre — the series

9 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. One envelope, and the three places a listener might be said to hear it. The amplitude envelope of a note with a 90 millisecond exponential attack, with the three criteria the literature offers drawn across it. The heard moment is 8.0 ms at the detection criterion, 28 ms at the perceptual-onset criterion and 94 ms at the perceptual-attack criterion. The physical onset is at zero on this axis and no criterion puts the heard moment there. The buttons play this attack against a two-millisecond one, started at the same instant.

    A note is heard after it starts

    Every rhythm essay until now has treated a note's onset as the moment it happens. It is not: the instant a listener aligns a note with a beat is later than its physical start by an amount the note's own attack decides, and for a sung or bowed note that amount is about thirty milliseconds — the size of the whole quantity six essays on microtiming set out to measure.

    part 1 · rhythm
  2. Which ensembles have this problem and which do not. The width of the heard-moment spread built into 7 standard instrumentations, before any player does anything. An ensemble drawn from one attack family has a spread of zero — every note is heard the same distance after it is started, so a common onset is a common heard moment, and this is true of a string quartet and of a gamelan for the same reason and in the same amount. A mixed ensemble carries between 11 and 33 milliseconds of it.

    The players who have to be early

    Ensembles have been measured for fifty years and found to be about forty milliseconds out of alignment, which has always been reported as the limit of human precision. Part of it is not: an ensemble that mixes attack families carries a heard-moment spread of ten to thirty-three milliseconds before anybody plays a note, and an ensemble drawn from one family carries none at all — which is true of a string quartet and of a gamelan for the same reason.

    part 2 · instruments
  3. How much earlier an accent is heard, by mechanism. An accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller.

    Playing louder is playing earlier

    An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.

    part 3 · perception
  4. The heard moment against the pitch, on an instrument whose own attack is 8 ms. A note cannot establish an amplitude in less than 4 of its own cycles, so the attack has a floor of 4 periods — 145 milliseconds at A0 and 1.9 at C7. Below A4 the floor is longer than the instrument's own attack and the pitch decides the heard moment; above it the instrument does. The lag runs from 46.0 milliseconds at the bottom to 2.5 at the top, a spread of 44 milliseconds that no player can play their way out of.

    A low note cannot start on time

    Three earlier essays have held the pitch at one value. A note cannot establish an amplitude in less than a few of its own cycles, so the attack has a floor that rises as the pitch falls — 146 milliseconds at the bottom of a piano and three at the top. On an instrument whose action takes eight milliseconds everywhere, that is a forty-three millisecond spread across the keyboard from the period alone, and no player can do anything about it.

    part 4 · rhythm
  5. Which notes of a scored chord have to be played early. Four parts of one chord, each with its own instrument, its own pitch and its own dynamic, and the perceptual centre that comes out of all three. piano, sforzando on E1: an attack family of 8 milliseconds against a pitch floor of 97, so the pitch is what limits it, shortened by the dynamic to 65, heard 20.5 after it starts and needing to be played 12.0 early; flute, quiet on A5: an attack family of 60 milliseconds against a pitch floor of 5, so the instrument is, shortened by the dynamic to 69, heard 21.8 after it starts and needing to be played 13.2 early; violin, mezzo forte on E4: an attack family of 90 milliseconds against a pitch floor of 12, so the instrument is, shortened by the dynamic to 90, heard 28.5 after it starts and needing to be played 19.9 early; trumpet, forte on A3: an attack family of 30 milliseconds against a pitch floor of 18, so the instrument is, shortened by the dynamic to 27, heard 8.6 after it starts and needing to be played 0.0 early. The spread is 19.9 milliseconds, which is well above the two or three a listener resolves, so a conductor asking for these four to sound together is asking for four different physical onsets.

    Which notes have to be played early

    There are three separate contributions to one quantity — the instrument's attack family, the dynamic it is played at, and the note's own period — and every figure so far varies one and holds the others. Added together for a real scoring they do not add: a sforzando low piano note is pitch-limited to a hundred-millisecond attack and the sforzando shortens it back to sixty-five, so flattening the dynamics makes the ensemble's spread larger rather than smaller.

    part 5 · rhythm
  6. An ensemble finding an asynchrony nobody told it about. An earlier essay produced a map of required leads — which notes of a scoring have to be played early, and by how much — and nothing tells the players those numbers, because they are a property of the instruments' attacks rather than of the music. So an ensemble has to find them, and the mechanism is already here: each player hears sounds rather than onsets and moves their next onset toward the mean of the others'. The spread of arrival times starts at 20 milliseconds and settles at 4, crossing 5 milliseconds after 5 beats — about 1.3 bars of four. The leads it converges on match that map to within 0.2 milliseconds, which is what makes this a convergence rather than a coincidence: the fixed point of players listening to each other is every player leading by their own attack.

    How many bars an ensemble needs

    The map of required leads is something nobody tells the players, because the leads are a property of the instruments' attacks. So an ensemble has to find them, and the mechanism is the one the microtiming essays describe: each player hears sounds rather than onsets and moves toward the others. It converges on the map to within a fifth of a millisecond, in five beats, and there is a best correction gain.

    part 6 · rhythm
  7. The passage that separates them, and a listener cannot hear it. A scoring changes at the halfway bar, and the two maps of required leads differ by 18.1 milliseconds at their widest. An ensemble that has internalised the map applies the new one on the first note of it and its spread never leaves zero. An ensemble that is listening to each other has to re-converge: its spread jumps to 11.8 milliseconds and takes 3 beats to get back under 5. The dashed line is twenty milliseconds, which is what a listener notices — and the disagreement never reaches it. So the two accounts are separable on a recording and very nearly not separable by ear, which is why nobody has noticed the distinction and why the measurement is worth making.

    The passage that separates two players

    An ensemble that has learnt where the asynchronies are applies them; one that is listening discovers them. In steady state the two are identical, which is why nobody has separated them. Change the scoring mid-phrase and they are not: one ensemble is wrong by twelve milliseconds for three beats and the other is not wrong at all — and twelve milliseconds is under what a listener notices and far above what a microphone resolves.

    part 7 · rhythm
  8. 14 players summed, against the one at their average onset. Each thin line is one player's rising envelope, started at its own moment, with a spread of 30 milliseconds about the beat and a 90-millisecond attack. The heavy line is the section: nominally identical sources add incoherently, so their powers add and the sum is the root of the mean of their squares, drawn here as a fraction of the section's own peak. The dashed line is the single player who started at the section's average onset. The section reaches the 6 dB below peak criterion at 18.1 milliseconds and that player at 27.6, a difference of 9.5. The section is early because the players who started first are already sounding while the average one is still building, and nothing a late player does can make the sum quieter.

    Twelve violins are more punctual than one

    Every essay until now treats a part as one player, and an orchestral part is a dozen. Sectioning does two things at once and only one of them was expected: it pulls the part's heard moment forward, by four milliseconds against a map spanning twenty-six, and it makes the part's arrival more accurate by very nearly the root of the number of players. So the map of required leads applies to an orchestra better than it applies to a quartet, and the case where it fails is three trumpets rather than fourteen violins.

    part 8 · rhythm
  9. One attack time, three shapes, 48 ms of disagreement. Three amplitude envelopes with the same 90-millisecond attack time, which is the only quantity the published tables report. a resonator from a step rises as one minus a decaying exponential; an excitation ramping rises in a straight line; a ramp through a resonator is a raised cosine. Each is normalised so that its own 10-to-90 per cent rise takes exactly 90 milliseconds, so all three are the same measurement. The horizontal rules are the three criteria a heard moment is read off. At 6 dB below peak the three shapes put the heard moment at 28.5, 56.4, 76.3 milliseconds — a spread of 48, on one attack time, from a property nothing in the table records.

    An attack time is not an attack

    Eight earlier essays read a heard moment off an envelope, and every one of them used the same envelope shape without saying so: the source table records one curve for all nine of its families, and the map's own arithmetic does not carry the parameter at all. A published attack time fixes a ten-to-ninety time and nothing else. Under the two other shapes the same measurement admits, every millisecond computed so far doubles — and the constructed passage called inaudible earlier becomes three times a listener's threshold.

    part 9 · perception

All series