The milliseconds that are the groove
Assumes: Swing is a ratio, and it is not two to one · Rhythm is a circle, and the bar line is a choice
A Viennese waltz played by a Viennese orchestra does not have three equal beats. The second beat arrives early — measurably, consistently, and by an amount large enough that a listener who has never been told notices when it is absent. The third beat arrives late to compensate.
The deviation is around fifty milliseconds — a quantity in the same range as the short note of a swung pair, and for related reasons. At a waltz tempo that is roughly a twelfth of a beat, which is nothing any notation has a symbol for and which is not a mistake.
What is being claimed
Two claims are worth separating, because only one of them is contentious.
The uncontentious one is that performances deviate from strict time. Nobody disputes it; every measurement of every performance of anything shows it, and a performance that did not would be identifiable as machine-made in seconds.
The contentious one is that some of those deviations are systematic — that they recur bar after bar, that they are shared between players in an ensemble, that they are specific to a repertoire, and that removing them removes something. That is a much stronger claim, and it is the one the measurements support.
Bengtsson and Gabrielsson’s work on Viennese waltz performance in the early 1980s found the second-beat anticipation to be consistent within orchestras, consistent across recordings, and characteristic of one performing tradition rather than of waltz playing generally. It is a dialect, in the sense a raga’s characteristic movement is.
Which computation produced the number
The left-hand side of that figure plots the published mean deviation of each beat, in milliseconds, against a strict grid. The right-hand columns do the conversion that makes the point.
A deviation of milliseconds at a tempo of beats a minute is a fraction of a beat equal to
and expressing that as a note value means taking its reciprocal. Fifty-five milliseconds at 80 beats a minute is a fraction of — about a fourteenth of a beat. The same 55 milliseconds at 200 beats a minute is — about a fifth.
So the same physical delay is a fourteenth of a beat in one column and a fifth in another. There is no note value that describes it, because it is not a note value: it is a duration, and every symbol in the notation is a proportion.
That is the argument, and it is arithmetic rather than aesthetics. The two quantities have different dimensions.
There is an awkward case for it, and reporting it makes the argument stronger rather than weaker. At the tempo a Viennese waltz is actually played, the deviation is almost exactly one sixth of a beat. A Viennese waltz runs at about sixty bars a minute, which is a hundred and eighty beats; fifty-five milliseconds against a beat of 333 is 0.165, and a sixth is 0.167. The fraction stays within ten per cent of a sixth from 164 to 200 beats a minute, which is essentially the whole tempo range of the form.
So the untranscribable quantity is, at its home tempo, very nearly transcribable — as a triplet subdivision of a triple beat, which is a thing the notation has a symbol for. Anybody wanting to write the Viennese lean down for one tempo could.
That does not rescue the notation and it locates the argument properly. The reason a sixth works is that a hundred and eighty is the tempo at which fifty-five milliseconds happens to land on a simple fraction, and the same fifty-five milliseconds is a ninth of a beat at 120 and a fourteenth at 80. The nearest simple fractions and the tempos they need are a fourth at 273 beats a minute, a fifth at 218, a sixth at 182, an eighth at 136, a twelfth at 91 — a different symbol for every tempo, and no symbol for the quantity.
The jazz profile has no such lucky tempo. Thirty milliseconds is a sixteenth of a beat at 120, a twelfth at 160, a tenth at 200 and an eighth at 240, and none of those is a value anybody writes. So one of the two profiles is coincidentally notatable in the range it is played in and the other is not, which is what a dimensional argument predicts: a coincidence is available at some tempo and is not available at all of them.
The three profiles
The figure carries three, and the third is the control.
The Viennese waltz. Second beat early by about 55 milliseconds, third beat late by about 25. The effect is that the bar’s first two beats are compressed and the last is stretched, which gives the characteristic forward lean.
The jazz soloist against the ride. Ashley’s 2002 measurements of jazz phrasing found soloists sitting consistently behind the drummer’s cymbal by around thirty milliseconds, across the whole phrase rather than at particular points. The lag is close to constant, which is what “laid back” turns out to mean quantitatively: not a rubato, not an accelerando and deceleration, but a steady offset.
The quantised grid. Zero everywhere, which is what a sequencer produces and what no player produces. It is included because a figure of deviations needs a zero line that is a real alternative rather than an abstraction, and because the audible difference between the third row and the other two is the whole subject.
What quantisation removes
Quantisation — snapping every recorded onset to the nearest grid position — is the operation that makes the argument concrete, because it is a lossy transformation that can be applied and undone in a studio and whose effect anyone can hear.
What it removes is exactly the material above. What is left is a rhythm that is correct in every respect a score can specify and that sounds mechanical, and the word “mechanical” is doing precise work: the result is what a mechanism produces, which is a performance with the systematic deviations set to zero.
The standard remedy in sequencing software is a “swing” or “shuffle” percentage, which shifts alternate grid positions by a proportion. That helps a little and fails in the same way the notated swing instruction fails: a proportion is not a duration, so the setting has to be changed whenever the tempo changes, and it applies to one subdivision rather than to the phrase.
More recent tools sample real performances and apply their deviation patterns to programmed material, which is a straightforward acknowledgement that the quantity being copied is a list of milliseconds.
Is thirty milliseconds even audible
The whole argument depends on these deviations being perceptible, so it is worth putting them against the thresholds.
The smallest timing change a listener can detect in an isochronous sequence is around 5 to 10 milliseconds under laboratory conditions, and perhaps 10 to 20 in ordinary listening with real instruments. The threshold for hearing two sounds as non-simultaneous is around 20 milliseconds for most timbres, and considerably less for sharp attacks.
A thirty-millisecond lag is therefore two to three times the detection threshold, and a fifty-five-millisecond anticipation is five times it. Neither is marginal. Both are comfortably above the point at which a listener would report that something had changed, while being far below the point at which either would be described as a different rhythm.
That window — above the threshold for noticing, below the threshold for renaming — is where all of expressive timing lives, and it is a window of a few tens of milliseconds. It is narrow, it is fixed in absolute terms, and it does not move when the tempo does.
Chords are not simultaneous either
There is a second kind of deviation, smaller in intent and larger in size, and it is the one most listeners are most surprised by.
The notes of a chord played by an ensemble, or by a pianist, do not arrive together. The spread between the earliest and latest note of a nominally simultaneous attack is routinely 30 to 50 milliseconds in ensemble playing, and comparable within a single pianist’s hands.
Rudolf Rasch’s measurements of ensemble performance in the late 1970s and 1980s established both the size of the spread and a systematic component within it: the part carrying the melody tends to lead the others by around 20 to 30 milliseconds. The effect is now generally called melody lead, and it is reliable enough to be a candidate explanation for why the melody in a chordal texture is heard as the melody at all.
That is a considerable claim, and it is not fully settled — some of the lead is explained by the mechanics of playing a louder note on a piano, since a harder key press produces an earlier sound, and separating intention from mechanism is exactly as hard here as in the waltz case. What is not in doubt is the magnitude: a chord is a spread of onsets a couple of tens of milliseconds wide, and a genuinely simultaneous chord is something only a sequencer produces.
The consequence for this essay is that the amount of timing information in a performance is much larger than the score’s rhythmic content. A four-part chord contains three inter-onset intervals the notation says nothing about, and there is one of those for every chord.
Why milliseconds and not fractions
The obvious question is why performance timing should be organised around absolute durations at all, when everything else about a piece of music scales with the tempo.
Two answers are available and both are partial.
The first is that the body does not scale. A bow change, a stick rebound, a breath and a finger lift all take about as long as they take, and none of them shortens proportionally when the tempo doubles. Any deviation that originates in the mechanics of playing will be measured in milliseconds by construction.
The second is that the auditory system has its own fixed timescales. Two onsets closer than about 20 to 30 milliseconds fuse into one event; the threshold for detecting that two sounds are not simultaneous is around 20 milliseconds for most timbres; and the region in which events are heard as rhythmically separate rather than as one articulation has a floor somewhere near 100 milliseconds — the same floor that makes a swing ratio fall with tempo. A deviation of thirty milliseconds is large enough to be perceived as an offset and small enough not to be perceived as a different rhythm, and that window is a fixed number of milliseconds wide.
Between them, those two accounts predict that expressive timing should sit in the tens of milliseconds and should not scale with tempo, which is what is observed.
Two evenly spaced pulses over one span coincide only where their grids agree, and an ensemble with a laid-back soloist is that arrangement with a very small offset: two layers sharing a tempo and not a phase. What makes it different from a genuine cross-rhythm is that the offset is a fixed number of milliseconds rather than a ratio — so as the tempo rises the two layers converge, and at some speed the lag is smaller than the fusion threshold and the ensemble simply plays together. A ratio has no such tempo.
Where the model stops
These are means over small corpora. Every number in the figure is an average, over a handful of performances, of a quantity that varies from bar to bar within a single performance. Drawing them as fixed offsets is a simplification of what was measured, and the variance around each mean is real and is not shown.
Systematic is not the same as intentional. A deviation that recurs every bar may be a stylistic choice, a consequence of how an instrument is played, or an artefact of the conductor’s beat. The measurements cannot separate those, and the literature is careful not to.
Not all deviation is groove. Performances also contain phrase-level rubato, ritardandi, and ordinary imprecision, and those are far larger than the effects here. Isolating the systematic part requires assuming a model of the non-systematic part, and the results depend on that assumption.
Removing them does not always hurt. The claim that quantisation destroys something is well supported for repertoire whose interest is rhythmic. For a great deal of other music the deviations carry very little, and a quantised performance is merely a tidy one — much as a rhythm drawn as a cycle loses nothing when its onsets are on the grid to begin with.
Two profiles is not a survey. The Viennese waltz and the jazz phrase are the two best-documented cases and they are two cases. Whether comparable systematic patterns exist in most repertoires is an open question, not an established fact, and the honest position is that they have been looked for in relatively few places.
Whose music, and when
Systematic performance timing has been measurable since the 1930s and studied seriously since the 1970s, and the studies cluster around the repertoires whose practitioners insisted there was something there.
Seashore’s laboratory at Iowa published measurements of vocal and violin performance from the 1930s, using photographic recording of pitch and time. The finding that struck contemporaries was the sheer size of the deviations — performers were nowhere near the notated values and audiences did not notice, which suggested that the notated values were not what was being heard.
The Swedish group around Sundberg, Bengtsson and Gabrielsson did the systematic work from the 1960s onward, including the waltz measurements, and framed the results as a set of performance rules that could be applied to a score to make a synthesis sound played rather than printed.
Jazz and popular-music timing research is mostly from the 1990s onward, and includes both the swing-ratio work and the studies of ensemble timing relationships that produced the “laid back” measurement. It arrived late partly because the repertoire was studied late and partly because the necessary onset-detection tools arrived with cheap computing.
Everything in a metre diagram is a proportion, at every level, including the unequal beats of an additive metre, whose long beat is three subdivisions rather than any number of milliseconds. That is what makes notation transposable across tempi and what makes it silent about the quantity this essay measures. The practical consequence runs the other way from what might be expected: notation’s inability to record absolute timing is not a failing to be fixed but the reason the notation works. A score whose rhythms were durations would be playable at one tempo; a score of proportions is playable at any.
The practical consequence runs the other way from what might be expected. Notation’s inability to record this is not a failing to be fixed but the reason the notation works. A score whose rhythms were absolute durations would be playable at one tempo; a score of proportions is playable at any. The timing information has to live somewhere else, and where it lives is in the aural transmission of a style — which is why it can be lost, and why it is repertoire-specific, and why an orchestra that has never played a Viennese waltz cannot produce one from the parts.
What is actually being transmitted
If the material cannot be written, the question of how it gets from one player to another becomes a real one, and the answer has consequences.
It is transmitted by listening and copying, which is the same channel by which a raga’s characteristic movement or an aksak dance’s step pattern travels. That has three observable consequences.
It is local. The Viennese second beat is Viennese. Other orchestras playing the same waltzes do not produce it, and the ones that try tend to overshoot, because the target is a specific number of milliseconds and the only available instruction is a description.
It can be lost. A performing tradition that stops being handed down stops existing, and no archive of scores recovers it. What can be recovered is what was recorded, which is why the twentieth century is the first period whose performance practice is knowable at this resolution and every earlier one is not.
It is what recordings are for. A recording preserves the layer notation cannot. That is a mundane observation with an immodest consequence: for any repertoire older than about 1900, the timing dialect is simply gone, and every historically informed performance of anything earlier is a reconstruction of pitch and rhythm with a modern timing practice underneath it.
The last of those is worth sitting with. A great deal of effort goes into recovering historical tuning, historical instruments and historical ornamentation, all of which leave documentary traces. Timing at this scale leaves none, and it is at least as large a component of what a performance sounds like.
The ladder from here
Later rungs on this anchor: the variance around these means, and what a distribution of deviations looks like within one performance. Ensemble timing relationships — who follows whom, and with what lag, in a string quartet and in a rhythm section. Beat-level asynchrony between instruments, which is systematic and larger than most people expect. The performance-rule systems that try to generate expressive timing from a score, and how far they get. And the question of transmission: what is actually being copied when a player learns a feel, given that the thing being copied cannot be written down.
Fifty-five milliseconds, every bar, for two hundred years, and no symbol for it anywhere.
Part 2 of 9
One essay in the series on microtiming. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 30.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Expressive timingGrooveMicrotimingNotationQuantisation
- How much of this is new notation, quantisation
- The shape a wall leaves behind groove, microtiming