Swing is a ratio, and it is not two to one
Jazz notation has a standard instruction at the top of the page: a pair of quavers, an equals sign, and a triplet crotchet followed by a triplet quaver. It means that written pairs of equal notes are to be played long-short, in the ratio two to one — a proportion, of the same kind every other rhythmic symbol carries.
It is a useful instruction and it is wrong in a specific and measurable way. Players do not produce a two-to-one ratio, except at one tempo, and the ratio they do produce changes continuously as the tempo changes.
What was measured
Anders Friberg and Andreas Sundström published the key measurement in 2002. They analysed ride-cymbal patterns from recordings by four jazz drummers across a wide range of tempi, extracted the onset times, and computed the ratio of the long note to the short note in each swung pair.
Two results came out of it.
The first is that the swing ratio falls steadily as the tempo rises — from about 3.5:1 at slow tempi down to about 1:1 at fast ones. Above roughly three hundred beats a minute there is no swing left to measure; the quavers are even, and the pattern is the one the grid would have produced anyway.
The second result is the one that explains the first. The short note of each pair holds a roughly constant absolute duration of about a tenth of a second, across a wide range of tempi. The long note absorbs everything else.
Which computation produced the number
The curve in that figure is not a fit to the ratio data. It is what the constant short note implies, worked out directly.
If the beat lasts milliseconds and the short note takes a fixed of them, the long note takes the rest, and the ratio is
With milliseconds this gives 5:1 at 100 beats a minute, 3:1 at 150, 2:1 at 200, 1.4:1 at 250 and exactly 1:1 at 300. The published measurements saturate around 3.5:1 at the slow end rather than continuing to rise, so the model carries a ceiling at that value — two parameters in total, and they reproduce both published endpoints.
The tempo at which the notated triplet is correct is then the solution of , which is
Two hundred beats a minute is a brisk medium-up tempo. Below it the notation understates the swing; above it, it overstates.
Three things about that model are worth stating that the two parameters do not announce.
The ceiling divides it into two regimes, and the headline mechanism holds in only one. The ratio reaches 3.5 at 133 beats a minute, so below that the ceiling binds and the short note is no longer constant — it is 111 milliseconds at 120 beats, 121 at 110, 133 at 100 and 167 at 80. So “the short note holds a constant tenth of a second” is a description of the range above 133 beats a minute, and everything slower is the other parameter doing the work. The 121-millisecond short note in the figure at 110 says so and does not say that it says so.
The crossing moves with the short note, in proportion. The tempo at which the triplet is exactly right runs from 250 beats a minute at a short note of 80 milliseconds to 167 at 120 — so a ten per cent uncertainty in a quantity measured across four drummers is a ten per cent uncertainty in the answer, and 200 is a round number standing for something between about 180 and 220.
And the notation is right over a very narrow band. Taking “right” to mean within ten per cent of the model’s ratio, the triplet reading holds from 188 to 214 beats a minute — twenty-six beats a minute out of the two hundred and forty a jazz group might use. Loosening it to twenty-five per cent widens the band only to 172 to 240.
So the instruction is not approximately right with exceptions at the edges. It is exactly right at one tempo, tolerably right across about a tenth of the usable range, and wrong by more than a quarter everywhere else — which is a stronger statement than the section above makes and is the one the arithmetic supports.
The same instruction at three tempi
The clearest way to see what the notation is failing to carry is to draw the same passage at different speeds.
The three pictures carry the same instruction and produce three different rhythms. That is not a defect in the instruction; it is a category error in what the instruction is being asked to do. Swing is a rule for placing an event in time, and the notation is a system for specifying proportions. A rule that fixes an absolute duration cannot be written as a proportion without also fixing the tempo — which is the same difficulty an unequal beat presents from the other direction.
Why a constant short note
The obvious question is why the short note should be the one that stays put, and the answer is not settled, but there are plausible accounts and they are testable.
One is perceptual: below about 100 milliseconds two events start to fuse into one, and a note shorter than that stops reading as a separate articulation — a threshold an instrument’s attack has to respect too. On that account the constant is a floor rather than a preference — the short note cannot get much shorter and stay audible as a note.
Another is motor: the physical action of a stick rebound or a tongued articulation has its own timescale, which does not scale with tempo. On that account the constant is a limit of the body rather than of the ear.
A third is that the constant is not really constant, and the flat region in the data is an average over players who differ. Friberg and Sundström’s four drummers do differ, and the 100-millisecond figure is a central tendency.
All three predict roughly the same curve, which is why the curve is not evidence for any of them in particular.
A prediction that can be checked
The model earns its keep by saying something falsifiable, and it says one thing sharply.
If the short note really is constant, then its duration in milliseconds should be the same at every tempo above the point where the ceiling stops binding — and, equivalently, the offbeat should fall a constant number of milliseconds before the next beat, regardless of how fast the music goes.
That is a stronger claim than “the ratio falls with tempo”, and it is easy to test on any recording with a steady tempo. Extract the onsets, find the beats, and measure the gap between each offbeat and the beat that follows it. If the gap is constant, the model holds; if the gap scales with the beat, the ratio is constant and the model is wrong; if the gap does something else, both are wrong.
Friberg and Sundström’s data behaves the first way above about 150 beats a minute and the second way below it, which is exactly why the model needs its ceiling. The ceiling is therefore not a free parameter dressed up as a finding — it marks the tempo below which the constant-short-note account stops applying, and that boundary is itself a measurement.
Two grids at once
There is a further consequence, and it is the one that makes swing structurally interesting rather than merely oddly notated.
A jazz rhythm section is not playing one rhythm. The ride cymbal is placing offbeats by the rule described above; the bass is playing quarter notes on the beat; the drummer’s hi-hat is on beats two and four. Those layers share their beats and nothing else, which is what two clocks running at once amounts to, and the offbeat is the only place where a timing decision has to be made at all.
That is why swing survives being played by people who have never discussed it. There is one degree of freedom, it is shared by everyone playing offbeats, and the rest of the ensemble is unaffected. It is also why a swung passage played by a machine at a fixed ratio sounds wrong at some tempi and not others: the machine is applying a proportion where a player applies a duration.
What a listener hears instead
None of this describes what the ratio sounds like, and there the picture is different again.
Listeners are markedly poor at reporting rhythmic ratios accurately and markedly good at sorting them into categories. Presented with a continuum of long-short pairs, listeners tend to hear them as one of a small number of types — even, two-to-one, three-to-one — rather than as a continuous quantity, and the boundaries between the categories are sharper than the underlying acoustic differences.
The consequence for the argument above is a pleasant one. A player producing 2.4:1 and a player producing 2.8:1 are doing measurably different things, and a listener will describe both as swung, because both fall inside the same perceptual category. The notation’s two-to-one is therefore not a description of what is played but a name for the category — and as a name it works fine.
Where it stops working is when the instruction is handed to something that takes it literally. A sequencer given a triplet grid produces a rhythm that is correct at 200 beats a minute and increasingly wrong away from it, and the characteristic stiffness of programmed swing is that error rather than a lack of some ineffable quality. It is the same failure as quantising a groove, arriving one level further down. Most sequencers now offer a swing amount as a percentage, which is a proportion, and which therefore has exactly the same problem in a more adjustable form.
Where the model stops
This is a ride cymbal, not a band. The measurement is of one instrument playing one pattern. A bass player walking quarter notes, a pianist comping and a soloist all have their own timing relationships to that cymbal, and they are not the same. Treating “the swing ratio” as a property of an ensemble goes well beyond what was measured.
Four drummers is four drummers. The corpus is small, it is specific to a style and a period, and individual variation within it is real. The curve is a summary of a handful of recordings, not a law.
The ceiling is a fudge. The constant-short-note model on its own predicts a ratio of 5:1 at 100 beats a minute, and the data does not go there. The ceiling at 3.5 is imposed to match the observation rather than derived from anything, and it is the weakest part of the model.
Ratio is not the only variable. A swung pair can have the same ratio and a different articulation, a different accent pattern and a different amount of overlap between the notes. Everything that makes one drummer’s swing recognisable and another’s not is outside this measurement.
Nothing here is about the beat itself. The ratio describes where the offbeat falls within a beat whose boundaries are assumed fixed. Whether the beats themselves are evenly spaced is a separate question with its own answer, and the answer is often no.
The 100-millisecond figure is not universal across styles. Shuffle patterns in blues and rock, and swung sixteenths in funk, have their own characteristic ratios and their own tempo behaviour, and the jazz ride figure should not be assumed to describe them.
Whose music, and when
The notation problem is older than jazz, and one earlier tradition handled it more honestly.
Notes inégales, in French Baroque practice, is the convention of playing written-even pairs unequally. The seventeenth- and eighteenth-century sources — Loulié, Hotteterre, Couperin and others — describe it, disagree with each other about how much inequality is appropriate, and are explicit that the amount varies with the tempo, the metre, the character of the piece and the taste of the performer. Some sources give ratios; the ratios given range from roughly 3:2 to 3:1.
That is a considerably better description of the phenomenon than the modern jazz instruction, because it says the ratio is variable and refuses to fix it. It is also much less convenient, which is presumably why the modern convention prefers a symbol.
The jazz notation itself dates from the swing era, when printed arrangements needed something to put on the page for players who might not have heard the style. The triplet symbol is a piece of engraving pragmatism: it is writable, it is roughly right in the middle of the tempo range, and it is unambiguous. A more accurate instruction would have to be a sentence, and sentences do not fit above a stave.
The shuffle in blues and rock is the same device at a much narrower range of tempi, which is why the triplet reading works better there. A shuffle is typically played between about 80 and 140 beats a minute, and within that band the ratio a player produces stays closer to a constant than the jazz figure does — partly because the range is narrow and partly because the shuffle sits on a subdivision that is genuinely triplet-based in the surrounding parts.
Swung sixteenths in funk and in later jazz introduce a further complication: the ratio applies at half the note value, and the same absolute short-note constraint therefore bites at half the tempo. A funk groove at 100 beats a minute has sixteenth-note pairs whose beat is 300 milliseconds, so a 100-millisecond short note gives a ratio of 2:1 — which is why swung sixteenths at moderate tempi sound much more strongly swung than swung quavers at the same tempo.
What a notation would need
It is worth asking what would actually fix this, because the answer is instructive about notation in general.
A symbol that carried swing correctly would have to specify a duration in milliseconds, or equivalently a duration plus a tempo. Notation has no way of writing that: every rhythmic symbol on a stave is a proportion of a beat, and the beat’s absolute length is given separately at the top of the page and is expected to vary.
The separation is deliberate and it is what makes notation useful. A piece can be played faster or slower and remain the same piece precisely because its rhythms are proportions. Fixing an absolute duration anywhere in the system breaks that, and would mean a passage stopped being playable at another tempo.
So the omission is not an oversight to be repaired by a better symbol. It is the price of the abstraction, and the same price is paid everywhere else the same abstraction is used — which is why rubato, groove and articulation are all handled by words rather than by symbols, and why every one of them is a source of the same complaint.
The historical evidence supports reading it that way. Every tradition that has tried to notate timing precisely has ended up with either a proportional system plus a verbal instruction, or a table of durations that only works at one speed. Nobody has produced a third option that musicians adopted.
The ladder from here
Later rungs on this anchor: what the rest of the ensemble does with respect to that cymbal, which is where “laid back” and “on top” become measurable. The relationship between swing ratio and dynamics, which is not independent — the long note is usually louder as well. Notes inégales in detail, and what can be recovered about a practice from sources that disagree. Swung sixteenths and the halving of the timescale. And the general problem this essay is a case of: what a notation for timing would have to look like, and why every attempt so far has been a proportional system with a rule bolted on the side.
Two to one is right at two hundred beats a minute. The instruction has been on the page for ninety years and is right about once a chorus.
Part 1 of 9
One essay in the series on microtiming. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 15.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
MicrotimingNotationSubdivisionSwingTempo
- A phrase is a number of seconds notation, tempo
- A silence long enough to be an ending notation, tempo
- A staccato is a dynamic mark notation, tempo
- How much music a page holds notation, tempo
- The axis that is not a time axis notation, tempo
- The time signature is a claim notation, subdivision