The beat is inferred, and sometimes wrongly
Metre is the most confidently held opinion in music and it has the least evidence behind it. A listener knows where the downbeat is, will tap it without hesitation, and will be certain — and there is frequently nothing in the sound that distinguishes their answer from someone else’s.
The reason is that metre is not a property of the signal. It is an inference the listener performs on the signal, and it can be modelled.
That the model ties is not a failure of the model. A pattern of evenly spaced onsets genuinely supports several readings, and which one arrives depends on what came before it.
The rules, and where they come from
The scoring above is a simplified version of the metrical preference rules in Fred Lerdahl and Ray Jackendoff’s A Generative Theory of Tonal Music (1983), which is the most-cited attempt to write down what a listener is doing when they find a beat.
Their proposal is a list of preferences, weighted against each other, of which the ones this figure implements are the strongest:
- Prefer a metre whose strong positions coincide with event onsets.
- Prefer a metre whose strong positions carry longer notes.
- Prefer a metre whose strong positions carry louder events, or events that begin a group.
- Prefer to keep the metre already established rather than switch.
The last is doing an enormous amount of work and is the reason the model’s ties are less common in practice than in isolation: a listener arrives at any given bar with an expectation, and the expectation is worth several points. Both halves of that sentence can be counted.
Running the same four candidates over every twelve-step pattern that begins on an onset and has three or more of them — 2,036 patterns — the top two candidates are exactly tied in 34.2 per cent of them, and 65.4 per cent are within two points, which is the next gap the rule set’s integers allow. So the tie in the figure above is not a curiosity chosen to make a point; it is what a third of the space does.
And the persistence term’s size can be read off the same census, which is which of the two is the beat asked over every pattern at once. Give an already-established metre a bonus and ask how often some pattern still overturns it:
| the bonus is worth | an established metre is overturned by some pattern in |
|---|---|
| 0 or 1 point | 100% of patterns |
| 2 or 3 | 68.8% |
| 4 to 6 | 45.5% |
| 8 | 21.9% |
The persistence preference has to be worth about eight points before it holds against three quarters of the patterns a listener could meet, and eight is nearly three times the strongest positive term in the whole rule set — the three points a strong position earns for carrying an onset.
That changes what kind of rule it is. Written as the last item in a list of four preferences it reads as a tie-breaker, something to fall back on when the surface is ambiguous. On these numbers it is the dominant term, larger than any evidence a single bar can supply, and the other three preferences are what decides the metre only in the first bar or two before it starts to bind.
Which is the right shape for what the rest of this essay is about. A frame that costs eight points to leave is a frame a listener commits to, and a metre has to be able to change its mind before that commitment is worth anything; it does not casually revise, and the figure-ground character of correcting a wrong downbeat is what a large persistence term feels like from the inside. A model whose persistence was worth one point would flicker between readings bar by bar, and nobody hears music that way.
The scores are the argument: 9 for the three-three-two reading against 1 for each of the two even ones, which is not a close call. The three-three-two reading catches all three onsets on strong positions and leaves none empty; the four-four reading catches two of four strong positions and strands an onset off the grid.
The tresillo is the interesting contrast with the opening figure. It is a Euclidean rhythm — three onsets spread as evenly as possible over eight steps — and its unevenness is what makes it unambiguous. A pattern that is perfectly even supports many metres; a pattern that is nearly even supports one.
What a wrong inference feels like
The strongest evidence that metre is inferred rather than perceived is that the inference can be wrong, and that being wrong has a distinctive character.
The everyday case is coming in late to a piece of music. A listener who joins a recording partway through commits to a downbeat, and if the commitment is off by one beat the music does not sound wrong — it sounds like a different piece, with different syncopations and a different character. Correcting it, when a cue eventually arrives, has the quality of a figure-ground reversal rather than of noticing a mistake.
The laboratory version has been run many times: the same recorded pattern, preceded by different priming contexts, produces different reported metres and different judgements about which events are accented. The physical accents in the pattern are unchanged. What moved was the frame.
The clearest musical exploitation is the hemiola, where a passage in three is written so that its accents suggest two, and the listener is caught between the metre they have established and the one the surface is proposing. It is a device that exists only because metre is committed to in advance and is expensive to revise.
The strong-position hierarchy
Metre is not just a period; it is a hierarchy, and the hierarchy is what the scoring above is really about.
In four-four the first beat is the strongest, the third is next, and the second and fourth are weak. Subdivide and the pattern continues downward: the first half of each beat is stronger than the second. That nested structure is what allows a listener to feel a syncopation, which is an event on a weak position where a strong one was expected.
Syncopation can then be defined precisely, and counted: an onset on a position weaker than a position it displaces. That definition supports a syncopation index, several of which have been published, and they broadly agree with musicians’ judgements of which rhythms are more syncopated.
They also produce the result that matters here. A rhythm’s syncopation is not a property of the rhythm — it is a property of the rhythm paired with a metre, and the same onsets under a different metre have a different count. There is no such thing as a syncopated pattern in isolation.
Entrainment, which is the other half
The preference-rule account describes which metre is chosen. It says almost nothing about the more striking fact, which is that having chosen one, a listener keeps time with it.
A person tapping along with music does not respond to each beat as it arrives — the taps precede the beats, typically by twenty to fifty milliseconds. That is not possible for a reactive system, and it is the standard evidence that what is happening is entrainment: an internal periodicity synchronised to the input and running slightly ahead of it, in the same family of behaviours as fireflies flashing together or a pendulum clock pulling its neighbour into step.
Entrainment explains several things preference rules do not.
The beat survives silence. Remove events for a bar and the metre continues; a listener taps through a gap without hesitation. A rule set that scores onsets against strong positions has nothing to say about a passage with no onsets in it.
Tempo changes are tracked smoothly. A gradual accelerando is followed rather than repeatedly re-inferred, which is what an oscillator locking to a drifting signal does and not what a scoring model does.
Only some periods are available. Entrainment is easiest near two events a second and gets progressively harder away from it, which is why the same rhythmic pattern played very fast or very slowly stops supporting a beat at all — and why a listener presented with an ambiguous pattern will usually settle on whichever candidate period is nearest that comfortable rate. That is a preference rule nobody wrote down, and it comes out of the mechanism for free.
That a listener entrained to one layer hears the other as syncopation, and can switch at will, is the everyday version of the tie above: the fit rule genuinely does not choose, and something outside it does.
The two accounts are complementary rather than rival. Preference rules describe the initial choice from the surface; entrainment describes the maintenance, the anticipation and the cost of switching. Contemporary models combine them, and the combination is what makes the fourth preference rule — keep the metre already established — mechanical rather than stipulated.
Where the model fails
Lerdahl and Jackendoff’s rules were written for common-practice European tonal music and they fail in ways that are worth naming, because the failures are informative rather than embarrassing.
Metres that are not built from twos and threes. The rules assume a hierarchy of equal divisions, and an aksak bar is unequal at the beat level — a bar of 2+2+3 has beats of two different lengths, which no tree of equal subdivisions produces. The model can be extended and the extension is not a small one.
Traditions with a non-metrical timeline. West African drumming is organised around a repeating bell pattern that functions as a reference without being a metre in this sense: it has no downbeat that all parts agree on, and different parts of the ensemble may relate to it with different orientations. A model that outputs one winning metre per passage is answering the wrong question about this music, and drawing the pattern as a cycle rather than a line is the first step towards asking a better one.
Expressive timing is treated as noise. The rules take a list of onsets and assume they are on the grid. Real performance is not on the grid: the deviations are systematic, they are large enough to be the whole of what a groove is, and a model that quantises them away has thrown out information a listener demonstrably uses. Worse, the deviations themselves carry metrical information — a performer lengthens the note before a strong beat — so the model is discarding evidence for the very thing it is trying to infer.
The rules are static and listening is not. The preference for continuing an established metre means the answer depends on history, and the published rule sets handle this by fiat rather than by dynamics. Models that treat metre finding as an entrainment problem — a bank of oscillators that lock onto periodicities in the input — capture that better, and they are the direction the field went after the mid-1990s.
The pieces that exploit it
Composers have known this without a model for a very long time, and the devices that depend on metre being an inference are among the oldest in the repertoire.
Anacrusis. A phrase beginning before the downbeat depends entirely on the listener having a downbeat to be before. Heard cold, an upbeat is indistinguishable from a downbeat, and the whole meaning of the gesture is supplied by the frame it arrives in.
The suspended barline. The opening of a piece is the one place where the listener has no established metre and the preference rules run unaided, which is why openings tend to be metrically explicit — a clear downbeat, a strong bass note, an obvious grouping — and why a composer who wants ambiguity puts it there. The first bars of Beethoven’s Ninth are the standing example: an open fifth with no clear downbeat, which resolves into an unambiguous metre only when the theme arrives.
Metric modulation. Reinterpreting a subdivision of the old metre as the beat of a new one — three quavers of the old bar becoming a beat of the new — works because the listener’s entrainment can be captured by a periodicity that was already present in the signal. It is a device that exists only because there is something to capture.
The rewritten bar line. A great deal of twentieth-century notation practice is about writing down a metre that a listener would not infer, and the fact that this requires special notation is the point. Stravinsky’s constantly-changing time signatures in The Rite of Spring are instructions to the players; what the listener does with the result is a separate question, and often a different answer.
Each of these is a way of putting the listener’s inference under strain, and each of them stops working if the inference is removed. That is as close to a controlled experiment as the repertoire offers.
What the model gets right
Set against the failures, three things the preference-rule account gets right and that no simpler account does.
It predicts which patterns are ambiguous, and the predictions match. Evenly spaced onsets are ambiguous; unevenly spaced ones with a clear grouping are not. That is testable and it holds — and it is the reason the Euclidean rhythms are as widespread as they are, since a pattern spread as evenly as possible without being perfectly even is exactly one that defines its own metre while remaining interesting.
It explains why a metre survives contradiction. Once established, a metre is worth points, so a passage can contradict it for some time before the reading flips — which is exactly what syncopation depends on, and what would be impossible if metre were read off the signal afresh each bar.
And it explains why notation works at all. A time signature is a communication of the intended reading, and it is necessary precisely because the reading is underdetermined. If metre were in the sound, a score would not need to specify it — and the things notation records and does not record are a good index of what is genuinely ambiguous.
The weights, and why they are a choice
The three numbers in the scoring — three points for a hit, minus two for an empty strong position, minus one for an offbeat onset — are the part of this figure with the least justification behind it, and saying so is more useful than defending them.
Lerdahl and Jackendoff give their rules as an ordered list of preferences without numerical weights, on the grounds that the weights are not known and that ordering is the most that the evidence supports. Turning an ordered list into a score requires inventing numbers, and any set of numbers that preserves the ordering will produce the same winner in clear cases and may differ in close ones.
What the figure’s numbers are chosen to do is preserve two properties. A strong position carrying an onset must be worth more than an offbeat onset costs, or a metre could win by having no strong positions at all. And an empty strong position must cost something, or a metre with very few strong positions would always win by never being contradicted. Those two constraints are what pin the signs and the rough magnitudes; the exact values are not pinned by anything.
The consequence is worth taking seriously. A tie in this model — like the three-way tie in the opening figure — is a genuine statement that the pattern does not distinguish the readings under any reasonable weighting. A narrow win is not a statement about anything, and should be read as a tie that the arithmetic happened to break.
What the picture cannot show
A score of candidate metres is a snapshot, and metre finding is a process with a time course: a listener takes a few seconds to settle, revises, and becomes progressively harder to shift. None of that dynamics is visible in a table of totals.
The figure also cannot show confidence. The model outputs a ranking, and what a listener has is a ranking plus a strength of commitment, which is the quantity that decides whether a syncopation is thrilling or merely confusing. Two patterns can have the same winner and feel completely different because one won narrowly and the other by a mile.
The ladder from here runs into what happens when the inferred beat and the played events deliberately disagree by small amounts — which is the twenty to fifty milliseconds that make a groove, and which are meaningless without a beat to be measured against.
Part 1 of 9
One essay in the series on metre induction. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 38.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
DownbeatEntrainmentMetreOnset patternSyncopation
- No term for an unequal beat metre, onset pattern, syncopation
- Syncopation is a number about the metre metre, onset pattern, syncopation
- A dancer who comes in late needs the downbeat marked downbeat, metre
- Asked for the rate, it answers a multiple downbeat, metre
- Expectation is a curve, not a list metre, syncopation
- The change reading follows the chords, not the bar downbeat, metre