Form and structure

A metre has to be able to change its mind

Replace the scoring function with a phase-corrected oscillator and the two failures reported earlier separate. The beat now survives two silent bars, drifting 23 milliseconds of a 560-millisecond beat. But it does not lose a moving tempo so much as lag behind it, by the rate over the correction gain — and a cadential ritardando that halves the tempo in eight beats is faster than the model can follow.

Assumes: The beat is inferred, and sometimes wrongly

The last two rungs found the same missing thing twice. A scoring function has nowhere to keep a commitment, so it cannot hold a beat through a bar that does not mark it; and its candidates are single numbers, so it cannot propose a beat that is not the same length every time.

The first of those has a standard repair and it is old: make the metre a state rather than a reading. A period and a next beat time, carried forward, nudged by whatever arrives near the beat. That is a coupled oscillator, it is what the field moved to after the mid-1990s, and it is what the rung that introduced the rules named as the alternative without running one.

A beat kept through the bars that do not mark it. A 16-step pattern played for 6 bars, with every onset removed from the beat itself in the middle two, and seeded timing noise on every onset. The line is how far the tracker's beat sits from where it started, in milliseconds of a 560 ms beat. It coasts through beats 9 to 16, keeps its period, and comes out at most 23 ms from true — 4 per cent of a beat. Handed the silenced bars on their own, the preference rules move the downbeat from phase 2 to phase 0.
Fig. 1 Six bars — two intact, two with every onset removed from the beat, two intact again, all with seeded timing noise. The tracker coasts through the eight silenced beats, keeps its period, and comes out at most 23 milliseconds from where it went in, on a beat of 560. Handed the silenced bars alone, the preference rules move the downbeat.

The whole mechanism, in three lines

The tracker holds two numbers: a period, and the time of the next beat.

At each predicted beat it looks for an onset nearby. If it finds one, it corrects: the beat time moves toward the onset by a fraction α, and the period changes by the error times a fraction β. If it finds nothing, it does the only other thing available — adds the period to the beat time and carries on.

That is the entire model. There is no scoring, no candidate list, and no representation of a bar. What makes it able to do the thing the rules cannot is not sophistication; it is that the beat time is a variable that persists between beats, so there is somewhere for a commitment to live.

The two gains have published ranges from tapping experiments: phase correction is fitted at roughly 0.5 to 0.9 and period correction at roughly 0.05 to 0.3. Every figure here uses 0.6 and 0.2 unless it says otherwise, and where the answer depends on the choice, that is stated.

Coasting is not free

The tracker’s silence behaviour is entirely determined by the period it was holding when the silence began, and a period estimate carries whatever error the preceding bars left in it.

A beat kept through the bars that do not mark it. A 16-step pattern played for 8 bars, with every onset removed from the beat itself in the middle two, and seeded timing noise on every onset. The line is how far the tracker's beat sits from where it started, in milliseconds of a 560 ms beat. It coasts through beats 9 to 24, keeps its period, and comes out at most 29 ms from true — 5 per cent of a beat. Handed the silenced bars on their own, the preference rules move the downbeat from phase 2 to phase 0.
Fig. 2 The same experiment with four silent bars rather than two. The tracker coasts through sixteen beats and its worst excursion grows from 23 milliseconds to 29 — it degrades, and it degrades slowly, because the error is a fixed period bias accumulating rather than a random walk.

Two silent bars cost 23 milliseconds; four cost 29; eight cost 57. The growth is roughly linear in the number of silent beats, which is what a constant period error looks like, and it is very much better than the random walk two independent clocks produce.

It is also entirely dependent on the quality of the estimate going in.

A beat kept through the bars that do not mark it. A 16-step pattern played for 6 bars, with every onset removed from the beat itself in the middle two, and seeded timing noise on every onset. The line is how far the tracker's beat sits from where it started, in milliseconds of a 560 ms beat. It coasts through beats 6 to 24, keeps its period, and comes out at most 185 ms from true — 33 per cent of a beat. Handed the silenced bars on their own, the preference rules move the downbeat from phase 2 to phase 0.
Fig. 3 The same two silent bars with the timing noise on the onsets raised from 12 milliseconds to 40. The tracker enters the silence with a much worse period and comes out 185 milliseconds off — a third of a beat, which is a lost beat rather than a wobble.

So the model predicts something specific and checkable: a beat is easier to keep through a break after a tightly played passage than after a loose one, and the difference is not small. Whether that is true of listeners is not established, and the experiment is obvious.

Two values are not a curve, though, and sweeping the jitter shows the degradation is not a gradient at all. From 4 milliseconds up it is orderly — 8, 15, 23, 38, 42 at 30 — and then between 34 and 36 milliseconds it jumps from 48 to 191 and stays there. There is a cliff, not a slope, and past it the tracker has latched onto the wrong onset rather than kept the right one imprecisely.

Past the cliff the outcome is also a lottery. Re-running at 40 milliseconds with eight different noise seeds gives worst drifts of 89, 90, 147, 154, 185, 187, 334 and 487 milliseconds — a factor of five, with the largest most of a beat. At 12 milliseconds the same eight seeds give 14 to 38, a factor of under three around a small number. So the tight case is reliable and the loose case is not merely worse but unpredictable, which is a different claim and a better one: what a loose passage costs is not a known amount of drift but the possibility of a lost beat.

A moving tempo is not lost, it is lagged

The other thing the earlier rung listed as a scoring model’s failure was tempo change: “a gradual accelerando is followed rather than repeatedly re-inferred, which is what an oscillator locking to a drifting signal does”.

It does, and the way it does is worth drawing, because it is not the way the sentence suggests.

A tracker following a tempo that will not stay still. The tracker's error as a fraction of a beat, against beat number, for 3 rates of tempo change with a period-correction gain of 0.2. It settles at 0.025 of a beat behind at 0.5 per cent a beat, 0.092 of a beat behind at 2.0 per cent a beat, 0.206 of a beat behind at 5.0 per cent a beat. It does not lose the beat; it lags, by very nearly the rate divided by the period-correction gain, and the lag reaches a quarter of a beat at 6.3 per cent a beat — at which point the tracker is nearer the wrong onset than the right one.
Fig. 4 The tracker’s error, as a fraction of a beat, following three rates of tempo change. It does not lose the beat and it does not catch up either: at each rate it settles at a fixed distance behind and stays there, because a first-order corrector following a ramp has a constant steady-state error.

A tracker that corrects its period in proportion to its error can only reduce the error to the point where the correction it makes equals the change it has to keep up with. That leaves a constant lag, and the lag is very nearly the ramp rate divided by the period-correction gain.

At 0.5 per cent a beat with a gain of 0.2, the lag settles at 0.025 of a beat. At 2 per cent, 0.092. At 5 per cent, 0.206. Divide each by 0.2 and the arithmetic is visible.

“Very nearly” is doing measurable work in that sentence, and the measurement is worth having, because the law is an upper bound that loosens as the rate rises. Against rate over gain the measured lag is 98 per cent of it at 0.5 per cent a beat, 92 at 2 per cent, 83 at 5 and 74 at 8 — and at the bottom of the gain range it is worse still, 46 per cent at 8 per cent a beat with a gain of 0.05. The tracker always follows better than the law says, by more the harder it is being pushed, because the period is compounding while the correction is proportional to the error it has already reduced.

That matters for the limit rather than for the lag. Setting rate over gain to a quarter would put the quarter-beat boundary at 25 times the gain; the true boundary is about 31 times it — 1.50 per cent at a gain of 0.05, 6.32 at 0.2, 9.53 at 0.3 — so the model tolerates about a quarter more tempo change than the closed form predicts. The scaling with the gain is nearly exact; the constant is not the obvious one.

And there is a rate at which lagging becomes losing

A lag of a tenth of a beat is a following listener slightly behind the music. A lag of half a beat is a listener on the offbeat, and somewhere between those the tracker stops being behind the right onset and starts being ahead of the wrong one.

Taking a quarter of a beat as the boundary — the point at which the nearest onset is as likely to be the next one as the current one — the limit at a gain of 0.2 is 6.3 per cent a beat.

A tracker following a tempo that will not stay still. The tracker's error as a fraction of a beat, against beat number, for 3 rates of tempo change with a period-correction gain of 0.2. It settles at 0.048 of a beat behind at 1.0 per cent a beat, 0.133 of a beat behind at 3.0 per cent a beat, 0.248 of a beat behind at 6.3 per cent a beat. It does not lose the beat; it lags, by very nearly the rate divided by the period-correction gain, and the lag reaches a quarter of a beat at 6.3 per cent a beat — at which point the tracker is nearer the wrong onset than the right one.
Fig. 5 Three rates approaching the limit, the last of them at it. At 6.3 per cent a beat the settled lag is a quarter of a beat, which for a beat of 560 milliseconds is 140 — a whole subdivision, and the point at which the tracker’s next correction may pull it onto the wrong onset entirely.

The limit scales with the gain, because the lag does. At the bottom of the published range for period correction it is 1.5 per cent a beat; at the top it is about 9.5.

A tracker following a tempo that will not stay still. The tracker's error as a fraction of a beat, against beat number, for 3 rates of tempo change with a period-correction gain of 0.05. It settles at 0.093 of a beat behind at 0.5 per cent a beat, 0.315 of a beat behind at 2.0 per cent a beat, 0.588 of a beat behind at 5.0 per cent a beat. It does not lose the beat; it lags, by very nearly the rate divided by the period-correction gain, and the lag reaches a quarter of a beat at 1.5 per cent a beat — at which point the tracker is nearer the wrong onset than the right one.
Fig. 6 The same three rates with the period-correction gain at 0.05, the bottom of the published range. The lags are four times larger — 0.315 of a beat at 2 per cent, which is already past the quarter-beat boundary — so a listener who corrects slowly loses a tempo change that a listener who corrects quickly follows without noticing.

A closing ritardando is faster than that

Here is where the number becomes a claim about music rather than about a simulation.

A cadential ritardando that halves the tempo over the last eight beats — an ordinary, unremarkable ending — requires the period to grow by 9.05 per cent a beat. That is above the limit at every gain in the published range but the very top of it, and which gain decides it turns out to be the one no figure here varies.

Every limit quoted above is computed at a phase-correction gain of 0.6, and phase correction has a published range of its own — roughly 0.5 to 0.9. Running the corners:

β = 0.05 β = 0.30
α = 0.5 1.46% 8.59%
α = 0.9 1.65% 11.16%

At the fast end of the period gain the phase gain decides the ritardando, and the 9.05 per cent falls between the two corners. A listener at α = 0.9 and β = 0.3 follows a tempo halved over eight beats with a quarter-beat of lag to spare; one at α = 0.5 with the same period gain has lost it. At the slow end of the period gain α barely matters — 1.46 against 1.65 — because there the correction is too small for the phase term to do anything with.

So the two gains are not independent knobs of equal standing. β sets the scale of the limit and α modulates it by up to a fifth, and the modulation is largest exactly where the musical question is decided.

The same halving spread over sixteen beats requires 4.43 per cent, which is inside the limit at gains of 0.15 and above. And a modest slowing, 120 to 100 over four beats, needs 4.66 per cent, which is the same territory.

One, two, four and eight bars are spans of a few seconds each across any ordinary tempo, and a ritardando occupies a span of that order. The span is what decides whether the change is followable: the same halving of tempo is easy over eight bars and impossible over one, because what a tracker needs is not a small change but a slow one.

So the model says something a concertgoer already knows and nobody had a number for: a big final rallentando is not followable, and it is not meant to be. The beat dissolves, which is exactly the effect — an ending is a place where the machinery is allowed to stop, and a tempo change fast enough to break entrainment is one of the ways of stopping it.

A slower ritardando, spread over four bars rather than two, stays inside the limit and reads as expressive rather than final. Conductors distinguish these routinely and by feel.

Which listener is doing the following

There is one more thing the gain does, and it is the most interesting consequence in this essay.

The period-correction gain is fitted per person, and it varies. Musicians correct faster than non-musicians; attention raises it; the published range spans a factor of six from bottom to top.

Put that beside the ritardando arithmetic and it says that two people in the same room lose the beat at different points in the same rallentando. At a gain of 0.3 the limit is 9.5 per cent a beat and a tempo halved over eight beats is just barely followable; at 0.1 the limit is 3.1 and it is long gone. Neither listener is wrong and neither is inattentive; they have different constants.

the levels available while the tempo is changing. The range of inter-onset intervals that can be heard as a beat at all, from about 100 to 2000 milliseconds, with the preferred rate near 550. Each mark is one metrical level of a piece at 120 beats a minute. Which of them a listener taps is decided by which falls nearest the preferred rate, not by which one the notation calls the beat.
Fig. 7 The metrical levels at 120 to the minute against the window. A ritardando drags every one of these to the right at once, and a listener whose tracker cannot follow the beat may still be able to follow the level above it — which is slower, moves the same proportion, and starts from a different place in the window.

That is testable in the cheapest possible way — play a ramp, ask people to tap, find where they fall off — and it predicts that where they fall off correlates with how fast they correct on a separate synchronisation task. Whether anybody has run it in that form is not something this site can establish.

What it still cannot do

The oscillator fixes the memory problem completely and it does not fix the other one.

A tracker holds one period, so it can lock onto a metre whose beats are all the same length and cannot represent one whose beats are not. 9/8 as 2+2+2+3 has no single number for a beat, and the tracker’s beat is a single number that moves only slowly — so an additive metre is not a hard case for it, it is outside what it can say at all.

An oscillator has one period. A bar of unequal beats needs a list, and swapping a scoring function for an oscillator does nothing about that. What would be needed is two coupled oscillators at a fixed ratio, or one oscillator with a phase-dependent period, and neither is in the figures on this page.

Nor does the tracker have any notion of a bar. It finds a beat; the level above the beat, the phase of the bar, and the group above the bar are all outside it. Real models run a bank of these at related periods and let them constrain each other, which is a straightforward extension and a different essay.

One pattern, four metres. The same 16-step onset pattern read under 3 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are a beat every 2 12, a beat every 4 -4, a beat every 8 0. At a bar of 2240 ms each candidate's beat also has a rate — a beat every 2 at 280 ms, a beat every 4 at 560 ms, a beat every 8 at 1120 ms — and weighting the fit by how near that rate is to the preferred 550 ms gives a beat every 2 0.91, a beat every 4 -0.99, a beat every 8 0.00, so a beat every 2 wins on the two rules together. Nothing about the sound differs between these readings; the bar line is supplied by the listener.
Fig. 8 What the scoring model is still good at: choosing among periods, over a whole pattern, with a rate term. The tracker has no opinion about which level of a hierarchy to lock onto — it locks onto whatever it was started at — so the two models are complementary rather than rivals.

That is the honest summary of the pair. The rules choose and cannot keep; the tracker keeps and cannot choose. A model of a listener needs both and the ones that work have both.

The accelerando, which is not symmetrical

Everything above has been a slowing. A speeding-up runs the same arithmetic with the sign reversed and the tracker settles ahead of the music by the same fraction of a beat, which is a different experience.

Running behind, the listener hears each onset slightly before they expected it, which is a series of small surprises in the direction of urgency. Running ahead, they expect each onset slightly before it arrives, which is a series of small delays. The magnitudes are identical and the reports are not, and this model has no term for the difference because a signed error is all it carries.

There is a second asymmetry that is in the model. A ritardando moves every metrical level toward the slow edge of the window and an accelerando moves them toward the fast edge — and the two edges are different distances away from the preferred rate. A beat at 550 milliseconds has a factor of 3.6 of slowing before it falls off the slow edge and a factor of 5.5 of speeding before it falls off the fast one, so there is more room to accelerate than to slow down before the level a listener is holding stops being available at all.

That is a prediction about where a listener switches level rather than where they lose the beat, and it says the switch should come sooner in a ritardando. Which is what happens: a slow movement is felt in subdivisions and a fast one in bars, and a piece that slows down enough recruits its subdivision as the new beat well before the old beat becomes untrackable.

What the picture cannot show

It cannot show a real onset stream. Every figure feeds the tracker onsets on an even grid with Gaussian noise. Real onsets are systematically displaced rather than randomly, and the displacements carry metrical information, which this tracker treats as error to be corrected away.

It cannot show the gains varying. α and β are constants here. In tapping studies they vary with tempo, with training and with whether the listener is attending, and a model in which the correction gain is itself adaptive would behave differently at every point on these curves.

It cannot show a ritardando that is not exponential. Every ramp here multiplies the period by a constant factor per beat. Measured ritardandi are closer to linear in tempo than in period, and the difference matters most at exactly the end of the ramp, where the limit is being approached.

It cannot show what happens after the beat is lost. The tracker keeps correcting past the quarter-beat boundary and simply locks onto a different onset. A listener does something else — usually stops tracking and waits — and there is no term for stopping in this model.

It cannot show the ensemble. A tracker following a recording is one listener. Players following each other are correcting toward one another, which is a coupled system with different dynamics, and is why an orchestra can execute a ritardando that no individual member could follow from outside — each is correcting toward the others rather than tracking a fixed stimulus.

And the quarter-beat boundary is a stipulation. It is a reasonable one, and it is not a measurement. The lag is analytic; the point at which a lag counts as a loss is a choice, and moving it to a third of a beat moves the ritardando limit from 6.3 per cent to about 8.4.

Where the ladder ends

Eight rungs. The first three established that the beat is inferred rather than given, that the inference has a preferred rate, and that the same computation returns the bar above the bar. The next four took the machinery apart: it answers the period confidently and the phase not at all; the quantity it makes computable — syncopation — belongs to a pair rather than to a rhythm; it cannot keep a beat through a bar that does not mark it; and it cannot propose a beat that is not always the same length.

This one supplies the missing half. A metre is a state, carried forward, corrected by what arrives — which is why it survives silence, why it follows a tempo at a computable distance behind, and why a fast enough change ends it.

One thing this rung inherits and does not repair is the phase. The tracker is started at a beat time, and where that beat time came from is the question the fourth rung could not answer: an oscillator locks onto whatever it was started at and has no more opinion about the phase than the rules had. What it adds is that once a phase is chosen it is kept, which turns a bad guess into a persistent one.

What the ladder has not got, and what it now knows it has not got, is a model in which the beat is a list rather than a number. That is one rung’s worth of arithmetic and several papers’ worth of combinatorics, and it is the obvious place for a ninth.

Part 8 of 9

One essay in the series on metre induction. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 11.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

BeatClosureEntrainmentExpectationMetreTempo