Rhythm and metre

A quantity resting against a wall

The short note of a swung pair was found sitting exactly on the fast edge of the tempo window. If that edge is a floor rather than a fitted number, then every swing statistic computed until now was computed on a censored sample — and a censored sample has a mean that is 0.80 standard deviations too high, a spread that is 40 per cent too low, and a correction gain that can come out twice what the players actually have.

Assumes: The short note is sitting on the floor · Two players and no clock

The short note is sitting on the floor ended by naming its own next rung and saying it needed no corpus. The finding it ended on was that swing’s fitted hundred-millisecond short note and the fast edge of the tempo window are two independently measured constants that agree to within the precision either is known to — and that if they are the same number, the short note is not a parameter of a swing model at all but a wall the model has been resting against.

This rung takes that seriously and asks what a wall does to a measurement.

The distribution that can be measured is not the one the player hasA player aiming at a short note of 100 milliseconds with a standard deviation of 15, against a floor at 100. The pale curve is what the player is doing and the heavy one is what can be recorded, because the 50 per cent of the parent below the floor arrives at the floor instead. The measured mean is 112.0 milliseconds rather than 100 and the measured standard deviation is 9.0 rather than 15. At 160 beats per minute that turns an intended swing ratio of 2.75 into a measured 2.35.intended 100measured 112608010012014000.020.040.06the short note's length, millisecondsdensitythe floor, 100 ms50% of the parentis on the wrong sidespread 15 → 9.0 msratio 2.75 → 2.35at 160 bpm
Fig. 1 A player aiming at a hundred-millisecond short note with a standard deviation of fifteen, against a floor at a hundred. The pale curve is what the player is doing; the heavy one is what can be recorded, because the half of the distribution below the floor arrives at the floor instead. The measured mean is 112 milliseconds rather than 100 and the measured standard deviation is 9 rather than 15. The buttons play the intended pair and the measured one.

The measured mean is high, the measured spread is low, and neither error is noise. Both are exact consequences of the wall, both have closed forms that have been known since Pearson, and both point in the direction that makes a published finding look stronger than it is.

The reason this is a rung rather than a footnote is that all three errors go the same way. A wall makes a quantity look larger than it is, steadier than it is, and more tightly coupled to its neighbour than it is — and steadier and more tightly coupled are the two things performance research is usually looking for.

The arithmetic, which is one line each

Write the intended value as μ, the player’s own variability as σ, and the wall at a. Put α = (a − μ)/σ and λ = φ(α)/(1 − Φ(α)), where φ and Φ are the normal density and its integral. Then:

The observed mean is μ + σλ and the observed variance is σ²(1 − λ(λ − α)).

At the case that matters — the intention sitting exactly on the wall, α = 0 — λ is 0.798, so the observed mean is 0.80 standard deviations high and the observed standard deviation is 0.603 of the true one.

Those are the formulas for a sample with the sub-floor attempts thrown away rather than piled up at the floor, which is not what a player does, and the next rung of this ladder is the correction: the mean bias for a process that piles up is exactly half of the figure above. The direction of every argument below survives it and the sizes halve. The formulas are left standing here because the rung that follows is about the difference between them, and because every relative claim in this essay — that the errors all point one way, that the spread is the worst affected — is unchanged by a common factor.

Both errors, against how far the intended value clears the wall. The two biases a lower wall introduces, in units of the parent standard deviation, against the gap between the intended value and the wall in the same units. Where the intended value sits exactly on the wall the measured mean is 0.78 standard deviations high and the measured spread is 0.61 of the true one — a 40 per cent understatement with nothing added to or removed from the process. Both errors are under a tenth once the gap reaches two standard deviations, so this is a question about one ratio and not about tempo, style or instrument.
Fig. 2 Both errors, against how far the intended value clears the wall, in units of the player’s own spread. The whole question is one ratio: at a gap of zero the mean is 0.80σ high and the spread is 0.60 of the truth; by two standard deviations of clearance both errors are under a tenth. Nothing here depends on tempo, on style or on instrument — only on where the intention sits relative to the floor.

A forty per cent understatement of the spread, with nothing added to the process and nothing taken away. That is the number to carry, because consistency is the thing timing studies report most often and it is the thing a wall corrupts most.

What it does to the swing ratio

The ladder’s own headline is that swing is not two to one: the measured ratio of the long note to the short is well below the notated triplet at every tempo above a slow ballad, and falls further as the tempo rises.

If the short note is on a floor, part of that is the wall.

The measured short note is 12 milliseconds longer than the intended one, at every tempo, and the ratio is the rest of the beat divided by the short note — so a longer measured short note is a smaller measured ratio. At 160 beats per minute an intended 2.75 comes out as 2.35. At 240 it comes out as 1.23 against an intended 1.50.

The distribution that can be measured is not the one the player hasA player aiming at a short note of 118 milliseconds with a standard deviation of 15, against a floor at 100. The pale curve is what the player is doing and the heavy one is what can be recorded, because the 12 per cent of the parent below the floor arrives at the floor instead. The measured mean is 121.3 milliseconds rather than 118 and the measured standard deviation is 12.4 rather than 15. At 200 beats per minute that turns an intended swing ratio of 1.54 into a measured 1.47.intended 118measured 1218010012014016000.020.040.06the short note's length, millisecondsdensitythe floor, 100 ms12% of the parentis on the wrong sidespread 15 → 12.4 msratio 1.54 → 1.47at 200 bpm
Fig. 3 The same figure with the intention placed 1.2 standard deviations above the floor rather than on it, which is what a player with a slightly longer target and the same variability would do. A fifth of the distribution is still clipped, the mean bias falls to 4 milliseconds and the spread recovers most of the way. The size of the effect is entirely a question of where the intention is, and nothing in a recording says where the intention is.
How much swing there is, against how fast the music goes. The ratio of the long note to the short note of a swung pair, as a function of tempo, derived from the finding that the short note holds a roughly constant hundred milliseconds. The notated readings are horizontal lines, and each is correct at exactly one tempo. Above about three hundred beats a minute the ratio reaches one and the swing has gone.
Fig. 4 The established curve: the swing ratio against tempo, on the model that holds the short note at a constant hundred milliseconds. Every point on it is a measured ratio, so every point on it carries the twelve-millisecond bias — which is a downward shift of the whole curve, largest at the fast end where the short note is the largest share of the beat. The shape of the curve is the model’s and is unaffected; its height is not.

This does not overturn the finding. The measured ratio is below 2:1 by far more than 12 milliseconds of short note accounts for, and the ratio falls with tempo for the reason the ladder gave, which is that the short note holds a constant length while the beat shortens. What it does is put a bias of known sign on every number, and the bias is largest exactly where the ladder’s claim is strongest.

It is also worth being precise about which of the ladder’s claims is at risk. That the deviations are not noise is untouched: a systematic offset survives any monotone distortion of the measurement. That the ratio falls with tempo is untouched, because the bias is nearly the same at every tempo. What is at risk is the value of the ratio at any given tempo, and the reported consistency of it — and the second is the more damaged, because it is off by a factor rather than by an additive constant.

What it does to the two-player result

The worse consequence is not in the swing ladder at all.

Two players and no clock built a model in which each player nudges their next onset toward the other, and measured the resulting asynchrony. That rung recorded its own limitation honestly: only the sum of the two gains is observable, so the measurement cannot say who is following whom.

A wall adds a second thing it cannot say, and this one is worse because it changes the number rather than hiding a decomposition.

What a one-sided wall does to a recovered correction gain. Four pairs of true correction gains, each simulated twice: once with both players free to correct in either direction, and once with the first player unable to shorten an interval that is already at the floor. The open mark is the lag-one autocorrelation of the asynchrony recovered from the free pair and the filled mark is the same statistic from the walled pair. The walled pair returns a larger number in every case — up to 3.0 times larger — from players whose gains have not changed at all.
Fig. 5 Four pairs of true correction gains, each simulated twice: freely, and with the first player unable to shorten an interval that is already at the floor. The recovered lag-one autocorrelation is larger in every case, by up to a factor of three, from players whose gains have not changed at all. The proportion of corrections that could not be made is printed under each row and runs from 59 to 73 per cent, because a player on the floor is refused about half of their corrections and then drifts, which makes the next one larger.

A correction that cannot be made in one direction is not a correction, and a series of half-refused corrections looks like a stronger coupling than the players have. Two players with gains of 0.3 each return a lag-one autocorrelation of 0.23 when free and 0.43 when one of them is against the wall, which would be read as a substantially tighter duo.

The direction is the alarming part. Every method for recovering coupling from performance data reads a larger autocorrelation as stronger mutual correction, and a floor produces a larger autocorrelation from less correction.

The wall in the other ladders

Once the shape is named it turns up in more than one place in this collection, and it is worth marking them so that the next rung to meet one recognises it.

The tempo window is itself two walls — a fastest and a slowest rate at which a listener will take something as a beat — and every measurement of a preferred tempo is a measurement inside them. A preferred-rate distribution reported from a population that includes people trying to tap faster than the window allows is censored at the fast end in exactly the way above.

The duration categories have edges, and the previous rung found that each ratio has a last tempo. A measurement of where a category’s centre sits, made near that tempo, is a measurement against a wall.

And the aksak metres’ unequal beats are bounded below by the same window at one end and by the inference of a metre at the other. The bound there was found to be arithmetic where it was expected to be perceptual, which is a different result, but it is the same picture: a quantity with a hard edge near where it is usually measured.

None of those is wrong. What they share is that the edge is part of the object rather than part of the noise, and a statistic that treats it as noise reports a narrower, later, tidier world than there is.

The metrical levels of 120 bpm, against the window. The range of inter-onset intervals that can be heard as a beat at all, from about 100 to 2000 milliseconds, with the preferred rate near 550. Each mark is one metrical level of a piece at 120 beats a minute. Which of them a listener taps is decided by which falls nearest the preferred rate, not by which one the notation calls the beat.
Fig. 6 The window the floor comes from: the rates at which a division of the beat can be taken as a beat, with the metrical levels of a bar at 120 laid inside it. The fast edge at a hundred milliseconds is the wall this whole essay is about, and it is drawn here as what it was measured as — a limit on the listener rather than on the player.

Which computation produced the numbers

Three things, of which two are closed-form and one is a simulation.

The truncated-normal mean and variance are evaluated directly, with the normal integral coming from the site’s own error function — the same one the categorical-hearing ladder uses to count how many categories fit in an octave. Nothing is sampled and nothing is fitted.

The swing consequence is that arithmetic applied to this site’s own swing model: the short note is held at a constant absolute length, the pair fills one beat, and the ratio is what remains divided by what is measured.

The duet result is the previous rung’s own simulation with one line added. Each player’s next interval is its own timekeeper plus a correction of minus the gain times the asynchrony; for the walled player, a correction that would shorten the interval is set to zero and counted. Everything else — the timer noise at 12 milliseconds, the motor noise at 6, the two hundred beats, the sixty seeds — is exactly what the previous rung used, so the comparison is between two runs that differ in one clause.

The soft-wall table is quadrature over the same normal parent under two transformations, each with one parameter and each reducing to hard censoring as that parameter goes to zero. The resistive wall is a softplus, which is the standard smooth approximation to a maximum; the blurred wall is the hard maximum against a floor that is itself normally distributed, integrated over both. Neither is offered as a model of a hand. They are chosen because they are the two ways a wall can stop being sharp — the barrier bends, or the barrier moves — and because they bracket the question rather than answer it.

The reason to compute both is that they disagree in sign on two of the three moments, which is the whole result. A single softening would have produced one set of numbers and an unearned impression that “the soft case” had been priced. What was actually available to be priced is that softening is not one thing, and that the diagnostic the ladder is heading for survives one kind and not the other.

Where the model stops

Nothing here shows the floor is a floor. That is the previous rung’s conditional and it remains one. Two constants agreeing to within their own precision is weaker evidence than it reads as, and this rung’s entire content is if it is a floor, then these are the errors. Both halves are worth having: an unconditional claim would be dishonest and a conditional claim with no arithmetic in it would be useless.

The wall is modelled as hard and it will not be, and softening it does not do what this rung originally guessed. A player asked for a note shorter than the window’s edge does not produce exactly the edge; they produce something near it. That was recorded here as giving smaller biases in the same directions. Running it gives larger biases, and the two natural readings of “softer” disagree about everything else.

There are two, and they are different processes rather than two descriptions of one. A resistive wall makes approaching the floor progressively harder, so attempts near it are pushed up. A blurred wall leaves the floor hard but lets its position vary from attempt to attempt, which is what a fluctuating motor limit would do. Both reduce to hard censoring as the softness goes to zero. With the parent’s mean on the wall and the softness in units of the parent’s spread:

softness resistive: bias / spread / skew blurred: bias / spread / skew
0 (hard) 0.399 / 0.584 / 1.64 0.399 / 0.584 / 1.64
0.25 0.437 / 0.567 / 1.64 0.411 / 0.602 / 1.31
0.5 0.534 / 0.545 / 1.51 0.446 / 0.653 / 0.68
1.0 0.806 / 0.521 / 1.15 0.564 / 0.826 / 0.14

The mean bias grows under both, so softening the wall makes the first error worse rather than better. The two then part company on the spread: a resistive wall understates it further, from 0.584 of the truth to 0.521, and a blurred wall recovers most of it, from 0.584 back to 0.826. So “smaller biases in the same directions” is wrong about the direction of the first and wrong about the existence of a common direction at all.

The consequential column is the third. A blurred wall very nearly erases the skew — 1.64 falls to 0.14 at a softness of one spread, which is indistinguishable from a free sample. That matters because the skew is the diagnostic the last section of this essay proposes and the next rung builds on: a pile-up at the floor is what a censored sample looks like, and a floor that wobbles by about the player’s own variability produces no pile-up to see. So the shape test can only detect a wall that is sharp compared with the player’s spread, and a negative result from it is evidence of a sharp wall’s absence rather than of any wall’s absence.

And the parent is assumed normal. Motor timing is close to normal over a few hundred milliseconds and is not exactly so; a parent with a longer lower tail would be clipped harder and biased more. That one has not been computed, and after the soft-wall result above it should not be guessed at either: the lesson of that table is that an intuition about the direction of a correction is worth very little until the correction is evaluated, and the intuition recorded here was wrong about a case that looked easier than this one.

Whose music, and what to do about it

The claim is about a measurement rather than about a repertoire, so it applies to every corpus of performance timing that has ever been analysed — which is mostly jazz, mostly Viennese classical, and mostly piano, because those are the recordings that came with a mechanism for extracting onsets.

The practical consequence is a test rather than a correction, and it is available in data that already exists.

A censored distribution is skewed and a free one is not. The wall produces a pile-up on one side, so the recorded short notes should be asymmetric about their own mean, with a longer tail upward. That is a shape rather than a summary statistic, it survives every uncertainty above, and it needs nothing but the histogram of the short notes — which every one of these studies computed and none of them published as a shape.

And the skew should disappear at slow tempi. A ballad’s short note is long enough to clear the floor by several standard deviations, so its distribution should be symmetric, and the same players at a fast tempo should show the pile-up. Same musicians, same instrument, same room, one variable.

Two kinds of systematic timing, which share a word. Each measured profile split into a constant offset from the grid and a pattern that varies by position in the bar. Viennese waltz, second beat is −10.0 ms of offset and 33.4 ms of pattern; jazz soloist against the ride is 28.8 ms of offset and 1.9 ms of pattern; quantised is 0.0 ms of offset and 0.0 ms of pattern. A motor deviation anywhere in the 8 to 20 ms range published for skilled performers leaves 74–95% of Viennese waltz, second beat's variation systematic, 1–5% of jazz soloist against the ride's variation systematic. The two quantities are independent and no single deviation figure distinguishes them.
Fig. 7 The published profiles, split into their offset and their pattern. The correction argued for here applies only to the first column, because a wall does not know which beat of the bar it is on — so the Viennese waltz’s early second beat is untouched by any of this, and the jazz soloist’s near-constant offset is the case where the whole of it applies.

What the picture cannot show

Whether players aim at the floor or near it. Everything above turns on the gap between the intention and the wall, and the gap is unobservable: what is recorded is the censored distribution, and a censored distribution with a small gap and a large σ looks very like one with a large gap and a small σ. The skew is the only thing that separates them.

Nor whether the parent is stationary. The arithmetic treats the intended value as one number for a whole take. A player who drifts — and the note that has a length is about a quantity that is not constant within a phrase — has a parent that moves in and out of range of the wall, which mixes censored and uncensored stretches in one recording and makes the summary statistics a weighted average of two regimes.

And nothing here is about hearing. A floor at the fast edge of the tempo window is a floor on what a listener can hear as two events, and the argument above treats it as a floor on what a player can produce. Those are different claims about different systems that happen to have arrived at the same hundred milliseconds, and the previous rung said as much when it called the agreement a coincidence between two round numbers.

Where this ladder goes next

Seven rungs. Swing is a ratio and not two to one; the deviations that make a groove are unnotatable; they split into offset and pattern; two players correct toward each other and the measurement cannot say who follows; the deviations sit inside categories whose boundaries are the simple ratios; the categories have tempo ranges and the swing constant may be the window’s own edge; and now, if it is, three published statistics are biased in known directions by known amounts.

The rung after it is the one the last section names and it is a shape rather than a number. Every argument here is about the first two moments of a distribution, and the thing that actually distinguishes a censored sample from a free one is its third — the skew. A model of what a walled timing distribution looks like, rather than of what its mean and spread come out as, would be a prediction with a shape to it, and this collection has the machinery to compute the shape and no corpus to compare it against. That is the same debt the clave rung recorded, arriving at the same ladder from the other end.

Part 7 of 9

One essay in the series on microtiming. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CensoringJitterMicrotimingMotor delayRegression to the meanSwingTempoTiming deviation