Form and structure

A detector whose resolution the performance sets

The boundary detector lost its free parameter when the psychological present became a number of notes at a stated tempo, and what that held still was named at the time: a performance slows into a phrase end, so the number of notes inside the present is not the same everywhere in a tune — it falls exactly where a boundary is. Making the width follow the performance recovers half of what removing the parameter cost, and honestly leaves the other half.

Assumes: The level the tempo chooses · A boundary at a stated level

This ladder has spent two rungs removing one number.

A boundary at a stated level introduced a scale parameter — the width of the Gaussian that smooths the boundary-strength curve — and used it to turn a threshold into a hierarchy, which removed the threshold and replaced it with a σ. The level the tempo chooses then removed the σ: the psychological present is a number of seconds, a tempo converts seconds into notes, and the level a listener reads is therefore not free at all.

That rung ended by naming what it had done to get there. It held the tempo constant. A real performance does not: it slows into a phrase end, by twenty to fifty per cent on the phrase-final note, which is the one thing every measurement of expressive timing agrees about. So the number of notes inside the present is not the same everywhere in a tune, and it falls exactly where a boundary is.

How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.
Fig. 1 How far the detector looks, note by note, once the performance has lengthened its phrase-final notes by thirty per cent. The window runs from five notes to ten against a constant seven for the unperformed version, and it dips at the notated boundaries — because a boundary is where the performance slows, and a slower stretch fits fewer notes into a fixed number of seconds.

What a locally adaptive scale is

Scale-space analysis has a standard form: one signal, one smoothing width, and a family of readings indexed by the width. Everything about it assumes the width is the same everywhere, which is what makes the family one-dimensional and what makes persistence — the largest width at which a feature survives — a well-defined ordering.

A width that varies along the signal is a different object. There is no single family and no persistence in the usual sense, because two neighbouring points are being read at two different resolutions. What there is instead is one reading, in which the detector’s own resolution is a property of the data rather than a parameter over it.

That is a stronger claim than it sounds. It says the listener does not choose a level: the performance chooses it for them, moment by moment, by controlling how much music fits inside the window the listener has.

What it buys, and what it costs

Three readings can be compared and they have to be, because two of them are answering different questions.

fixed width, fitted to the tune       F = 0.71
fixed width, at the present           F = 0.36
adaptive width                        F = 0.55

The first is the free parameter the fourth rung was trying to get rid of: tune σ until the detector agrees best with the notation. It wins, and it wins for a reason that is not to its credit — it has been given the answer to fit to.

The second is the fifth rung’s principled setting, with no freedom in it at all: the present is 3.5 seconds, the tempo is 120, so the window is seven notes and there is nothing to choose. It is much worse.

The third is this rung’s, with the same lack of freedom and one extra piece of information — the performance. It recovers about half of what the second reading gave up.

That is the honest result: removing a free parameter costs accuracy, letting the performance set the resolution recovers part of the cost, and neither principled reading matches a fitted one. A rung that claimed the adaptive detector wins outright would be claiming it had got something for nothing.

How much rubato it takes before adapting the scale pays. Three readings of one tune against how much a phrase-final note is lengthened. The flat upper line is a fixed-width detector at whatever width suits this tune best, which is the free parameter an earlier essay was trying to remove. The flat lower line is a fixed-width detector at the width the psychological present dictates, which is the principled setting arrived at later. The third is the position-dependent detector. The adaptive reading beats the principled fixed one at every amount drawn, including none at all — the window already varies without any rubato, because the tune's own written durations are not equal — and it reaches 0.55 against 0.36. Published phrase-final lengthening is 20 to 50 per cent, marked, and the adaptive curve is at its highest inside that band. Neither reading reaches the fitted one, which is the honest cost of removing a free parameter.
Fig. 2 The three readings against how much rubato the performance has. The two fixed lines are flat by construction — neither knows about the timing at all — and the adaptive one is not. The shaded band is the published range of phrase-final lengthening, and the adaptive reading is at its best inside it.

The circularity, and what is left after it

There is an obvious objection and it has to be met head on: the lengthening was applied at the notated boundaries, so of course a detector that narrows where the tune slows finds the notated boundaries.

That is exactly the objection the fifth rung met when it constructed a performance rather than measuring one, and its answer applies here. The circularity would be fatal if the reported quantity were whether the detector finds the boundaries. It is not. What is reported is by how much, against how much rubato, and the comparison is with a fixed-width detector reading the same performed timing — which has the same information available and does not use it.

There is also a piece of the effect that owes nothing to the construction. At zero rubato the adaptive window still varies, from five notes to eleven, because the tune’s own written durations are not equal. A tune with long notes in it has stretches where fewer notes fit inside three and a half seconds whether or not anybody plays it expressively. The adaptive reading beats the principled fixed one even there, and that part of the result is entirely free of the construction.

What the circularity does cost is any claim about how much rubato is needed. The curve rises with lengthening because the lengthening is at the right places by construction, so its slope is not evidence.

Frère Jacques, phrased at the level each tempo selectsThe number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 7 boundaries, drawn as the flat line; the model matches it best at 40 beats per minute, where the present holds 2.3 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.the page's own phrasing402.3n603.5n804.7n1207.0n1609.3n20011.7n26015.2n02468boundaries the model findsbeats per minute, with the notes inside the present under itbest at 40 bpmF = 0.71the scale is the presentdivided by the note length
Fig. 3 The earlier figure, which is the reading this essay modifies: the boundaries the tempo-selected level gives, across tempos. Every column here uses one width for the whole tune, and the prediction that made it interesting — that the same tune at two tempos is heard as differently phrased — survives unchanged, because the adaptation is about variation within a tune rather than between tempos.

What the window does between the boundaries

The dips are the obvious feature and they are not the only one. Between the boundaries the window widens, and it widens most where the notes are shortest.

That has a consequence the fixed reading cannot have. A fast run in the middle of a phrase puts many notes inside the present, so the detector smooths over many of them and does not report boundaries inside it — which is right, because a fast run is one gesture. A slow passage puts few notes in, the detector looks at a handful at a time, and it becomes willing to find structure at a finer grain.

So the same tune is segmented at different grains in different places, and the grain follows the note density rather than the bar line. That is a claim about listening that a fixed-width model cannot make and that matches an ordinary observation: a listener asked where the units are in a fast passage names larger units than in a slow one, in the same piece.

It is also the same shape as the finding in one of these eight-bar phrases accelerates, which is that a phrase is a number of seconds rather than of bars. Both say the clock is in seconds and the notation is not.

Twinkle, twinkle at 140, across the width of the present. The boundaries the model finds when its scale is set by the psychological present, at the three ends of that present's published range — 2, 3.5, 8 seconds — with the notation's own phrase ends on the top row. The scale in notes is the present divided by the mean note length, so the same band is a different number of notes at every tempo. At 3.5 seconds the model finds 5 boundaries against the page's 5.
Fig. 4 The other earlier figure: the same tune at one tempo read across the width of the present, from two seconds to eight. Every column is a fixed width and the range between them is the model’s remaining uncertainty — the one parameter that has not been removed and, being a fact about listeners rather than a choice, cannot be.

The tune where it does nothing

Running the adaptation on a tune whose written durations are nearly all equal produces almost no variation in the window, and therefore almost no difference from the fixed reading.

That is the control this rung needs and it is worth stating as a result rather than as a caveat. The adaptation has an effect exactly in proportion to how much the note durations vary, from any source — written or performed — and where they do not vary it degenerates gracefully into the fifth rung’s fixed reading. A modification that quietly changed the answer on a tune with nothing to adapt to would be a modification doing something other than what it says.

How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 6 notes against a constant 6 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.91 with the notated phrasing, against 0.91 for the fixed width the present dictates and 1.00 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.
Fig. 5 The same computation on a tune of nearly uniform durations. The window barely moves except at the phrase ends the constructed performance lengthens, and the adaptive reading is within a hundredth of the fixed one. Where there is nothing to adapt to, the adaptation does nothing.

All three tunes, and the split the caveats said could not be made

Two of the claims above are stated qualitatively and both can be given numbers, because the collection has three melodies and running all three costs nothing.

tune fitted at the present adaptive share of the gap recovered window, in notes
Frère Jacques 0.714 0.364 0.545 52% 5 to 10
Twinkle 1.000 0.909 0.909 0% 5 to 6
Ode to Joy 0.667 0.333 0.333 0% 5 to 7

The dependence is not a proportionality; it is a threshold. The two tunes that gain nothing have a window varying by nine per cent of its own mean, and the one that gains half the gap has one varying by twenty-six. Below some amount of variation the adaptive reading is not merely a smaller improvement — it finds exactly the same boundaries as the fixed reading, to the last note. Three tunes cannot say where the threshold is, and they can say that there is one rather than a slope.

That also puts the headline in its place. The 0.55 in the summary above is one tune’s number, and the average over the three melodies this collection has is 0.60 against the fixed reading’s 0.54 — a gain of six points rather than nineteen, carried entirely by one melody.

The second claim is the one the caveats give up on: that the window varies for two reasons, the written durations and the performed lengthening, and that the figure draws only their sum. Separating them takes one run rather than a curve. Set the rubato to zero and Frère Jacques’ window still runs from five notes to eleven, purely from its own written durations, and the adaptive reading scores 0.400 against the fixed 0.364.

So the accounting is: of the 0.350 gap between the principled fixed reading and the fitted one, the adaptive detector recovers 0.181 — and 0.036 of that, a fifth of it, owes nothing at all to the constructed performance. Four fifths does.

That is a less comfortable number than the section on circularity implies, and it is the right one to publish. The construction-free part of the result is real, it is the part the objection cannot touch, and it is small: a tenth of the way from the principled reading to the fitted one. The rest is the performance, placed at the boundaries by assumption, doing what it was built to do.

Which computation produced the numbers

The boundary-strength curve is melodyBoundaries, unchanged since where a phrase ends: a weighted sum of the three cues that rung’s model uses, evaluated between every pair of adjacent notes.

The performance is performanceBoundaries’ own construction — every phrase-final note lengthened by a stated fraction of itself, with everything else at its notated duration.

The local window is a count: starting at each note, how many notes fit inside the present, taken forward and then backward until the seconds run out. The smoothing width is that count over four, so a window of n notes is about two standard deviations either side, which is the same convention the fifth rung used to turn its note count into a σ.

The smoothing itself is a Gaussian whose width is read from an array rather than from a scalar, normalised pointwise so the varying width does not change the curve’s total. The peaks are found the same way boundaryHierarchy finds them, with the same minimum separation of three, so nothing about the peak-picking differs between the three readings.

Agreement is the F score against the notated phrasing at a tolerance of one note, which is the fourth rung’s measure. The fitted control sweeps σ from 0.3 to 8 in forty steps and reports the best it finds, so it is given every advantage.

Why a boundary narrows the window rather than widening it

There is a direction to check here, because the opposite would have been just as plausible in advance.

A phrase-final note is longer, so fewer notes fit inside the present, so the window is narrower in notes. A narrower smoothing keeps more detail, so the detector is more willing to report a peak — and a peak is what a boundary is. So the performance’s own gesture makes the detector more sensitive precisely where the gesture is.

The alternative story — that slowing down should make a listener take in more structure, not less — is about a window measured in bars rather than in seconds, and this ladder’s first rung is the argument against it: a phrase is a number of seconds, and a listener’s present does not stretch when the music slows.

So the sign of the effect is a consequence of the ladder’s own first result, which is a small piece of internal consistency worth having. A model in which the window were counted in bars would predict the opposite and would have to explain why a ritardando makes an ending harder to hear.

The boundaries that survive each amount of smoothing. The local boundary strengths of Frère Jacques read at every scale: the curve is smoothed with a Gaussian of the width on the horizontal axis and the peaks that survive are counted. Small scales give 7 boundaries and large ones give one, and the notation marks 7. The level with that many falls at a width of 0, where the model finds 71 per cent of the notated boundaries and 71 per cent of what it finds is notated — a comparison with no threshold in it, which is what the scale parameter buys.
Fig. 6 The earlier hierarchy for the same tune: how many boundaries survive each amount of smoothing, with the level whose count matches the notation’s marked. Everything in this picture is one width per column, and the adaptive reading is a single point that does not lie on any of them.

Where the model stops

The performance is constructed and the ladder has now said so four times. This collection has no performance timing data, and the whole apparatus above is a model of expressive timing rather than a measurement of one. The published range of phrase-final lengthening is real; its placement at exactly the notated boundaries is an assumption, and a real performer lengthens at some boundaries and not others.

The present is one number. Three and a half seconds is the middle of a range that runs from two to eight, and the fifth rung’s own figure is about that width being the model’s remaining uncertainty. The adaptive width inherits it: a listener with a two-second present gets a window that varies over a different range and a different reading.

And it is one tune. Two of the three melodies this collection carries show no gain from the adaptation at all — one of them is drawn above, and the other is the tune the whole ladder has been tested on — because their notated durations are nearly uniform and the constructed rubato is the only source of variation. The gain is real where the window varies and absent where it does not, which is at least the right dependence — but three tunes is three tunes.

What the picture cannot show

It cannot show a listener adapting. The model changes its resolution because the arithmetic of “how many notes fit in three and a half seconds” changes. Whether a listener’s segmentation window works that way — a fixed duration, filled with whatever arrives — is the fifth rung’s assumption, and it is an assumption about a mechanism nobody has measured directly.

Nor can it produce a hierarchy. The whole apparatus of persistence, which is the fourth rung’s contribution, needs a family of readings at increasing widths. A locally adaptive width gives one reading, so the ordering that rung produced is not available here — which is a real loss and is the reason this is a rung beside that one rather than after it.

The two sources of variation separate at one point and not along a curve. The window varies because the notated durations vary and because the performance lengthens; the figures draw their sum, and the section above splits it at zero rubato. What one point cannot give is how the split changes with the amount of lengthening, which would need the whole sweep recomputed against a written-duration-only baseline at each step.

Whose performances, and when

Phrase-final lengthening of twenty to fifty per cent is a measurement of Western art-music performance, mostly of piano playing and mostly of the nineteenth-century repertoire, and it is one of the most robust findings in the whole literature of expressive timing. It is also not universal: a great deal of music is played to a grid on purpose, and in that repertoire the adaptive detector and the fixed one are the same detector.

That is a useful boundary on the claim. The argument here is that a performance can hand a listener the scale at which to segment, and it applies exactly where performances are free to do so. In a tradition where the timing is fixed — by a click, by a dance, by an ensemble too large to bend — the information is simply not there, and a listener has to fall back on the fixed present the fifth rung computed. That is a repertoire boundary rather than a modelling one, and it is the same boundary the microtiming ladder draws for a different quantity.

Where this ladder goes next

Six rungs. A phrase is a number of seconds; the sentence and the period differ by a computable ratio; a boundary detector agrees with the page where the tune uses the cue it weights; the detector has a scale, which removes its threshold; the scale has a unit, which removes the parameter; and now the unit is a function of position, set by the performance.

What is owed after this is the hierarchy the adaptation costs, and a family of readings is the rung that pays it. A varying width gives one reading and no persistence, and persistence is what made the fourth rung’s boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than a family of fixed ones — a present of two seconds, of three and a half, of eight, each producing its own varying window — which is a family indexed by the one parameter the ladder has left and cannot remove, since the width of the psychological present is a fact about listeners rather than a choice.

Part 6 of 8

One essay in the series on phrase. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Perceptual presentPhraseRubatoScale-spaceSegmentationTempoTiming deviation