A detector whose resolution the performance sets
Assumes: The level the tempo chooses · A boundary at a stated level
This ladder has spent two rungs removing one number.
A boundary at a stated level introduced a scale parameter — the width of the Gaussian that smooths the boundary-strength curve — and used it to turn a threshold into a hierarchy, which removed the threshold and replaced it with a σ. The level the tempo chooses then removed the σ: the psychological present is a number of seconds, a tempo converts seconds into notes, and the level a listener reads is therefore not free at all.
That rung ended by naming what it had done to get there. It held the tempo constant. A real performance does not: it slows into a phrase end, by twenty to fifty per cent on the phrase-final note, which is the one thing every measurement of expressive timing agrees about. So the number of notes inside the present is not the same everywhere in a tune, and it falls exactly where a boundary is.
What a locally adaptive scale is
Scale-space analysis has a standard form: one signal, one smoothing width, and a family of readings indexed by the width. Everything about it assumes the width is the same everywhere, which is what makes the family one-dimensional and what makes persistence — the largest width at which a feature survives — a well-defined ordering.
A width that varies along the signal is a different object. There is no single family and no persistence in the usual sense, because two neighbouring points are being read at two different resolutions. What there is instead is one reading, in which the detector’s own resolution is a property of the data rather than a parameter over it.
That is a stronger claim than it sounds. It says the listener does not choose a level: the performance chooses it for them, moment by moment, by controlling how much music fits inside the window the listener has.
What it buys, and what it costs
Three readings can be compared and they have to be, because two of them are answering different questions.
fixed width, fitted to the tune F = 0.71
fixed width, at the present F = 0.36
adaptive width F = 0.55
The first is the free parameter the fourth rung was trying to get rid of: tune σ until the detector agrees best with the notation. It wins, and it wins for a reason that is not to its credit — it has been given the answer to fit to.
The second is the fifth rung’s principled setting, with no freedom in it at all: the present is 3.5 seconds, the tempo is 120, so the window is seven notes and there is nothing to choose. It is much worse.
The third is this rung’s, with the same lack of freedom and one extra piece of information — the performance. It recovers about half of what the second reading gave up.
That is the honest result: removing a free parameter costs accuracy, letting the performance set the resolution recovers part of the cost, and neither principled reading matches a fitted one. A rung that claimed the adaptive detector wins outright would be claiming it had got something for nothing.
The circularity, and what is left after it
There is an obvious objection and it has to be met head on: the lengthening was applied at the notated boundaries, so of course a detector that narrows where the tune slows finds the notated boundaries.
That is exactly the objection the fifth rung met when it constructed a performance rather than measuring one, and its answer applies here. The circularity would be fatal if the reported quantity were whether the detector finds the boundaries. It is not. What is reported is by how much, against how much rubato, and the comparison is with a fixed-width detector reading the same performed timing — which has the same information available and does not use it.
There is also a piece of the effect that owes nothing to the construction. At zero rubato the adaptive window still varies, from five notes to eleven, because the tune’s own written durations are not equal. A tune with long notes in it has stretches where fewer notes fit inside three and a half seconds whether or not anybody plays it expressively. The adaptive reading beats the principled fixed one even there, and that part of the result is entirely free of the construction.
What the circularity does cost is any claim about how much rubato is needed. The curve rises with lengthening because the lengthening is at the right places by construction, so its slope is not evidence.
What the window does between the boundaries
The dips are the obvious feature and they are not the only one. Between the boundaries the window widens, and it widens most where the notes are shortest.
That has a consequence the fixed reading cannot have. A fast run in the middle of a phrase puts many notes inside the present, so the detector smooths over many of them and does not report boundaries inside it — which is right, because a fast run is one gesture. A slow passage puts few notes in, the detector looks at a handful at a time, and it becomes willing to find structure at a finer grain.
So the same tune is segmented at different grains in different places, and the grain follows the note density rather than the bar line. That is a claim about listening that a fixed-width model cannot make and that matches an ordinary observation: a listener asked where the units are in a fast passage names larger units than in a slow one, in the same piece.
It is also the same shape as the finding in one of these eight-bar phrases accelerates, which is that a phrase is a number of seconds rather than of bars. Both say the clock is in seconds and the notation is not.
The tune where it does nothing
Running the adaptation on a tune whose written durations are nearly all equal produces almost no variation in the window, and therefore almost no difference from the fixed reading.
That is the control this rung needs and it is worth stating as a result rather than as a caveat. The adaptation has an effect exactly in proportion to how much the note durations vary, from any source — written or performed — and where they do not vary it degenerates gracefully into the fifth rung’s fixed reading. A modification that quietly changed the answer on a tune with nothing to adapt to would be a modification doing something other than what it says.
All three tunes, and the split the caveats said could not be made
Two of the claims above are stated qualitatively and both can be given numbers, because the collection has three melodies and running all three costs nothing.
| tune | fitted | at the present | adaptive | share of the gap recovered | window, in notes |
|---|---|---|---|---|---|
| Frère Jacques | 0.714 | 0.364 | 0.545 | 52% | 5 to 10 |
| Twinkle | 1.000 | 0.909 | 0.909 | 0% | 5 to 6 |
| Ode to Joy | 0.667 | 0.333 | 0.333 | 0% | 5 to 7 |
The dependence is not a proportionality; it is a threshold. The two tunes that gain nothing have a window varying by nine per cent of its own mean, and the one that gains half the gap has one varying by twenty-six. Below some amount of variation the adaptive reading is not merely a smaller improvement — it finds exactly the same boundaries as the fixed reading, to the last note. Three tunes cannot say where the threshold is, and they can say that there is one rather than a slope.
That also puts the headline in its place. The 0.55 in the summary above is one tune’s number, and the average over the three melodies this collection has is 0.60 against the fixed reading’s 0.54 — a gain of six points rather than nineteen, carried entirely by one melody.
The second claim is the one the caveats give up on: that the window varies for two reasons, the written durations and the performed lengthening, and that the figure draws only their sum. Separating them takes one run rather than a curve. Set the rubato to zero and Frère Jacques’ window still runs from five notes to eleven, purely from its own written durations, and the adaptive reading scores 0.400 against the fixed 0.364.
So the accounting is: of the 0.350 gap between the principled fixed reading and the fitted one, the adaptive detector recovers 0.181 — and 0.036 of that, a fifth of it, owes nothing at all to the constructed performance. Four fifths does.
That is a less comfortable number than the section on circularity implies, and it is the right one to publish. The construction-free part of the result is real, it is the part the objection cannot touch, and it is small: a tenth of the way from the principled reading to the fitted one. The rest is the performance, placed at the boundaries by assumption, doing what it was built to do.
Which computation produced the numbers
The boundary-strength curve is melodyBoundaries, unchanged since where a phrase ends: a weighted sum of the three cues that rung’s model uses, evaluated between every pair of adjacent notes.
The performance is performanceBoundaries’ own construction — every phrase-final note lengthened by a stated fraction of itself, with everything else at its notated duration.
The local window is a count: starting at each note, how many notes fit inside the present, taken forward and then backward until the seconds run out. The smoothing width is that count over four, so a window of n notes is about two standard deviations either side, which is the same convention the fifth rung used to turn its note count into a σ.
The smoothing itself is a Gaussian whose width is read from an array rather than from a scalar, normalised pointwise so the varying width does not change the curve’s total. The peaks are found the same way boundaryHierarchy finds them, with the same minimum separation of three, so nothing about the peak-picking differs between the three readings.
Agreement is the F score against the notated phrasing at a tolerance of one note, which is the fourth rung’s measure. The fitted control sweeps σ from 0.3 to 8 in forty steps and reports the best it finds, so it is given every advantage.
Why a boundary narrows the window rather than widening it
There is a direction to check here, because the opposite would have been just as plausible in advance.
A phrase-final note is longer, so fewer notes fit inside the present, so the window is narrower in notes. A narrower smoothing keeps more detail, so the detector is more willing to report a peak — and a peak is what a boundary is. So the performance’s own gesture makes the detector more sensitive precisely where the gesture is.
The alternative story — that slowing down should make a listener take in more structure, not less — is about a window measured in bars rather than in seconds, and this ladder’s first rung is the argument against it: a phrase is a number of seconds, and a listener’s present does not stretch when the music slows.
So the sign of the effect is a consequence of the ladder’s own first result, which is a small piece of internal consistency worth having. A model in which the window were counted in bars would predict the opposite and would have to explain why a ritardando makes an ending harder to hear.
Where the model stops
The performance is constructed and the ladder has now said so four times. This collection has no performance timing data, and the whole apparatus above is a model of expressive timing rather than a measurement of one. The published range of phrase-final lengthening is real; its placement at exactly the notated boundaries is an assumption, and a real performer lengthens at some boundaries and not others.
The present is one number. Three and a half seconds is the middle of a range that runs from two to eight, and the fifth rung’s own figure is about that width being the model’s remaining uncertainty. The adaptive width inherits it: a listener with a two-second present gets a window that varies over a different range and a different reading.
And it is one tune. Two of the three melodies this collection carries show no gain from the adaptation at all — one of them is drawn above, and the other is the tune the whole ladder has been tested on — because their notated durations are nearly uniform and the constructed rubato is the only source of variation. The gain is real where the window varies and absent where it does not, which is at least the right dependence — but three tunes is three tunes.
What the picture cannot show
It cannot show a listener adapting. The model changes its resolution because the arithmetic of “how many notes fit in three and a half seconds” changes. Whether a listener’s segmentation window works that way — a fixed duration, filled with whatever arrives — is the fifth rung’s assumption, and it is an assumption about a mechanism nobody has measured directly.
Nor can it produce a hierarchy. The whole apparatus of persistence, which is the fourth rung’s contribution, needs a family of readings at increasing widths. A locally adaptive width gives one reading, so the ordering that rung produced is not available here — which is a real loss and is the reason this is a rung beside that one rather than after it.
The two sources of variation separate at one point and not along a curve. The window varies because the notated durations vary and because the performance lengthens; the figures draw their sum, and the section above splits it at zero rubato. What one point cannot give is how the split changes with the amount of lengthening, which would need the whole sweep recomputed against a written-duration-only baseline at each step.
Whose performances, and when
Phrase-final lengthening of twenty to fifty per cent is a measurement of Western art-music performance, mostly of piano playing and mostly of the nineteenth-century repertoire, and it is one of the most robust findings in the whole literature of expressive timing. It is also not universal: a great deal of music is played to a grid on purpose, and in that repertoire the adaptive detector and the fixed one are the same detector.
That is a useful boundary on the claim. The argument here is that a performance can hand a listener the scale at which to segment, and it applies exactly where performances are free to do so. In a tradition where the timing is fixed — by a click, by a dance, by an ensemble too large to bend — the information is simply not there, and a listener has to fall back on the fixed present the fifth rung computed. That is a repertoire boundary rather than a modelling one, and it is the same boundary the microtiming ladder draws for a different quantity.
Where this ladder goes next
Six rungs. A phrase is a number of seconds; the sentence and the period differ by a computable ratio; a boundary detector agrees with the page where the tune uses the cue it weights; the detector has a scale, which removes its threshold; the scale has a unit, which removes the parameter; and now the unit is a function of position, set by the performance.
What is owed after this is the hierarchy the adaptation costs, and a family of readings is the rung that pays it. A varying width gives one reading and no persistence, and persistence is what made the fourth rung’s boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than a family of fixed ones — a present of two seconds, of three and a half, of eight, each producing its own varying window — which is a family indexed by the one parameter the ladder has left and cannot remove, since the width of the psychological present is a fact about listeners rather than a choice.
Part 6 of 8
One essay in the series on phrase. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Perceptual presentPhraseRubatoScale-spaceSegmentationTempoTiming deviation
- The parameter that did not decide the answer perceptual present, phrase, tempo
- A form is sharp at the bottom and vague at the top perceptual present, phrase
- A proportion is only as fine as its two durations perceptual present, phrase
- A quantity resting against a wall tempo, timing deviation
- A silence long enough to be an ending perceptual present, tempo
- An ending is a deceleration perceptual present, tempo