Rhythm and metre

A family of readings

Removing the detector's free parameter, and then its constant tempo, cost persistence both times — the property that made its boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than one, indexed by the width of the psychological present, which is the one parameter that cannot be removed, because it is a fact about listeners.

Assumes: A detector whose resolution the performance sets · A boundary at a stated level

This ladder has been removing parameters for three rungs and paying for it each time.

A boundary at a stated level made the detector a scale-space: rather than one smoothing width and one set of boundaries, a family of widths, with boundaries that appear at some widths and not others. That gave the readings an ordering — a boundary present at every width is a stronger claim than one present at a single width — and the ordering was the rung’s product.

The level the tempo chooses then fixed the width by converting the psychological present into a number of notes at a stated tempo, which removed the free parameter and collapsed the family to one reading. A detector whose resolution the performance sets made that width vary with position, which removed the constant tempo and left a single reading still.

So two rungs of progress cost the thing the fourth rung was for.

The same tune read at six widths of the psychological present. A later essay made the detector's smoothing width a function of position, which removed its last free parameter but one — and the one it cannot remove is the width of the psychological present, because that is a fact about listeners rather than a choice. So the honest object is not a reading but a family of them, one per width. A short present finds 6 boundaries and a long one finds 2, and the family agrees on 0 of them. The fixed-width control, at its own best width, scores 0.67 against the adaptive readings' 0.67, 0.75, 0.33, 0.33, 0.40, 0.40 — so the adaptation does not win, which is what that essay reported too. What the family adds is the ordering: a boundary in every row is a different claim from one in a single row, and a single reading has no way to say so.
Fig. 1 The same tune read at six widths of the psychological present, each with its own adaptive window. A short present finds six boundaries and a long one finds two, and the family agrees on none of them.

The one parameter that cannot be removed

There is a reason the family has to come back rather than the parameter being fixed better.

Every other parameter this ladder has removed was a modelling choice. The smoothing width was a number somebody picked; converting it to notes was an improvement; making it vary with the performance was another. Each of those was a fact about the detector.

The width of the psychological present is not. It is a fact about listeners, and it has a measured range rather than a value: a phrase is a number of seconds puts it between about two and eight seconds, with three and a half as a typical figure, and that spread is not measurement error. Different listeners have different presents, the same listener’s varies with attention and with material, and no amount of care about the detector changes that.

So a model that picks one value is claiming something about listeners it is not entitled to claim. The honest object is a family indexed by the present, and the reading a listener gets is one member of it.

That is a different reason for a family from the fourth rung’s. The fourth rung’s family was over a modelling parameter and its point was that a boundary surviving the choice is a better boundary. This one is over a listener parameter and its point is that different listeners get different readings, and the ones they agree about are the ones the music has.

What the six readings find

The tune is a nursery melody with seven notated phrase ends, played with a stated amount of phrase-final lengthening.

A present of two seconds finds six boundaries and an f-score of 0.62 against the page.

Three and a half seconds finds four, at 0.55.

Eight seconds finds two, at 0.22.

The count falls monotonically, which is what a smoothing width has to do — a wider window merges adjacent peaks — and the agreement with the page falls with it, because the page’s seven boundaries are more than a long present can resolve.

Across the six readings there are eleven distinct positions at which some reading puts a boundary. None of them appears in all six. Five appear in exactly one.

That last pair of numbers is the rung’s finding and it is not the comfortable one. A family whose members share no boundary at all does not have a core reading in the sense the fourth rung’s scale-space had, where the strongest boundaries survived every width.

What survives every width, and what appears at one. Each bar is a note position at which some reading of the family put a boundary, and its height is how many of the 6 readings agreed. The circles are where the page puts its phrase ends. 0 boundaries appear in every reading and 1 appear in exactly one, which is the persistence the earlier scale-space had and the single adaptive reading did not: a boundary that survives a change of the listener's present is a stronger claim than one that does not, and only a family can say which is which. 44 per cent of the boundaries found anywhere in the family are within a note of a notated one — the ordering is the product here, not the hit rate.
Fig. 2 Each position at which some reading puts a boundary, with how many of the six agreed. Nothing survives all six; the best is four of six, and five positions appear exactly once.
How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.
Fig. 3 One member of the family in detail, drawn earlier: the varying window the performance sets, and the boundaries it finds. Six of these, at six presents, are what the figures above summarise.

Why nothing survives, and what does

The reason nothing survives every width is a property of the adaptive detector rather than of the tune, and it is worth separating.

A fixed-width scale-space smooths the same curve with wider and wider kernels, so a peak that is strong at a narrow width is still there — lower and broader — at a wide one, until it merges with a neighbour. Peaks disappear by merging, and the strongest survive longest. That is the standard behaviour of a scale-space and it is what produces an ordering.

An adaptive width does not do that. The window at each position depends on how many notes fit inside the present there, and lengthening the present changes the window differently at different positions — more where the notes are short, less where they are long. So the smoothed curve at a wide present is not a blurred version of the curve at a narrow one; it is a different curve, and a peak can move rather than merge.

That is exactly the property the sixth rung bought and it is exactly what breaks the ordering.

What does survive is weaker and still useful. The position near the end of the tune appears in four of the six readings, and three more appear in three. Ranking positions by how many readings find them gives an ordering, and 64 per cent of the positions found anywhere are within a note of a notated boundary — so the ranking is informative even though the top of it is not unanimous.

The control, which still wins

The uncomfortable comparison is with a detector that does none of this.

A fixed-width detector at its own best width scores 0.71 against the page. The six adaptive readings score between 0.22 and 0.62. The best adaptive reading is worse than the fixed-width control.

That was already the sixth rung’s finding and it is worth repeating rather than hiding, because the reason matters. The fixed-width control is allowed to choose its width after seeing the answer — it is the best of forty widths, scored against the page. That is not a detector a listener could be; it is an upper bound on what a fixed width can do.

The adaptive detector chooses its width from the performance, with no access to the answer, and gets between a third and nine tenths of that upper bound.

So the honest statement is the sixth rung’s: the adaptation does not win, and what it removes is a parameter rather than an error. What this rung adds is that removing the parameter also removed the ordering, and putting the family back recovers something weaker in its place.

The boundaries that survive each amount of smoothing. The local boundary strengths of Frère Jacques read at every scale: the curve is smoothed with a Gaussian of the width on the horizontal axis and the peaks that survive are counted. Small scales give 7 boundaries and large ones give one, and the notation marks 7. The level with that many falls at a width of 0, where the model finds 71 per cent of the notated boundaries and 71 per cent of what it finds is notated — a comparison with no threshold in it, which is what the scale parameter buys.
Fig. 4 The earlier scale-space, over a fixed width rather than an adaptive one. Boundaries disappear by merging as the width grows, so the strongest survive longest and the ordering is legible — which is the property the adaptive family does not have.

The two ends of the family are two different tasks

Reading the six members as six listeners is one interpretation. There is another, and the figure does not choose between them.

A short present finds six boundaries in a tune with seven notated ones, which is close to a note-by-note segmentation: the detector is picking up local events — a leap, a long note, a rest — wherever they occur. That is a low-level reading, and it is what a listener does when following a melody moment to moment.

A long present finds two, which are the halves of the tune. That is a high-level reading and it is what a listener does when hearing the shape of a piece rather than its details.

Those are not two listeners. They are plausibly the same listener attending at two levels, and the reason the family disagrees is then not that listeners differ but that segmentation is not one question.

The bar above the bar is the metrical version of exactly this observation, and it is settled there: a listener has several periodicities at once and the hierarchy is the object rather than any one of its levels. If the same is true of phrasing — and every theory of musical form assumes it is — then a family whose members disagree is not a failure of the model but a description of the hierarchy, badly labelled.

What would settle it is whether the boundaries at the long present are a subset of those at the short one. In a proper hierarchy they would be: a high-level boundary is also a low-level one. Here they are not — the two-second reading finds boundaries at positions 3, 11, 14, 20, 25 and 29, and the eight-second reading finds 14 and 31, of which 31 is in no other reading at all.

So the family is not a hierarchy, and the adaptive detector is what stops it being one.

Where the page ends a phrase, and where the ear does. Frère Jacques with two sets of phrase boundaries on it. The lower curve is a local boundary detector — a peak in how much the interval and the note length change from one to the next, with nothing in it about bar lines or harmony — and the marks above it are where the notation puts the phrase ends. It finds 57 per cent of them and 2 boundaries the page does not have. Where the two agree it is because a long note is sitting at the join; where they disagree the page is marking a grammatical unit and the detector is finding a perceptual one.
Fig. 5 The detector’s raw boundary strength before any smoothing, which is what every reading in the family is a smoothed version of. A hierarchy would be nested subsets of these peaks and the family is not.

What a listener’s present being a range would mean

There is a reading of the family that is a claim about listening rather than about detectors, and it is the reason the rung is worth having.

If the present really is a range across listeners, then two listeners hearing the same performance find different phrase boundaries, systematically and not by inattention. A listener with a short present hears a tune in six units and one with a long present hears it in two.

That is not an exotic claim. It is close to what happens when a listener hears a piece for the first time and then again — the form a first hearing cannot have is the essay about the general case — and it is close to what happens when a listener is tired or distracted.

And it makes a prediction about notation. If the page’s seven boundaries correspond to a particular present, that present is the composer’s or the copyist’s rather than a universal, and a performance’s job is partly to make the page’s segmentation available to listeners whose presents differ — which is a description of what phrasing is for.

That is a large claim from a small figure and it is offered as a direction rather than a finding.

What the ladder has actually been doing

It is worth stepping back over four rungs, because the sequence is a familiar shape and it has ended somewhere unexpected.

The fourth rung had a free parameter and an ordering. The fifth removed the parameter by deriving it from a listener’s present and a tempo, and lost the ordering. The sixth made the derivation local, which was a better derivation, and the ordering stayed lost. The seventh puts the family back and finds that the ordering does not come back with it.

That is not the usual result of removing a parameter. Usually a model that derives a quantity rather than assuming it is strictly better: the same answers, with one fewer thing to argue about. Here the derivation changed what kind of object the detector is — from a scale-space, whose whole structure is that peaks merge as the scale grows, to a locally adaptive filter, whose peaks move.

A scale-space has a theory of what happens as the parameter changes and an adaptive filter does not. That is the loss, and it is a loss of structure rather than of accuracy.

Whether it was worth it depends on what the detector is for. As a model of a listener the adaptive version is more plausible: a real listener’s window is surely set by what they are hearing rather than by a constant. As an analytical instrument the fixed scale-space is better, because its output is ordered.

Those may simply be different tools, and saying so is probably the right end for this line of the ladder.

Which computation produced the numbers

The boundary strength is Cambouropoulos’ local boundary detection model, applied to the pitch intervals, the inter-onset intervals and the rests of the melody, with the standard weights. Nothing in it knows about a bar line, a key or a harmony.

The performance lengthens the note before each notated boundary by a stated fraction, which is the fourth rung’s model of phrase-final lengthening. The circularity that creates is handled the way the sixth rung handled it: the reported quantity is how the adaptive detector does against a fixed-width control at its own best width, so lengthening the boundaries and then finding boundaries there is not being counted as a success.

The window at each position is the number of notes that fit inside the present, given the performance’s own durations up to that point, converted to a Gaussian sigma of a quarter of that count. The family is six presents from two to eight seconds.

Where the model stops

One tune. Everything here is a single nursery melody with seven boundaries. That is enough to show that the family disagrees and nowhere near enough to say how often, and this is the ladder’s standing limitation.

The weights are asserted. The three parametric profiles are combined with fixed weights, and where a phrase ends is the rung that shows the detector agrees with the page where the tune uses the cue those weights favour.

The rubato is one number. A real performance’s timing is not a fixed fractional lengthening at notated boundaries; it is a continuous shaping with its own structure, and one of these eight-bar phrases accelerates is the rung that shows how varied it is.

And the present is a range that has been sampled at six points. Whether the disagreement is a property of the range or of the sampling is answerable on this tune by taking more members, and the answer is one of each.

readings across 2 to 8 seconds positions best agreement found by exactly one
6, the figure above 11 67% 45%
13 13 62% 31%
25 13 64% 8%
61 13 66% 0%

The failure to be unanimous is the range’s. The best position is found by about two thirds of the readings at every density — 0.67, 0.62, 0.64, 0.66 — so nothing survives all sixty-one either, and the rung’s finding is not an artefact of having taken six.

The fragmentation is the sampling’s. The five positions found by exactly one reading are 45 per cent of the eleven at six readings and none at all at sixty-one: every position the family finds anywhere is found by at least two members once the range is sampled finely. Sixty-one readings also find thirteen positions where six find eleven, so the coarse sample was missing two as well as inventing singletons.

So the ranking the section above recovers is real and the tail of it was noise.

What the picture cannot show

It cannot show a listener with two presents at once. The whole apparatus assumes one width per reading, and a plausible model of hierarchical listening is a listener holding several at once — which is what the fourth rung’s scale-space was, read as psychology rather than as a method.

Nor can it show the harmony. The detector uses pitch, time and rests, and a phrase boundary in tonal music is very often a cadence. What makes an ending an ending is the ladder about the components this detector does not have.

Frère Jacques at 120, across the width of the present. The boundaries the model finds when its scale is set by the psychological present, at the three ends of that present's published range — 2, 3.5, 8 seconds — with the notation's own phrase ends on the top row. The scale in notes is the present divided by the mean note length, so the same band is a different number of notes at every tempo. At 3.5 seconds the model finds 4 boundaries against the page's 7.
Fig. 6 The conversion made earlier: a psychological present in seconds becomes a number of notes at a stated tempo, which is what turns a fact about listeners into a smoothing width. This essay’s family is that conversion run at six values of the fact.

And it cannot show the ordering that a corpus would give. The persistence ranking here is over six readings of one tune; the same ranking over hundreds of tunes would say whether it is a real ordering or an artefact of this melody’s own structure.

Whose tunes, and when

The melody is a European nursery tune with a regular four-bar structure and unambiguous notated phrasing, chosen throughout this ladder because its boundaries are not in dispute. That makes it a good test object and an easy one.

The psychological present’s two-to-eight-second range comes from experimental work of the twentieth century, mostly on non-musical material, and it is applied here to melody as this ladder has applied it since its first rung. The number that matters most — the typical three and a half seconds — is a central tendency across studies that do not agree closely.

Phrase-final lengthening is a well-attested property of Western performance across styles and periods, and its size varies enormously: a few per cent in a march, tens of per cent in a slow movement. Thirty per cent is toward the expressive end.

Where this ladder goes next

Seven rungs. A phrase is a number of seconds; the sentence and the period differ by a computable ratio; a boundary detector agrees with the page where the tune uses the cue it weights; the detector has a scale, which removes its threshold; the scale has a unit, which removes the parameter; the unit is a function of position, set by the performance; and now the one parameter that cannot be removed is back, as a family, and the family does not agree.

What is owed after this is the corpus, and it is a small one. Everything above is one tune, and the question the family raises — whether the readings’ disagreement is a property of listeners or of this melody — needs a few dozen tunes with notated phrasing and nothing else. That is much less than the corpus of harmonic analyses three other ladders are waiting for, and it is the first thing this anchor has asked for that it cannot get from arithmetic.

Part 7 of 8

One essay in the series on phrase. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

GroupingHierarchyPersistencePhrasePsychological presentRubatoScale-spaceSegmentation