A boundary at a stated level
Assumes: Where a phrase ends · A phrase is a number of seconds
The third rung of this ladder ran Cambouropoulos’ boundary detector over the three tunes this collection carries and found that whether it agrees with the notated phrasing depends entirely on which cue each tune uses. Twinkle’s phrases all end on a long note, so a duration-weighted detector finds five of five with no false alarms; Ode to Joy’s run on in crotchets and its phrasing is in the intervals, where a duration detector finds one of three.
It ended by naming two debts. A version with a scale parameter — a boundary model that reports a hierarchy rather than a list, so a comparison with the notation is a comparison at a stated level. And a version run on performance timings rather than on notated ones, which is where the missing information in the hardest case was guessed to be.
Both are cheap to build and they do different jobs.
The threshold was a free parameter and now it is not
The third rung’s model has a threshold in it: a local maximum in the boundary-strength curve counts as a boundary if it clears the mean. That is a decision, it is not in Cambouropoulos’ model, and it was made here in order to turn a curve into a list.
A scale removes it. Smooth the strength curve with a Gaussian of width σ and read off the peaks. At σ = 0 every wobble is a boundary; as σ grows, small peaks are absorbed into larger neighbours and the count falls monotonically to one. There is no threshold anywhere: a peak either survives the smoothing or it does not.
What that buys is a comparison the third rung could not make. The notation marks some number of phrase ends — five for Twinkle, three for Ode to Joy, seven for Frère Jacques — so the model can be read at the level that produces that many boundaries and asked whether they are the right ones. Nothing about the answer depends on a threshold, because the number is supplied by the notation rather than chosen.
For Twinkle the answer at that level is perfect: five boundaries, all five notated, nothing spurious. That is not new — the third rung got the same result with a threshold — but it is now got without one.
The hierarchy is real and is not much use here
Once there is a scale there is also a persistence: the largest σ at which a boundary is still a peak. Order the boundaries by it and the result is a hierarchy — the boundaries that survive most smoothing first, which is the ordering a listener’s sense of a major and a minor join ought to correspond to.
Asked whether that ordering carries information the flat curve does not, the answer is mixed and mostly negative.
For Twinkle the top five by persistence get three of the notated five; the top five by raw strength get all five. For Ode to Joy both get one of three. For Frère Jacques persistence gets six of seven and raw strength three of seven, which is the one clear win.
One thing that pattern says about the three tunes is worth noting, because it names when to reach for a hierarchy. Persistence beats raw strength on exactly the tune whose strength curve has the most competing peaks, and loses on the one whose peaks are already few and correct. So a hierarchy is a sorting device rather than a detecting one: it does not find anything the flat curve did not, it decides which of the things the flat curve found are the important ones, and that is only worth doing when there are more of them than the answer needs.
So the hierarchy is a genuine structure and it is not a better detector. On a tune whose phrases are marked by long notes, the raw strengths are already correctly ordered and smoothing only blurs them; on a tune with a varied rhythm and many competing peaks, persistence sorts the important ones from the incidental ones and does considerably better.
That is a real finding about when a hierarchy helps, and it is smaller than the debt implied. What the scale parameter was owed for was the comparison, and that it does deliver.
Two per cent of rubato
The second debt is the interesting one, and the third rung guessed right about where the information is.
This collection has no performance data, so the performance has to be constructed. That is legitimate if the construction is the thing every measurement of expressive timing agrees on, and there is exactly one such thing: a phrase-final note is lengthened. Todd’s and Repp’s measurements and everything since put the amount at roughly twenty to fifty per cent for an ordinary phrase ending.
Lengthening the notated boundaries and then finding boundaries there would be circular. The answer offered to the circularity is that the amount is the result: if the model needs more rubato than any performer applies, then performance timing is not where the information is. That answer turns out not to work, for a reason the last section of this rung demonstrates rather than argues — but the test is worth running before the objection is stated, because the number it produces is the evidence for the objection.
Two per cent. The tune the third rung found nearly impossible reaches full recall at a rubato twenty-five times smaller than the smallest a performer would apply — though exact is the wrong word for it, as the section below shows: the detector finds all three notated boundaries and six others at the same time, and it would find any three notes that had been lengthened instead.
The reason is worth stating because it is a property of the model rather than of the music, and it cuts both ways. Cambouropoulos’ detector measures change: its inter-onset profile scores each interval by how much it differs from its neighbours. Ode to Joy is a tune in almost uniform crotchets, so its duration profile is nearly flat and contributes nothing — and any perturbation whatever becomes the only signal there is.
A detector of change on a rhythm with no change in it is infinitely sensitive to the first change introduced. That is why two per cent is enough, and it is also why the number should not be read as a measurement of anything about the tune. It is the same sensitivity a metre induction shows to a cue nobody supplied: on a tie, any evidence decides.
That argument can be demonstrated rather than asserted, and the demonstration is the answer to the circularity worry above. Lengthen three notes chosen at random rather than the three notated phrase ends, and score the detector against those:
| lengthening | recall on the notated ends | recall on three random notes |
|---|---|---|
| 2% | 1.00 | 0.74 |
| 5% | 1.00 | 0.97 |
| 10% | 1.00 | 0.97 |
By five per cent the detector finds arbitrary lengthened notes as reliably as it finds phrase endings. So the two per cent is not a fact about where the phrasing is; it is the amount of perturbation this detector needs to notice a perturbation, and it would be the same number for a lengthening applied anywhere. The circularity the section above tried to escape by measuring the amount is not escaped: the amount is a property of the detector’s sensitivity and not of the tune’s phrasing.
There is a second number the recall column hides and it should be printed beside it. At every lengthening from two per cent upward the detector finds nine boundaries where the notation marks three — recall 1.00 and precision 0.33. So it gives all three is true and it finds the phrasing is not: it finds the phrasing and six other places, and the rubato bought recall while leaving precision exactly where it was at zero rubato.
Which changes what the second debt was paid with. The scale parameter delivered a comparison without a threshold and did not rescue the hard case; the rubato delivers recall on the hard case and does not deliver precision, and what it demonstrates is a property of the detector rather than of performance. Neither debt bought what it was owed for, and the honest result of paying both is a clearer account of what this detector is: an instrument that reports where the largest local change is, whatever put it there.
The tune that stays hard
The third tune goes the other way and it is the honest limit of the whole exercise.
Four of seven at no rubato, five at thirty per cent, and no further inside a doubling of the note. The tune the notation-based model already did best on is the one that performance timing cannot finish, which is the opposite of what the third rung’s guess would predict if the guess were about tunes rather than about models.
The pattern across the three is clean and is not the pattern anybody was looking for:
| tune | rhythm | boundaries found on the page | rubato needed for all |
|---|---|---|---|
| Twinkle | phrases end long | 5 of 5 | none |
| Ode to Joy | uniform crotchets | 1 of 3 | 2% |
| Frère Jacques | varied | 4 of 7 | more than 100% |
The rubato a tune needs is inversely related to how much rhythmic variety it already has. A tune with none is solved by a whisper of it; a tune with a lot cannot be solved by any amount, because the lengthening is one change among many and the detector has no way to prefer it.
What the level is, if it is not chosen
There is a third thing the scale makes possible and it is the one that would make this a model rather than a measurement.
At present the level is chosen by counting the notation’s marks, which is a comparison and not a prediction: the notation is being used to select the level and then to score it. That is a fair test of whether the right boundaries are in the hierarchy and no test at all of whether a listener would read them.
The level a listener reads is not free either, and this ladder already measured what fixes it. A phrase is a number of seconds — two to eight, the psychological present — and a tempo turns seconds into notes. So a smoothing width of two notes at 120 beats a minute is a second, and the level the arithmetic would select is whichever σ puts the mean gap between surviving boundaries inside the window.
That closes the loop and this rung does not close it, because the conversion has one more decision in it than looks likely — a σ in notes is a σ in seconds only if the notes are equal, and the tunes whose notes are not equal are exactly the ones the scale mattered for. It is named here as the next rung rather than attempted.
Which computation produced the numbers
The boundary strengths are lbdm, unchanged: each of the interval, inter-onset and rest profiles gets a degree of change between consecutive values, an interval’s strength is its own size times the change either side of it, and the three are combined with the published weights.
The scale is a Gaussian convolution of the resulting strength curve, with peaks read as local maxima and a minimum separation of three notes. That separation is not decoration: a smoothed curve has plateaux, and a bare local-maximum test walks along one reporting every point on it — which the first version of this did, producing “boundaries” at notes 14, 15 and 16 and a persistence ordering that was an artefact of the plateau. Three is the same gap this site’s own peak-finder uses elsewhere.
The persistence of a boundary is the largest σ at which it is still a peak, and the hierarchy is that ordering. Selecting the top n by persistence enforces the same three-note separation, for the same reason.
The performance lengthens the inter-onset interval of each notated phrase-final note by a stated fraction and leaves everything else alone. The tunes and their notated phrasings are this site’s own three, entered as three melodies and three lists of phrase ends, and both are as small a corpus as a corpus can be.
Every number here is over three tunes. The third rung said so about its own numbers and it is more true of these, because this rung reports a relation between three tunes rather than a property of each — and a relation over three points is a shape rather than a finding.
Whose music, and when
The three tunes are European, tonal, metrical and short, and two of them are children’s songs — the same three this collection has used since it began encoding melodies at all. The claim about final lengthening is about performance practice in the same tradition, measured mostly on Western art music by pianists in laboratories.
The boundary model is not about a repertoire and is worse for it. It has no bar line, no key, no harmony and no memory, so what it can find is a local change in a stream — which is a fair model of the first thing an ear does with an unfamiliar tune and a poor one of what a listener who knows the piece does. The phrase’s own duration bound is a fact about a listener that nothing in this detector contains: the model would happily report a boundary every two notes or once in eighty, and the window that says a phrase is a few seconds long is not in it at any scale. The sentence and the period differ by one ratio computed from the lengths of their parts, and that ratio is a hierarchy statement the detector could in principle be asked for and has not been.
That is the honest reason the scale parameter matters more than the rubato result. A scale is where a duration bound could be put in. A σ of two notes at a tempo of 120 is about a second, and the psychological present is two to eight — so the level a listener reads is a level, in principle, that the tempo fixes rather than the analyst.
What the picture cannot show
The performance is constructed and its shape is one number. A real ritardando is a curve over several notes — a square root rather than a step — and lengthening one note is the crudest possible version of it. A distributed slowing would spread the boundary evidence over the approach and might be found more easily or less.
The rubato is applied at the notated boundaries and nowhere else. A performer also lengthens notes that are not phrase ends, and the model has no way to tell the two apart; adding that noise would raise every number’s requirement.
Persistence over three tunes is a shape rather than a result, and the one clear win for the hierarchy is a single tune.
And the scale is in notes rather than in seconds. Smoothing by two notes is a different amount of time at every tempo, which is exactly the collision the microtiming ladder is about and the same one a fixed physical delay makes unwritable, and converting the scale to seconds is the obvious next step and is not taken here.
Where this ladder goes next
Four rungs. A phrase is a number of seconds and the bar count is whatever the tempo makes it; the sentence and the period differ by one ratio computed from the lengths of their parts; a boundary detector agrees with the page only where the tune uses the cue the detector weights; and now the detector has a scale, which removes its threshold, and a constructed performance, which finishes the case the page could not.
The rung after it is the one the last paragraph names and this rung deliberately did not take. The scale is in notes and the ladder’s first rung is in seconds. Convert one to the other and the level a listener reads is no longer a free parameter at all: the psychological present is two to eight seconds, a tempo turns that into a number of notes, and the hierarchy is then to be read at whichever level the tempo selects. That is a prediction with teeth — the same tune at two tempos should be heard as differently phrased, at levels the arithmetic can name in advance — and it is a computation, because both halves are already on this site.
Part 4 of 8
One essay in the series on phrase. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Final lengtheningLbdmNotationPhraseRubatoScale-spaceSegmentation
- A piece is mostly itself again notation, segmentation
- A proportion is only as fine as its two durations notation, phrase