Where a phrase ends
Assumes: A phrase is a number of seconds · The boundary is where the neighbourhood changes
This ladder has established that a phrase is a number of seconds rather than a number of bars, and that an eight-bar phrase does not run at a constant tempo. Both take the phrase’s boundaries as given: the page says where they are and the measurements are made between them.
The page saying where they are is itself a claim, and it is the same kind of claim a time signature makes — an assertion about how the music divides that the sound may or may not support. There is a standard model for finding boundaries in a melody with nothing but the melody, it is fifty years old in one form or another, and running it against the notation is a two-line experiment.
The model
The detector is Cambouropoulos’ local boundary detection model, and it has no music theory in it at all.
For each parameter of the melody — the size of each interval, the length of each note, the length of each rest — compute how much that parameter changes from one event to the next, as a proportion: |a − b| / (a + b). An event’s boundary strength is its own size multiplied by the change on either side of it. Combine the three profiles with weights and take the local peaks.
That is the whole model. It knows nothing about bar lines, keys, harmony, cadences or repetition. It says: a boundary is where something changes.
Two things about that are worth noticing before any results. The first is that it is a local rule — the same shape as the boundary that is where the neighbourhood changes — each event is compared with its two neighbours and nothing else — so it cannot in principle find a boundary that is only a boundary because of something four bars earlier. The second is that the multiplication by the event’s own size is what makes it a boundary detector rather than a change detector: a large interval preceded and followed by small ones scores highly, and a small interval in a field of small intervals does not, however proportionally different it is.
Five of five is a good result and it is worth being suspicious of, because Twinkle is the easiest possible case: its phrases are all the same length, they all end with a long note, and the long notes occur nowhere else.
The tune it fails on
One of three, and it is the one with the long note.
That is not a subtle failure and it locates the cause exactly. Twenty-eight of Ode to Joy’s thirty notes are one beat long. A detector that puts half its weight on duration has almost nothing to work with, and its answer comes from the interval profile alone with a quarter of the weight.
Give it the interval profile with all of the weight and the result reverses.
And the reverse is true of Twinkle.
So the two tunes need opposite detectors, and each does badly on the setting that suits the other. The published weighting — a quarter on pitch, a half on duration, a quarter on rests — is a compromise that is worse on each tune than the single cue that tune uses: it gets 33 per cent recall on Ode to Joy where pitch alone gets 100, and 71 per cent precision on Twinkle where duration alone gets 100.
The table, so the pattern is visible at once
| notated | default weights | pitch only | duration only | |
|---|---|---|---|---|
| Ode to Joy | 3 | 1 found, 7 marked | 3 found, 11 marked | 1 found, 3 marked |
| Twinkle | 5 | 5 found, 7 marked | 2 found, 10 marked | 5 found, 5 marked |
| Frère Jacques | 7 | 4 found, 6 marked | 5 found, 10 marked | 3 found, 5 marked |
Read down the columns. The default weighting is never the best on any tune. Pitch alone has the best recall on two of the three and the worst precision on all three. Duration alone is perfect on one tune and useless on another.
A single detector cannot be right about all three, and the reason is not that the model is too simple. It is the same division of labour an ending that can be heard coming runs into. It is that the tunes carry their phrasing in different places, and a fixed weighting is a bet about which place.
The compromise is not the compromise it looks like
The published weighting is a quarter on pitch, a half on duration and a quarter on rests, and the third of those is spent on nothing here. All three tunes are stored as continuous runs of notes, so the rest profile is identically zero before it is weighted.
The natural worry is that this handicaps the default setting — that a quarter of its attention is being paid to a silent channel, and that the comparison above is therefore unfair to it. It is not, and the reason is worth following, because it is a property of the peak-picking rule rather than of the weights.
Each profile is normalised to its own maximum and the three are then summed with the weights. A profile that is everywhere zero stays zero, so the combined curve is the pitch and duration terms alone, scaled to a maximum of three quarters rather than one. The threshold is the mean of that same combined curve, and the peak test asks whether a value exceeds the threshold and both its neighbours. Every one of those comparisons is between two numbers carrying the same factor, so the factor cancels: multiplying the whole curve by three quarters cannot move a single peak.
So on these tunes the published weighting is exactly the weighting [1/3, 2/3, 0] — one part pitch to two parts duration — and running the model both ways returns identical boundary lists on all three. The dead channel is not a handicap; it is not there at all. The compromise fails on Ode to Joy for the reason already given and for no additional one: two thirds of its weight is on a profile with almost no variation in it, and the remaining third has to carry the answer alone.
That is a small result and it matters for what can be concluded. Had the rest channel diluted the default setting, the finding would have been about an artefact of running a three-profile model on rest-free tunes, and the fix would have been to renormalise. It does not, so the finding is about the cues.
What chance would have scored
None of the numbers so far has a floor under it. A detector marking eleven of a possible twenty-nine positions, scored with a tolerance of one note either side, is claiming a great deal of the tune, and “it found all three” needs to be read against how often that happens by accident.
The null model is the same detector stripped of its content: mark the same number of positions, chosen uniformly at random, and score them the same way. For a boundary away from the ends, the chance of missing it is the chance that none of the k marks lands on it or on either neighbour, which for n positions is the fraction of k-subsets avoiding those three — so the expected recall is
1 − (n−k)(n−k−1)(n−k−2) / n(n−1)(n−2).
| marks | model recall | chance recall | |
|---|---|---|---|
| Ode to Joy, default | 7 of 29 | 33% | 58% |
| Ode to Joy, pitch only | 11 of 29 | 100% | 78% |
| Ode to Joy, duration only | 3 of 29 | 33% | 29% |
| Twinkle, default | 7 of 41 | 100% | 44% |
| Twinkle, pitch only | 10 of 41 | 40% | 58% |
| Twinkle, duration only | 5 of 41 | 100% | 33% |
| Frère Jacques, default | 6 of 31 | 57% | 49% |
| Frère Jacques, pitch only | 10 of 31 | 71% | 70% |
| Frère Jacques, duration only | 5 of 31 | 43% | 42% |
Read the last two columns together and most of this essay’s results disappear.
Two of the nine settings are worse than chance, and both are the ones already identified as mismatched: the default weighting on Ode to Joy, and pitch alone on Twinkle. That is a stronger statement than “it does badly” — a detector that marks seven positions and finds one of three notated boundaries has done worse than marking seven positions with no information at all.
All three Frère Jacques settings sit within a point or two of chance. The tune that looked like the interesting mixed case, where no setting was much better than any other, turns out to be a tune where no setting is much better than nothing. The earlier reading — that its phrases end sometimes on a long note and sometimes on a leap, so no single cue serves — was too generous. On this scoring the model is not choosing badly between cues on Frère Jacques; it is not detecting anything.
Three results survive, and they are the ones the essay should have been about. Duration alone on Twinkle is 100% against a chance of 33, and it is the only setting anywhere that marks exactly as many boundaries as the notation has and gets every one of them. The default weighting on Twinkle is 100 against 44, which is the same result with two spurious marks attached. Pitch alone on Ode to Joy is 100 against 78 — real, but a much smaller margin than the bare recall suggests, because eleven marks with a tolerance of one cover thirty-three of the twenty-nine available slots’ worth of coverage before the overlaps are taken out.
The general lesson is about the tolerance rather than the model. A window of one note either side triples the target, and on a thirty-note tune a detector that marks a third of the positions is nearly guaranteed to hit any given boundary. Recall without a mark count beside it is not a measurement, which is why every figure above states how many boundaries the model drew as well as how many it got right.
What the disagreement is about
There are two readings of these numbers and they lead to different conclusions.
The first is that the detector is bad. That would be too quick. A model with three parameters and no theory is not expected to reproduce a notated phrasing, and on the tune whose phrasing is carried by the cue it weights most, it is exactly right with no errors — which is more than a model this simple has any business being.
The second is that the page and the detector are answering different questions, and the evidence for that is in the false alarms.
The extra boundaries the model finds are not random. In Ode to Joy under interval weighting it marks after notes 3, 5, 11, 13, 18, 20 and 25 as well as the three notated ones, and those are half-bar divisions — the sub-phrase groupings that a performer would breathe over but not mark. Precision measured against a notated phrasing counts them as errors, and a listener would not.
So the page marks one level of a hierarchy and the detector finds several. Notation has one symbol for a phrase and no symbol at all for the level below it, so a comparison that treats every unmarked peak as a false alarm is measuring the notation’s resolution rather than the model’s accuracy.
That is the missing parameter and it is the same one. The novelty scan takes a width; the boundary model does not, and its implicit width is set by the weighting. Reading the two figures together, the disagreement between the page and the detector is mostly a disagreement about scale, and the page’s scale is a convention rather than a measurement.
And the constraint that decides which level the page marks is not in the page or in the model. A phrase is two to eight seconds long because that is how much can be held as one present event, so the marked level is the one that fits the window and the levels below it are sub-groupings. Twinkle’s phrases at a hundred and twenty are three and a half seconds and Ode to Joy’s are four; both tunes’ half-phrases are near the window’s bottom.
Which gives a way to score the model that does not depend on the notation at all: prefer the boundaries whose spacing falls inside the window. On Ode to Joy under pitch weighting the eleven marked boundaries divide the tune into stretches of two to five notes, which is one to two and a half seconds — below the window. Keeping only the strongest peaks until the mean spacing is four seconds recovers the notated three. That is a post-hoc filter and it is offered as a suggestion rather than a result, because it was chosen after seeing the answer.
What the notation is actually marking
Which leaves the question of what a slur or a phrase mark is for, if not this.
Three things, and only the first is what the detector looks for.
Where the music groups. Some of the time the page is recording the same thing the detector finds, and where it is, the two agree.
Where the performer should breathe or bow. That is an instruction about execution, and it is constrained by physiology and by the instrument as much as by the music — a string player’s phrase mark is a bowing and a singer’s is a breath, and the two are not always in the same place in the same piece.
Where the grammar divides. An antecedent and a consequent are a syntactic pair and the division between them is at the half cadence, which is a harmonic event. Nothing in a melodic boundary detector can see a half cadence.
An ending is a bundle of independent cues — harmonic, metrical, durational, registral — and a phrase mark on the page sits where enough of them coincide. A melodic boundary detector reads one of them, and nothing in it can see a half cadence, which is the harmonic event an antecedent and consequent divide at.
That last point is the one to take away. The tune whose notated durations carry no phrasing has a performance whose durations do, because that is what a performer is for. The score’s flat run of crotchets is an instruction to be shaped, and the shaping is where the boundary information is.
And the crudest cue of all is one none of these tunes uses: a rest long enough to be heard as a gap rather than as an articulation. That threshold is a fixed number of milliseconds and the tunes here are continuous, so nothing in them is available to it.
Which computation produced the numbers
The model is implemented as published: degree of change r = |a − b| / (a + b) between consecutive values of each profile, strength s = value × (r before + r after), each profile normalised to its own maximum, then combined with the stated weights. Peaks are local maxima above the mean of the combined strength.
A boundary is counted as found if the model marks within one note of the notated one, and the tolerance is stated in every figure rather than being chosen per tune.
Tightening it to zero is worth doing, because it separates a detector that is right from one that is nearly right. It changes nothing on two of the three tunes: Ode to Joy and Twinkle score identically at both tolerances under all three weightings, so every boundary the model finds on them it finds exactly, and the window is not propping up the headline results. The entire effect of the tolerance is on Frère Jacques, where the default weighting falls from four of seven to two, pitch alone from five to two, and duration alone from three to two — which is another way of saying what the chance baseline said about that tune. A detector whose hits are all off by a note is not locating boundaries; it is marking near them.
Combining the two figures into an F-measure ranks the settings without hiding either, and the ranking is the one the separate columns imply: the default weighting is never the best on any tune — 20 against pitch alone’s 43 on Ode to Joy, 83 against duration alone’s 100 on Twinkle, and 62 against pitch alone’s 71 on Frère Jacques. The claim that the compromise is worse on each tune than the single cue that tune uses holds under a combined measure as well as under the two separate ones.
Recall is the fraction of notated boundaries found. Precision is the fraction of found boundaries that are notated. Both are reported and neither is optimised: no threshold, weighting or tolerance was tuned to improve a number.
The notated boundaries are the phrase ends of each tune’s own structure — Twinkle’s six seven-note phrases, Ode to Joy’s four bars of eight, Frère Jacques’ eight short phrases — entered once and used for every weighting.
The intervals are absolute semitone distances and the durations are the notated beats. There are no rests in any of the three tunes, so the third profile is identically zero and contributes nothing under any weighting; the “duration only” runs are therefore the same as “duration and rests”.
What the picture cannot show
Three tunes, and they are the wrong three. All three are strophic melodies with regular phrase lengths and no rests. The interesting cases for a boundary model are the ones with irregular phrases, and this collection has none.
The notated phrasing is my reading, not an editorial marking from a source. For these three tunes the phrase structure is not in dispute — they are four-square — but it is an input rather than a measurement, and a different reading would change every number.
The model has no memory. It compares each event with its two neighbours and nothing else, so it cannot use the fact that a phrase has happened before. Repetition is the strongest segmentation cue there is, and in a strophic tune it settles the question by itself; leaving it out is what makes the task hard enough to be interesting and also what makes the model unlike a listener.
The weights are the only thing varied. The threshold, the peak-picking rule and the normalisation are all held at their published forms, and each of them would change the numbers as much as the weights do. Nothing here searches over them, and a model with four free choices evaluated on three tunes could be made to say almost anything.
And there is no harmony. A half cadence is a boundary, it is where the page most reliably puts a phrase mark, and nothing in a melodic profile can detect one. The one case here where the notation and the detector disagree most — Ode to Joy — is the tune whose phrase structure is most clearly harmonic.
The ladder from here
This rung asked whether the page’s phrase marks are where the sound’s boundaries are and found that it depends which cue the tune uses, with no fixed weighting serving two tunes. What the ladder owes is the version with a scale parameter — a boundary model that reports a hierarchy rather than a list, so that a comparison with the notation is a comparison at a stated level — and the version run on performance timings rather than on notated ones, which is where the missing information in the hardest case turns out to be.
Part 3 of 8
One essay in the series on phrase. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 11.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Boundary conditionGroupingInter-onset intervalMelodic intervalNotationNoveltyPhraseSegmentation
- A note lasts until the next one starts inter-onset interval, notation
- A piece is mostly itself again notation, segmentation
- A proportion is only as fine as its two durations notation, phrase
- A silence long enough to be an ending grouping, notation
- Leaps do not fall where offbeats do melodic interval, notation
- One number for a page melodic interval, notation