The level the tempo chooses
Assumes: A boundary at a stated level · A phrase is a number of seconds
A boundary at a stated level gave the boundary detector a scale, which removed its threshold and turned a list of boundaries into a hierarchy. It ended by naming the thing it had deliberately not taken: the scale is in notes and this ladder’s first rung is in seconds.
A phrase is a number of seconds established the other half four rungs ago. The psychological present — the span of time a listener holds as a single now — is two to eight seconds, and a phrase is a phrase because it fits inside one. A tempo converts notes to seconds, so the two halves have been on this site since that rung was written and have not been multiplied together.
The level is no longer free, and the tune that was the fourth rung’s easy case comes out exactly right: at 160 beats per minute the present holds 8.2 notes, the model finds five boundaries, and the page has five in the same places.
The conversion, which is one division
The present is a duration. The scale parameter is a width in notes. The bridge is the mean note length, which a tempo fixes.
At 60 beats per minute a crotchet is a second and a tune of crotchets and minims averages perhaps 1.2 seconds a note, so a 3.5-second present holds about three notes. That is very nearly the number of chords a modulation needs and very nearly the number of bars an ordered key-finder needs, which is a coincidence worth noting and not more than that: all three are quantities of a few musical events, and a few musical events is what fits inside a present. At 200 the same tune averages 0.35 seconds a note and the present holds ten.
The smoothing width is then a quarter of that, because a Gaussian of standard deviation σ spans about four σ — so “a window of the present” and “a kernel of that width” are the same statement written two ways.
It is worth pausing on what has happened to the model. A scale-space model with a free scale is a family of readings and refuses to choose among them, which is honest and is also a way of not answering. Supplying the scale from outside turns it into a single reading with a stated commitment — and the outside supply is not arbitrary, because the psychological present is a measured quantity with its own literature.
That is the shape a model becomes a claim. Nothing about the detector changed; what changed is that a parameter is now a consequence of two published numbers rather than a dial — and the last section of this rung is about how much a consequence of two published numbers is worth when one of them is a band four wide.
What it predicts
The prediction is not that the model gets a tune right. It is that the same tune at two tempos is heard as differently phrased, and it names the levels.
At 40 beats per minute the model finds eleven boundaries in Twinkle, twinkle — every two-bar figure separately — because the present holds two notes and nothing is smoothed. At 260 it finds three, because the present holds thirteen notes and the tune is three arcs. At 160 it finds the five the page has.
That is a strong claim and it is checkable by ear rather than by argument: a tune sung slowly enough should break into more phrases than the same tune sung quickly, at ratios the arithmetic gives.
It also has a corollary that is easier to test and harder to explain away. The relationship is not merely monotone but quantitative: doubling the tempo halves the note length, doubles the notes inside the present and doubles the smoothing width — and in a scale-space, doubling σ removes the boundaries that survive only to half of it. So a factor of two in tempo should remove a specific, nameable set of boundaries, and which ones is printed on the fourth rung’s persistence ranking.
Every boundary in a tune therefore has a tempo above which it disappears, and the tempo is its persistence times the note length times four. That is a per-boundary prediction rather than a per-tune one, and it is the sharpest thing in this rung — sharper than the per-tune counts, because a ratio of two tempos does not contain the present at all. The count at any one tempo depends on the present and the tempo at which one named boundary goes depends on the present too, but the ratio between two boundaries’ disappearance tempos is their ratio of persistences and nothing else. So the ordering of the boundaries by when they vanish is a prediction with no free quantity in it, and it is the one form of the claim the last section does not weaken.
Where it does not rescue anything
The fourth rung was careful about which of its results were rescues and which were not, and this rung has to be as careful.
The scale removes a free parameter and cannot supply missing information. Where a phrase ends is the rung that established which cues the detector weights, and it found the agreement with the page holds only where the tune uses them. Frère Jacques is a tune of equal notes with no rests and no leaps at its phrase ends, so the boundary-strength curve has almost nothing in it at the places the page marks — and smoothing a flat curve at any width gives a flat curve.
That is exactly what the fourth rung found, and its own next step was the remedy: two per cent of final lengthening puts the boundaries where the page has them. The tempo conversion and the rubato are separate repairs to separate defects, and only the second addresses this one.
The uncertainty that is left
The present is a band and not a number, and the band is wide: two to eight seconds is a factor of four. Running the counts at both ends of it says what that costs, and the cost is comparable to the effect this rung is about.
| Twinkle, boundaries found | present 2 s | 3.5 s | 8 s |
|---|---|---|---|
| at 40 to the minute | 11 | 11 | 6 |
| at 80 | 11 | 7 | 5 |
| at 160 | 6 | 6 | 2 |
| at 260 | 6 | 4 | 1 |
Read across, the present’s own width moves the count by a factor of three at 160 and by six at 260. Read down, the tempo moves it from eleven to four across a factor of six and a half in speed. The two are the same size — so the uncertainty in the parameter that was supposed to fix the level is as large as the level’s whole dependence on tempo.
That is a real qualification of the level is no longer free. It is less free than it was: a free parameter admitted any reading and this one admits a band about three wide. But it is not fixed, and a prediction that names five boundaries at 160 is really a prediction of somewhere between two and six.
What survives the band is the ordering, and it survives it completely. Under every value of the present, and on all three tunes, the count at a slow tempo exceeds the count at a fast one — 11 against 6 at the narrow end, 11 against 4 in the middle, 6 against 1 at the wide end. So the prediction that the same tune breaks into more phrases when sung slowly is safe; the prediction of how many is not.
Which is exactly the limitation the fourth rung recorded about its own free parameter, arriving again one level up. That rung’s argument was that the ordering by persistence is what matters and the absolute level is not determined; this rung has replaced a free level with a measured one and has inherited the same shape of answer, because the measurement it imported has a factor of four in it. A parameter supplied from outside is only as determined as the outside quantity, and three and a half seconds is a mode rather than a value.
So the level is determined and the determination is only as good as the present is known. That is a real weakening and it is worth stating in the same breath as the result: the parameter has moved from being free to being measured badly, which is progress and is not the same as being fixed.
What can be said is that the typical value — three and a half seconds, which is the middle of every published estimate — is the one that works, and it works without having been chosen for it.
The three tunes, and what their disagreement is worth
The three tunes give three different answers to the same question, and the differences are informative rather than noise.
Twinkle, twinkle matches the page at 140 to 160 beats per minute with a perfect score. Ode to Joy never matches it: at the tempo where the count is right the boundaries are in the wrong places, and its best reading recovers one of three. Frère Jacques is worst of all.
The ordering is the fourth rung’s ordering exactly, which is reassuring — this rung did not change which tunes are hard — and the reason is the same. Twinkle has rests and repeated notes at its phrase ends, so its boundary-strength curve has real peaks where the page has boundaries; the other two do not. A scale decides which peaks survive and cannot invent one.
What the tempo conversion adds is that Twinkle’s match is now at a predicted level rather than a chosen one, which is the difference between a model that fits and a model that predicts.
Which computation produced the numbers
Three steps.
The mean note length is melodyTiming at the stated tempo, which is the site’s own conversion from a list of note values to seconds. It is a mean over the tune rather than a per-note quantity, which is a simplification: a tune of very unequal notes has a present that holds a different number of notes at different points in it.
The scale is the present divided by that mean and then by four. The four is the only arbitrary number in the chain and it comes from a Gaussian’s effective width; using two instead doubles every σ and shifts every curve left by an octave of tempo without changing its shape.
The boundaries are boundaryHierarchy at that single σ, which is the fourth rung’s own function called at one point on its axis rather than swept across it. Nothing else changed.
The comparison is with NOTATED_PHRASES, the phrase ends as the page marks them, with a tolerance of one note — the same tolerance the fourth rung used and for the same reason, which is that a boundary “at” a note is ambiguous between the note that ends one phrase and the one that starts the next.
Where the model stops
The tempo is a number and a performance is not. Everything here holds the tempo constant through the tune, and an ending is a deceleration established that a real performance slows at exactly the places this model is trying to find. So the present holds fewer notes at a phrase end than in the middle of one, which would sharpen the model at the moments that matter and is not in it.
The present is a span and the model uses it as a kernel. A window and a Gaussian of a quarter of its width are not the same object; a window has edges and a Gaussian has tails. The choice matters least where the boundary strength is spiky and most where it is flat, which is the hard case again.
And nothing here is heard. The prediction — that a tune sung slowly is phrased more finely — has never been tested, and the test is a listening experiment of exactly the kind this collection cannot run. What it can do is say what the experiment would have to find, and the figures are that.
Whose music, and what a tempo marking is for
The arithmetic is about a listener and applies to any monophonic line. The consequence is about notation, and it is a claim about a practice.
A tempo marking is usually read as an instruction about speed, and this rung says it is also an instruction about phrasing. Two performances of the same page at 60 and at 160 do not differ only in how long they take; if the model is right, they present different numbers of phrases to a listener, and the composer who wrote the marking chose which.
That reframes a familiar argument in early-music performance. The disputes about whether a movement goes at one tempo or at twice it are usually conducted as disputes about the character of the music or about the meaning of a time signature. The arithmetic here says that the two tempos give different phrase structures — not different renderings of one structure — so the question has a consequence beyond character.
It also gives a reading of the practice this collection has been circling since the phrase ladder’s first rung. Phrases in the European tradition are written in bars, and bars are a notational unit that does not know about tempo; but the tunes that are sung at slow tempos have short phrases in bars, and the tunes taken fast have long ones. A chorale’s phrase is two bars and a scherzo’s is eight, and the arithmetic of the ladder’s first rung says that both are about three and a half seconds. The bar count is what varies so that the seconds can stay put, which is the first rung’s finding arriving from the other end.
What the picture cannot show
Which tempo a tune belongs at. The figures report the tempo at which the model matches the page, and that is not evidence that the tune is sung there — it is evidence about the model. Turning it round and using the model to infer a tempo from a phrasing would be a genuinely useful thing and it would need the model to be right first.
Nor whether a listener has one present or several. The model uses a single width, and the metrical hierarchy is several levels at once — a listener attending to a bar and to a four-bar group simultaneously is attending at two scales, and there is no reason the present should pick out one of them and no other.
Nor how a listener knows the tempo before the phrasing. The conversion needs a tempo, and a tempo is inferred — the beat is inferred, and sometimes wrongly. So the level is set by a quantity the listener is also working out, and the two inferences run at once.
And the tunes are three. Everything here runs on the three melodies this collection carries, and three is not a corpus. The mechanism is general and the numbers are anecdotes.
Where this ladder goes next
Five rungs. A phrase is a number of seconds and the bar count is whatever the tempo makes it; the sentence and the period differ by one computable ratio; a boundary detector agrees with the page only where the tune uses the cue it weights; the detector has a scale, which removes its threshold; and now the scale has a unit, which removes the parameter.
The rung after it is the one the constant-tempo limitation names, and it joins two things this ladder already has. A real performance slows into a phrase end, so the number of notes inside the present is not constant through a tune — it falls exactly where a boundary is. That makes the smoothing width a function of position, which is a different kind of model: not a scale-space with one axis but a locally adaptive one, where the detector’s own resolution is set by the performance it is listening to. Both halves are on this site, in this ladder, and they have never been put together.
Part 5 of 8
One essay in the series on phrase. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
GroupingHierarchyLbdmPerceptual presentPhraseScale-spaceSegmentationTempo
- A silence long enough to be an ending grouping, perceptual present, tempo
- The parameter that did not decide the answer perceptual present, phrase, tempo
- A form is sharp at the bottom and vague at the top perceptual present, phrase
- A proportion is only as fine as its two durations perceptual present, phrase
- An ending that exists so a bigger one can hierarchy, phrase
- How long a limping bar can be perceptual present, tempo