Form and structure

The level the tempo chooses

The boundary detector has a scale parameter and an earlier essay left it free, ending with the sentence that names this one: the scale is in notes and the psychological present is in seconds. The psychological present is two to eight seconds, a tempo converts one to the other, and the level a listener reads then stops being a parameter at all — which is a prediction with teeth, because the same tune at two tempos should be phrased differently at levels the arithmetic names in advance.

Assumes: A boundary at a stated level · A phrase is a number of seconds

A boundary at a stated level gave the boundary detector a scale, which removed its threshold and turned a list of boundaries into a hierarchy. It ended by naming the thing it had deliberately not taken: the scale is in notes and this ladder’s first rung is in seconds.

A phrase is a number of seconds established the other half four rungs ago. The psychological present — the span of time a listener holds as a single now — is two to eight seconds, and a phrase is a phrase because it fits inside one. A tempo converts notes to seconds, so the two halves have been on this site since that rung was written and have not been multiplied together.

Twinkle, twinkle, phrased at the level each tempo selectsThe number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 5 boundaries, drawn as the flat line; the model matches it best at 160 beats per minute, where the present holds 8.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.the page's own phrasing402.0n603.1n804.1n1206.1n1608.2n20010.2n26013.3n024681012boundaries the model findsbeats per minute, with the notes inside the present under itbest at 160 bpmF = 1.00the scale is the presentdivided by the note length
Fig. 1 The number of boundaries the model finds when its smoothing scale is set by the present rather than chosen, against the tempo the tune is taken at. Under each tempo is how many notes the present holds there. A slow tempo puts few notes inside it and smooths hardly at all, so every local dip in the boundary curve survives; a fast tempo puts many notes inside it and smooths the tune into two or three arcs.

The level is no longer free, and the tune that was the fourth rung’s easy case comes out exactly right: at 160 beats per minute the present holds 8.2 notes, the model finds five boundaries, and the page has five in the same places.

The conversion, which is one division

The present is a duration. The scale parameter is a width in notes. The bridge is the mean note length, which a tempo fixes.

At 60 beats per minute a crotchet is a second and a tune of crotchets and minims averages perhaps 1.2 seconds a note, so a 3.5-second present holds about three notes. That is very nearly the number of chords a modulation needs and very nearly the number of bars an ordered key-finder needs, which is a coincidence worth noting and not more than that: all three are quantities of a few musical events, and a few musical events is what fits inside a present. At 200 the same tune averages 0.35 seconds a note and the present holds ten.

The smoothing width is then a quarter of that, because a Gaussian of standard deviation σ spans about four σ — so “a window of the present” and “a kernel of that width” are the same statement written two ways.

The boundaries that survive each amount of smoothing. The local boundary strengths of Twinkle, twinkle read at every scale: the curve is smoothed with a Gaussian of the width on the horizontal axis and the peaks that survive are counted. Small scales give 11 boundaries and large ones give one, and the notation marks 5. The level with that many falls at a width of 2, where the model finds 100 per cent of the notated boundaries and 100 per cent of what it finds is notated — a comparison with no threshold in it, which is what the scale parameter buys.
Fig. 2 The earlier figure, which this one is putting a scale on: the boundaries against σ, with each surviving to its own height and the level whose count matches the notation marked. That figure’s horizontal axis is a free parameter and its argument was that the ordering by persistence is what matters. This essay says which point on that axis a listener is standing at.

It is worth pausing on what has happened to the model. A scale-space model with a free scale is a family of readings and refuses to choose among them, which is honest and is also a way of not answering. Supplying the scale from outside turns it into a single reading with a stated commitment — and the outside supply is not arbitrary, because the psychological present is a measured quantity with its own literature.

That is the shape a model becomes a claim. Nothing about the detector changed; what changed is that a parameter is now a consequence of two published numbers rather than a dial — and the last section of this rung is about how much a consequence of two published numbers is worth when one of them is a band four wide.

What it predicts

The prediction is not that the model gets a tune right. It is that the same tune at two tempos is heard as differently phrased, and it names the levels.

At 40 beats per minute the model finds eleven boundaries in Twinkle, twinkle — every two-bar figure separately — because the present holds two notes and nothing is smoothed. At 260 it finds three, because the present holds thirteen notes and the tune is three arcs. At 160 it finds the five the page has.

That is a strong claim and it is checkable by ear rather than by argument: a tune sung slowly enough should break into more phrases than the same tune sung quickly, at ratios the arithmetic gives.

It also has a corollary that is easier to test and harder to explain away. The relationship is not merely monotone but quantitative: doubling the tempo halves the note length, doubles the notes inside the present and doubles the smoothing width — and in a scale-space, doubling σ removes the boundaries that survive only to half of it. So a factor of two in tempo should remove a specific, nameable set of boundaries, and which ones is printed on the fourth rung’s persistence ranking.

Every boundary in a tune therefore has a tempo above which it disappears, and the tempo is its persistence times the note length times four. That is a per-boundary prediction rather than a per-tune one, and it is the sharpest thing in this rung — sharper than the per-tune counts, because a ratio of two tempos does not contain the present at all. The count at any one tempo depends on the present and the tempo at which one named boundary goes depends on the present too, but the ratio between two boundaries’ disappearance tempos is their ratio of persistences and nothing else. So the ordering of the boundaries by when they vanish is a prediction with no free quantity in it, and it is the one form of the claim the last section does not weaken.

Ode to Joy, phrased at the level each tempo selectsThe number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 3 boundaries, drawn as the flat line; the model matches it best at 40 beats per minute, where the present holds 2.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.the page's own phrasing402.2n603.3n804.4n1206.6n1608.8n20010.9n26014.2n01234567boundaries the model findsbeats per minute, with the notes inside the present under itbest at 40 bpmF = 0.67the scale is the presentdivided by the note length
Fig. 3 Ode to Joy, the same way. The count falls from six boundaries at the slowest tempo to two at the fastest, and the page has three. The curve’s shape is the same as the previous one — more notes in the present means fewer boundaries — and where it crosses the page’s own count is a different tempo.

Where it does not rescue anything

The fourth rung was careful about which of its results were rescues and which were not, and this rung has to be as careful.

Frère Jacques, phrased at the level each tempo selectsThe number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 7 boundaries, drawn as the flat line; the model matches it best at 40 beats per minute, where the present holds 2.3 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.the page's own phrasing402.3n603.5n804.7n1207.0n1609.3n20011.7n26015.2n02468boundaries the model findsbeats per minute, with the notes inside the present under itbest at 40 bpmF = 0.71the scale is the presentdivided by the note length
Fig. 4 Frère Jacques, which was the hard case before and is still hard. At every tempo the model finds boundaries in the wrong places — its best reading recovers five of the page’s seven and gets three wrong — and no setting of the scale fixes that, because the scale decides how many boundaries there are and not where.

The scale removes a free parameter and cannot supply missing information. Where a phrase ends is the rung that established which cues the detector weights, and it found the agreement with the page holds only where the tune uses them. Frère Jacques is a tune of equal notes with no rests and no leaps at its phrase ends, so the boundary-strength curve has almost nothing in it at the places the page marks — and smoothing a flat curve at any width gives a flat curve.

That is exactly what the fourth rung found, and its own next step was the remedy: two per cent of final lengthening puts the boundaries where the page has them. The tempo conversion and the rubato are separate repairs to separate defects, and only the second addresses this one.

How much rubato the notated phrasing needs before the model finds it. The same boundary detector run on Frère Jacques with each notated phrase-final note lengthened by the amount on the horizontal axis, which is the one thing every measurement of expressive timing agrees a performer does. At no lengthening the model finds 57 per cent of the notated boundaries; it never reaches all of them inside a doubling of the final note. The shaded band is the published range of final lengthening, 20 to 50 per cent, so the question is whether the crossing is inside it.
Fig. 5 The other half of that earlier work, and the one that does fix this tune: the same detector on a constructed performance with the phrase-final notes lengthened, against how much lengthening it takes. Two per cent is enough, and the published range for real final lengthening starts at twenty. The information the notated version does not carry is in the performance, and no amount of scale-setting substitutes for it.

The uncertainty that is left

The present is a band and not a number, and the band is wide: two to eight seconds is a factor of four. Running the counts at both ends of it says what that costs, and the cost is comparable to the effect this rung is about.

Twinkle, boundaries found present 2 s 3.5 s 8 s
at 40 to the minute 11 11 6
at 80 11 7 5
at 160 6 6 2
at 260 6 4 1

Read across, the present’s own width moves the count by a factor of three at 160 and by six at 260. Read down, the tempo moves it from eleven to four across a factor of six and a half in speed. The two are the same size — so the uncertainty in the parameter that was supposed to fix the level is as large as the level’s whole dependence on tempo.

That is a real qualification of the level is no longer free. It is less free than it was: a free parameter admitted any reading and this one admits a band about three wide. But it is not fixed, and a prediction that names five boundaries at 160 is really a prediction of somewhere between two and six.

What survives the band is the ordering, and it survives it completely. Under every value of the present, and on all three tunes, the count at a slow tempo exceeds the count at a fast one — 11 against 6 at the narrow end, 11 against 4 in the middle, 6 against 1 at the wide end. So the prediction that the same tune breaks into more phrases when sung slowly is safe; the prediction of how many is not.

Which is exactly the limitation the fourth rung recorded about its own free parameter, arriving again one level up. That rung’s argument was that the ordering by persistence is what matters and the absolute level is not determined; this rung has replaced a free level with a measured one and has inherited the same shape of answer, because the measurement it imported has a factor of four in it. A parameter supplied from outside is only as determined as the outside quantity, and three and a half seconds is a mode rather than a value.

Twinkle, twinkle at 140, across the width of the present. The boundaries the model finds when its scale is set by the psychological present, at the three ends of that present's published range — 2, 3.5, 8 seconds — with the notation's own phrase ends on the top row. The scale in notes is the present divided by the mean note length, so the same band is a different number of notes at every tempo. At 3.5 seconds the model finds 5 boundaries against the page's 5.
Fig. 6 One tune at one tempo, read at the three ends of the present’s published range, with the page’s own phrasing on the top row. At two seconds the model finds six boundaries and one of them is wrong; at three and a half it finds exactly the page’s five; at eight it finds two. The factor of four in the present is a factor of three in the number of boundaries.

So the level is determined and the determination is only as good as the present is known. That is a real weakening and it is worth stating in the same breath as the result: the parameter has moved from being free to being measured badly, which is progress and is not the same as being fixed.

What can be said is that the typical value — three and a half seconds, which is the middle of every published estimate — is the one that works, and it works without having been chosen for it.

The three tunes, and what their disagreement is worth

The three tunes give three different answers to the same question, and the differences are informative rather than noise.

Twinkle, twinkle matches the page at 140 to 160 beats per minute with a perfect score. Ode to Joy never matches it: at the tempo where the count is right the boundaries are in the wrong places, and its best reading recovers one of three. Frère Jacques is worst of all.

The ordering is the fourth rung’s ordering exactly, which is reassuring — this rung did not change which tunes are hard — and the reason is the same. Twinkle has rests and repeated notes at its phrase ends, so its boundary-strength curve has real peaks where the page has boundaries; the other two do not. A scale decides which peaks survive and cannot invent one.

What the tempo conversion adds is that Twinkle’s match is now at a predicted level rather than a chosen one, which is the difference between a model that fits and a model that predicts.

Which computation produced the numbers

Three steps.

The mean note length is melodyTiming at the stated tempo, which is the site’s own conversion from a list of note values to seconds. It is a mean over the tune rather than a per-note quantity, which is a simplification: a tune of very unequal notes has a present that holds a different number of notes at different points in it.

The scale is the present divided by that mean and then by four. The four is the only arbitrary number in the chain and it comes from a Gaussian’s effective width; using two instead doubles every σ and shifts every curve left by an octave of tempo without changing its shape.

The boundaries are boundaryHierarchy at that single σ, which is the fourth rung’s own function called at one point on its axis rather than swept across it. Nothing else changed.

The comparison is with NOTATED_PHRASES, the phrase ends as the page marks them, with a tolerance of one note — the same tolerance the fourth rung used and for the same reason, which is that a boundary “at” a note is ambiguous between the note that ends one phrase and the one that starts the next.

Where the model stops

The tempo is a number and a performance is not. Everything here holds the tempo constant through the tune, and an ending is a deceleration established that a real performance slows at exactly the places this model is trying to find. So the present holds fewer notes at a phrase end than in the middle of one, which would sharpen the model at the moments that matter and is not in it.

The present is a span and the model uses it as a kernel. A window and a Gaussian of a quarter of its width are not the same object; a window has edges and a Gaussian has tails. The choice matters least where the boundary strength is spiky and most where it is flat, which is the hard case again.

And nothing here is heard. The prediction — that a tune sung slowly is phrased more finely — has never been tested, and the test is a listening experiment of exactly the kind this collection cannot run. What it can do is say what the experiment would have to find, and the figures are that.

Whose music, and what a tempo marking is for

The arithmetic is about a listener and applies to any monophonic line. The consequence is about notation, and it is a claim about a practice.

A tempo marking is usually read as an instruction about speed, and this rung says it is also an instruction about phrasing. Two performances of the same page at 60 and at 160 do not differ only in how long they take; if the model is right, they present different numbers of phrases to a listener, and the composer who wrote the marking chose which.

That reframes a familiar argument in early-music performance. The disputes about whether a movement goes at one tempo or at twice it are usually conducted as disputes about the character of the music or about the meaning of a time signature. The arithmetic here says that the two tempos give different phrase structures — not different renderings of one structure — so the question has a consequence beyond character.

It also gives a reading of the practice this collection has been circling since the phrase ladder’s first rung. Phrases in the European tradition are written in bars, and bars are a notational unit that does not know about tempo; but the tunes that are sung at slow tempos have short phrases in bars, and the tunes taken fast have long ones. A chorale’s phrase is two bars and a scherzo’s is eight, and the arithmetic of the ladder’s first rung says that both are about three and a half seconds. The bar count is what varies so that the seconds can stay put, which is the first rung’s finding arriving from the other end.

What the picture cannot show

Which tempo a tune belongs at. The figures report the tempo at which the model matches the page, and that is not evidence that the tune is sung there — it is evidence about the model. Turning it round and using the model to infer a tempo from a phrasing would be a genuinely useful thing and it would need the model to be right first.

Nor whether a listener has one present or several. The model uses a single width, and the metrical hierarchy is several levels at once — a listener attending to a bar and to a four-bar group simultaneously is attending at two scales, and there is no reason the present should pick out one of them and no other.

Nor how a listener knows the tempo before the phrasing. The conversion needs a tempo, and a tempo is inferred — the beat is inferred, and sometimes wrongly. So the level is set by a quantity the listener is also working out, and the two inferences run at once.

And the tunes are three. Everything here runs on the three melodies this collection carries, and three is not a corpus. The mechanism is general and the numbers are anecdotes.

Where this ladder goes next

Five rungs. A phrase is a number of seconds and the bar count is whatever the tempo makes it; the sentence and the period differ by one computable ratio; a boundary detector agrees with the page only where the tune uses the cue it weights; the detector has a scale, which removes its threshold; and now the scale has a unit, which removes the parameter.

The rung after it is the one the constant-tempo limitation names, and it joins two things this ladder already has. A real performance slows into a phrase end, so the number of notes inside the present is not constant through a tune — it falls exactly where a boundary is. That makes the smoothing width a function of position, which is a different kind of model: not a scale-space with one axis but a locally adaptive one, where the detector’s own resolution is set by the performance it is listening to. Both halves are on this site, in this ladder, and they have never been put together.

Part 5 of 8

One essay in the series on phrase. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

GroupingHierarchyLbdmPerceptual presentPhraseScale-spaceSegmentationTempo