Series

Phrase — the series

8 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. A phrase is a number of seconds, and the bars follow the tempo. Phrase durations for 1, 2, 4, 8, 16-bar phrases at seven tempos, on a logarithmic seconds axis, with the 2 to 8 second window shaded. The window is a property of the listener and does not move; which bar count falls inside it is decided entirely by the tempo.

    A phrase is a number of seconds

    Musical phrases are described in bars, and four is the number everybody names. But the constraint that fixes a phrase is a property of the listener and is measured in seconds, so the bar count is whatever the tempo makes it. Across seven ordinary tempos the bar count that lands inside the window moves by a factor of eight, while the window itself does not move at all.

    part 1 · form
  2. Two ways to fill eight bars, and only one of them accelerates. The period against the sentence, drawn as the lengths of their constituent units against position on a grid of eight bars. The ratio beside each row is the mean unit length in its second half divided by the mean in its first: the period at 1.00, the sentence at 0.67. A ratio below one is an acceleration — the unit shortening as the phrase approaches its arrival — and a ratio of one is a plan whose unit never changes length.

    One of these eight-bar phrases accelerates

    The sentence and the period both occupy eight bars, both end with a cadence, and both are recognised by ear rather than counted. What separates them is arithmetic. One halves its unit halfway through and the other does not, and the difference comes out as a single ratio — 0.67 against 1.00 — computed from nothing but the lengths of the parts.

    part 2 · form
  3. Where the page ends a phrase, and where the ear does. Twinkle, twinkle with two sets of phrase boundaries on it. The lower curve is a local boundary detector — a peak in how much the interval and the note length change from one to the next, with nothing in it about bar lines or harmony — and the marks above it are where the notation puts the phrase ends. It finds 100 per cent of them and 2 boundaries the page does not have. Where the two agree it is because a long note is sitting at the join; where they disagree the page is marking a grammatical unit and the detector is finding a perceptual one.

    Where a phrase ends

    Run a boundary detector over the three tunes used throughout and it agrees with the notated phrasing on one of them perfectly and on another almost not at all. The reason is which cue each tune uses: Twinkle's phrases all end on a long note, so a duration-weighted detector finds five of five with no false alarms; Ode to Joy's run on in crotchets and its phrasing is in the intervals, where a duration detector finds one of three and a pitch detector finds all three and eight others. No fixed weighting serves both, and the published one is worse on each tune than the single cue that tune uses.

    part 3 · form
  4. The boundaries that survive each amount of smoothing. The local boundary strengths of Twinkle, twinkle read at every scale: the curve is smoothed with a Gaussian of the width on the horizontal axis and the peaks that survive are counted. Small scales give 11 boundaries and large ones give one, and the notation marks 5. The level with that many falls at a width of 2, where the model finds 100 per cent of the notated boundaries and 100 per cent of what it finds is notated — a comparison with no threshold in it, which is what the scale parameter buys.

    A boundary at a stated level

    A boundary detector run over three tunes agreed with the notation on one and barely at all on another, and left two things owing: a version with a scale parameter, and a version run on performance timings. Both are paid here, and they pay differently — the scale removes a free parameter from the comparison and does not rescue the hard case, while two per cent of rubato does.

    part 4 · form
  5. Twinkle, twinkle, phrased at the level each tempo selects. The number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 5 boundaries, drawn as the flat line; the model matches it best at 160 beats per minute, where the present holds 8.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.

    The level the tempo chooses

    The boundary detector has a scale parameter and an earlier essay left it free, ending with the sentence that names this one: the scale is in notes and the psychological present is in seconds. The psychological present is two to eight seconds, a tempo converts one to the other, and the level a listener reads then stops being a parameter at all — which is a prediction with teeth, because the same tune at two tempos should be phrased differently at levels the arithmetic names in advance.

    part 5 · form
  6. How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.

    A detector whose resolution the performance sets

    The boundary detector lost its free parameter when the psychological present became a number of notes at a stated tempo, and what that held still was named at the time: a performance slows into a phrase end, so the number of notes inside the present is not the same everywhere in a tune — it falls exactly where a boundary is. Making the width follow the performance recovers half of what removing the parameter cost, and honestly leaves the other half.

    part 6 · form
  7. The same tune read at six widths of the psychological present. A later essay made the detector's smoothing width a function of position, which removed its last free parameter but one — and the one it cannot remove is the width of the psychological present, because that is a fact about listeners rather than a choice. So the honest object is not a reading but a family of them, one per width. A short present finds 6 boundaries and a long one finds 2, and the family agrees on 0 of them. The fixed-width control, at its own best width, scores 0.67 against the adaptive readings' 0.67, 0.75, 0.33, 0.33, 0.40, 0.40 — so the adaptation does not win, which is what that essay reported too. What the family adds is the ordering: a boundary in every row is a different claim from one in a single row, and a single reading has no way to say so.

    A family of readings

    Removing the detector's free parameter, and then its constant tempo, cost persistence both times — the property that made its boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than one, indexed by the width of the psychological present, which is the one parameter that cannot be removed, because it is a fact about listeners.

    part 7 · rhythm
  8. The breath is the looser ceiling nearly everywhere. How long a trained singer can hold a phrase on one breath, across a compass and at four dynamics, against the 8-second ceiling the psychological present puts on the same phrase. The flow through the folds rises with pitch and with loudness, so the breath ceiling falls both ways: at 60 decibels it runs 32.6 seconds at the bottom of the compass to 21.2 at the top; at 70 decibels it runs 23.1 seconds at the bottom of the compass to 15.0 at the top; at 80 decibels it runs 16.4 seconds at the bottom of the compass to 10.6 at the top; at 90 decibels it runs 11.6 seconds at the bottom of the compass to 7.5 at the top. The shaded line is the listener's ceiling and it does not move. The breath binds only where the two lines cross — 1 of the 40 cells drawn, all of them loud and high. So the constraint everybody names when asked why a phrase is the length it is, is almost never the constraint that decides it.

    The ceiling everybody names is the loose one

    Ask why phrases are the length they are and the answer given is the breath. It is arithmetic — usable lung volume over the air a note costs per second — and it comes out between fifteen and twenty-three seconds at a comfortable dynamic and between seven and twelve at a loud one. The ceiling the present moment imposes, the two-to-eight seconds inside which a stretch is heard as one thing rather than as a series, is two to three times tighter at almost every note and dynamic. A singer in an adagio is not running out of breath at the phrase end. They are running out of present.

    part 8 · form

All series