Form and structure

The ceiling everybody names is the loose one

Ask why phrases are the length they are and the answer given is the breath. It is arithmetic — usable lung volume over the air a note costs per second — and it comes out between fifteen and twenty-three seconds at a comfortable dynamic and between seven and twelve at a loud one. The ceiling the present moment imposes, the two-to-eight seconds inside which a stretch is heard as one thing rather than as a series, is two to three times tighter at almost every note and dynamic. A singer in an adagio is not running out of breath at the phrase end. They are running out of present.

Assumes: A phrase is a number of seconds · A family of readings

The essay that put a phrase’s length in seconds put a phrase’s length in seconds rather than bars, and gave the reason as the psychological present: a stretch of two to eight seconds is heard as one thing rather than as a series of things, so the phrase is that long and the bar count is whatever the tempo makes it. Later essays replaced the fixed window with a scale the tempo chooses and then with a resolution the performance sets, and both of those are still ceilings on a listener.

That is a ceiling on a listener. It is not the ceiling anybody names. Ask a musician why phrases are the length they are and the answer is the breath — a phrase is what a singer can sing on one lungful, and instrumental phrasing inherits the shape from the voice.

Both ceilings are real and only one of them has ever been computed here. The second is arithmetic and the arithmetic is one line.

The breath is a division

A singer holding a note is emptying a reservoir at a rate. The reservoir is the usable volume above residual — about 1.8 litres for an untrained adult and 2.8 for a trained one — and the rate is the mean transglottal flow, which is what the folds let through per second.

seconds on one breath=usable volumemean flow\text{seconds on one breath} = \frac{\text{usable volume}}{\text{mean flow}}

The flow is not a constant. A comfortable middle-register note runs around 140 millilitres a second; a higher note needs more subglottal pressure and costs more; a louder one costs more still. Those two dependencies are what make the ceiling a surface rather than a number.

The breath is the looser ceiling nearly everywhere. How long a trained singer can hold a phrase on one breath, across a compass and at four dynamics, against the 8-second ceiling the psychological present puts on the same phrase. The flow through the folds rises with pitch and with loudness, so the breath ceiling falls both ways: at 60 decibels it runs 32.6 seconds at the bottom of the compass to 21.2 at the top; at 70 decibels it runs 23.1 seconds at the bottom of the compass to 15.0 at the top; at 80 decibels it runs 16.4 seconds at the bottom of the compass to 10.6 at the top; at 90 decibels it runs 11.6 seconds at the bottom of the compass to 7.5 at the top. The shaded line is the listener's ceiling and it does not move. The breath binds only where the two lines cross — 1 of the 40 cells drawn, all of them loud and high. So the constraint everybody names when asked why a phrase is the length it is, is almost never the constraint that decides it.
Fig. 1 How long a trained singer can hold a phrase on one breath, across a compass and at four dynamics, against the eight-second ceiling the psychological present puts on the same phrase. The heavy horizontal line is the listener’s ceiling and it does not move.

At a comfortable seventy decibels the breath ceiling runs from twenty-three seconds at the bottom of the compass to fifteen at the top. At ninety it runs from 11.6 to 7.5.

Which is two to three times looser than the other one

The listener’s ceiling is eight seconds and it does not depend on the note, the dynamic or the singer. At a comfortable dynamic the breath ceiling is between two and three times looser, and it stays looser everywhere except in one corner.

That corner is worth naming precisely because it is the only place the received explanation is right. A trained singer at ninety decibels — a full voice, not a comfortable one — crosses the eight-second line only above about C6. An untrained singer at the same dynamic and pitch has 5.2 seconds and crosses it comfortably. Everywhere else the two ceilings are not close.

The breath is the looser ceiling nearly everywhere. How long a untrained singer can hold a phrase on one breath, across a compass and at four dynamics, against the 8-second ceiling the psychological present puts on the same phrase. The flow through the folds rises with pitch and with loudness, so the breath ceiling falls both ways: at 70 decibels it runs 14.8 seconds at the bottom of the compass to 9.7 at the top; at 80 decibels it runs 10.5 seconds at the bottom of the compass to 6.8 at the top; at 90 decibels it runs 7.5 seconds at the bottom of the compass to 4.9 at the top. The shaded line is the listener's ceiling and it does not move. The breath binds only where the two lines cross — 14 of the 30 cells drawn, all of them loud and high. So the constraint everybody names when asked why a phrase is the length it is, is almost never the constraint that decides it.
Fig. 2 The same surface for an untrained singer, with 1.8 litres usable rather than 2.8. The whole family moves down by a third, and the loud rows cross the listener’s ceiling over most of the compass rather than at the top of it.

So the two constraints are not competing descriptions of one number. They are constraints on different things, two to three times apart, and the tighter one is almost always the listener’s.

What that does to a four-bar phrase

The first essay’s central figure is a phrase’s written length in seconds, at seven tempos, read against the listener’s window. Adding the second ceiling puts a floor under that reading rather than a second wall beside it.

The listener's ceiling binds at every tempo and every note. A 4-bar phrase at seven tempos and seven notes, sung at 85 decibels by a trained singer, with which of the two ceilings is the lower one in each cell and whether the written phrase fits under it. The written length depends only on the tempo — 18.5 seconds in a slow movement, 10.9 seconds in a sarabande, 13.3 seconds in a chorale, 6.0 seconds in a minuet, 7.0 seconds in an allegro, 4.0 seconds in a scherzo, 7.5 seconds in a house track — and the ceiling depends only on the note and the dynamic. Nowhere in the grid is the breath the binding constraint. A four-bar phrase in a slow movement overruns the listener's ceiling by a factor of two while sitting comfortably inside the singer's.
Fig. 3 A four-bar phrase at seven tempos and seven notes, sung at eighty-five decibels, with which of the two ceilings is the lower in each cell and whether the written phrase fits under it. The written length depends only on the tempo; the ceiling depends only on the note and the dynamic.

A four-bar phrase in a slow movement is sixteen seconds long. That is twice the listener’s ceiling and comfortably inside a trained singer’s breath at any dynamic below a forte. So a singer in an adagio, arriving at the end of a long phrase, is not running out of air. They are at the end of something the listener stopped hearing as one thing about halfway through.

The same phrase in a scherzo is four seconds — inside both ceilings with room to spare, which is why fast music can write longer phrases in bars and does.

A phrase is a number of seconds, and the bars follow the tempo. Phrase durations for 1, 2, 4, 8, 16-bar phrases at seven tempos, on a logarithmic seconds axis, with the 2 to 8 second window shaded. The window is a property of the listener and does not move; which bar count falls inside it is decided entirely by the tempo.
Fig. 4 The first essay’s own figure: phrase durations at seven tempos on a logarithmic seconds axis, with the listener’s window shaded. Which bar count lands inside it is decided entirely by the tempo, and nothing about the performer is in it.

The two ceilings move in opposite directions with tempo

There is a structural difference between the two constraints that the surfaces above hide, and it decides which one a composer is negotiating with.

The listener’s ceiling is in seconds and does not care about tempo. Eight seconds is eight seconds whether the music is an adagio or a scherzo, so the bar count it permits rises directly with the tempo: two bars of four at 60, three and a third at 100, five and a third at 160.

The breath ceiling is in seconds too, and so does not care either — but it is spent by a singer who is also singing notes, and a faster tempo puts more notes into the same air. A phrase of sixteen quavers at 60 and one of sixteen quavers at 160 cost nearly the same air, because the air is spent per second and not per note.

So both ceilings scale the same way and the gap between them is constant. That is why the ratio of two to three holds across the whole table and why no tempo brings the breath into play. A composer writing faster is buying bars from both ceilings at the same rate, and the one that runs out first at every speed is the listener’s.

The one thing tempo does change is the cost of being wrong. At a slow tempo the listener’s ceiling is exceeded by a bar and a half of a four-bar phrase; at a fast one the same phrase fits with room over. So the slow movements are where the constraint bites, and slow movements are also where phrase marks are longest and where singers are asked for the most — which is consistent with the constraint being real and with everybody attributing it to the wrong thing.

What the breath does instead

None of this says the breath is irrelevant, and what it actually does is more specific than what it is credited with.

It fixes where a rest goes, not how long a phrase is. A singer has to breathe somewhere, and a breath takes time — a quarter of a second snatched, a second taken comfortably. A phrase whose length is set by the listener’s ceiling still has to have a breath placed at one of its boundaries, and which boundary is a decision the composer makes for the singer. That is a question about where a boundary goes rather than how far apart boundaries are, and it is the essay that ran a boundary detector over three tunes question rather than the first’s.

It fixes the maximum, and maxima matter at the extremes. A vocal line that asks for fourteen seconds on a high forte is asking for something outside the ceiling, and composers who write such lines are known for it. The constraint is real and it binds in the repertoire that tests it.

And it distinguishes the instruments. A wind player has the same division with different numbers; a string player does not have it at all, because a bow change can be made inaudible and a bow is not a lung. If the breath were what set phrase length, string writing would have systematically longer phrases than vocal writing at the same tempo. That is a claim about repertoire, it is checkable, and there is no corpus to check it on — the same shortfall a boundary at a stated level recorded when it needed performance timings — but the prediction of the other ceiling is that the two should be the same, because the listener is the same.

Why the received answer is nearly right anyway

There is a reason the breath explanation survives despite being the loose constraint, and it is not that musicians are careless.

The two ceilings are not independent of each other. Both are, in the end, about a few seconds, and the reason both are about a few seconds may be the same reason: a voice and a memory both evolved in a creature whose utterances are a few seconds long. A ceiling of eight seconds and a ceiling of twenty are different numbers, and they are the same order of magnitude in a way that a ceiling of eight seconds and one of two hundred would not be.

It is the same shape one of these eight-bar phrases accelerates found in the sentence and the period: two things named by one word, distinguished by arithmetic nobody had done. So the breath is a good explanation of the scale and a bad explanation of the value. It gets a listener to within a factor of three of the right answer by a mechanism that is not the operative one — which is a common enough shape in this collection that it is worth naming as one.

The listener's ceiling binds at every tempo and every note. A 8-bar phrase at seven tempos and seven notes, sung at 70 decibels by a trained singer, with which of the two ceilings is the lower one in each cell and whether the written phrase fits under it. The written length depends only on the tempo — 36.9 seconds in a slow movement, 21.8 seconds in a sarabande, 26.7 seconds in a chorale, 12.0 seconds in a minuet, 13.9 seconds in an allegro, 8.0 seconds in a scherzo, 15.0 seconds in a house track — and the ceiling depends only on the note and the dynamic. Nowhere in the grid is the breath the binding constraint. A four-bar phrase in a slow movement overruns the listener's ceiling by a factor of two while sitting comfortably inside the singer's.
Fig. 5 An eight-bar phrase at a comfortable dynamic, which is the phrase everybody means when they say four or eight bars. At no tempo and no note is the breath the binding ceiling; at the three slowest tempos the written phrase overruns the listener’s by a factor of two or more.

The eight-bar case is the one to carry, because it is the phrase length the whole tradition names. Sung at a comfortable dynamic, the breath never binds it anywhere. The listener’s ceiling binds it at every tempo below about a hundred and twenty.

What a ceiling is not

Both of these are maxima and neither is a target, and the difference matters more here than it usually does.

A singer does not sing until the air runs out. The last part of a breath is the part with the least control over intonation, dynamic and vibrato, so the working length of a phrase is well under the ceiling — perhaps two thirds of it, though nothing here can say. That makes the breath figures above generous by a margin nobody has measured, and the margin runs in the direction that would narrow the gap this essay is about.

The listener’s ceiling behaves the other way. Eight seconds is the upper end of a range whose lower end is two, and a phrase is not better for being close to eight: a stretch at the top of the window is one thing only barely, and a stretch in the middle is one thing comfortably. So the working length there is also well under the ceiling, and it is under it for a reason that argues for shorter phrases rather than longer.

Both working lengths are therefore lower than both ceilings, and there is no reason to think they are lowered by the same fraction. That is the honest limit on the comparison: what has been computed is which wall is nearer, and the walls are not where anybody stands. It is enough for the essay’s claim, which is about which constraint a composer is negotiating with, and it is not enough to say what a phrase’s length actually is.

Which computation produced the numbers

The usable volume is the volume above residual, taken as 1.8 litres untrained and 2.8 trained, which are ordinary figures for an adult and are the only two measured constants here.

The mean transglottal flow at a reference of middle C and seventy decibels is taken as 140 millilitres a second, and it is scaled multiplicatively: 1.2 per cent per semitone above the reference and 3.5 per cent per decibel above it. Those two slopes are the model’s own choices, fitted to nothing, and they are chosen so that a loud high note costs roughly two and a half times a comfortable middle one — which is the size of the effect published measurements report rather than any particular measurement.

The ceiling is the division and nothing else. The listener’s ceiling is the upper end of the psychological present as the essays here carry it, eight seconds, which is the same number the first essay used.

Every number here is a ceiling rather than a typical value. A singer holding a note until the air runs out is not singing; the usable fraction in performance is well under one, and what the figures give is the wall rather than the working length.

Where the model stops

Flow is not the whole cost. A singer spends air on phonation and also on the parts of a phrase that are not sustained tone — consonants, onsets, and the wastage of an inefficient adduction. A breathy voice quality can double the flow at the same loudness, which moves the whole surface down and is not in the model.

The two slopes are the model’s weakest point and they are stated rather than hidden. Neither the 1.2 per cent per semitone nor the 3.5 per cent per decibel is a measurement; they are chosen so that the surface has the shape published work describes, and the conclusion rests on the size of the gap rather than on either slope. Doubling both — a far steeper dependence than anything reported — brings the ninety-decibel row under the listener’s ceiling over most of the compass and leaves the comfortable rows where they are, so what the slopes decide is how large the corner is, not whether the gap exists at a normal dynamic.

Nor is the reservoir emptied evenly. Lung recoil does work for a singer at high volume and against them at low, so the last half-litre costs more effort than the first even at constant flow. That is about difficulty rather than duration and it means the ceiling is reached with declining control rather than abruptly.

And a phrase is not one note. The figures above price a sustained tone, and a real phrase has articulation, rests within it and notes of varying pitch and dynamic. The mean flow over a phrase is lower than the flow at its loudest note, which makes the ceiling above conservative — in the direction that widens the gap the essay is about.

What the picture cannot show

It cannot show a wind player. The division is the same and the numbers are not: a flute at a fortissimo in its top register is famously expensive and a clarinet in its chalumeau is famously cheap, and a brass player’s limit is often the lip rather than the lung. Every one of those is the same arithmetic with a different flow, and there is no flow figures for them.

Nor a phrase mark. A slur in a score is a notational object with several meanings, one of which is a bowing and another of which is a breath. Where a phrase ends is a question a detector answers from the notes, and a breath mark is a different kind of evidence that the detector here has no term for.

And it cannot show what a listener does with a breath. A breath is audible, and an audible breath is a boundary cue in its own right — arguably the strongest one a vocal line has. Whether a listener uses it as a phrase boundary, or discounts it as a necessity of the medium, is exactly the sort of question the family of readings would need answered before its detector could be run on a sung performance.

Still open: whether a breath is a boundary the detector should see

This account’s detector reads a melody as notes and finds boundaries in their durations and intervals. A sung performance contains something the detector cannot see and a listener cannot miss: a silence with a sound in it.

Adding a breath to the detector is not a matter of adding a note. A breath occupies time that belongs to neither side, it is louder in some styles than others, and — the awkward part — it is placed by a performer at a boundary the detector is supposed to be finding independently. A detector that used breaths would be scoring its own answer sheet.

What would make the question tractable is a case where the two disagree. A singer takes a breath mid-phrase when the phrase is too long for the air, which by the arithmetic above happens where the written phrase exceeds the breath ceiling rather than where it exceeds the listener’s — so the disagreements should cluster in exactly the corner this essay identified: long phrases, high, loud. Finding those cases in a recording and asking whether a listener hears a boundary there would say whether a breath is a cue a listener uses or a cost they discount, and it is the one question about phrasing that could be settled with a microphone rather than a corpus.

Part 8 of 8

One essay in the series on phrase. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DurationDynamicsMusical formPerceptual presentPhraseRegisterSinging