Form and structure

A final chord is not made loud by adding to it

An earlier essay on closure said the loudest cue an ending has needs a corpus rather than an arithmetic. The arithmetic was built one essay ago, so it does not. Run four ending textures through it and two things come out backwards: a final chord three parts thicker than the rest arrives *quieter* against the running impression than the passage it ends, and a texture that drops a part a bar does not get quieter at all until the bar where there is one part left.

Assumes: A silence long enough to be an ending · The dynamics are in the score already

The fifth rung of this ladder closed on a silence and ended by naming what was still unmeasured:

Every ending in every repertoire is also a dynamic event — a diminuendo, a final chord louder than the rest, a texture thinning — and this collection has no model of dynamics in a form at all. That is the next rung, and it is the one that would need a corpus rather than an arithmetic.

It needed a corpus because the only route to a dynamic curve appeared to be a recording. The previous essay found another route: count the parts, realise them, put every partial in its band, sum, and smooth. That produces a dynamic reading of a page with no performance in it — so the debt is payable after all, and this rung is what it buys.

Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness.
Fig. 1 Four endings through the closing six bars of a thirty-two bar scheme, as short-term loudness computed from the part count alone. The gestures are the ones the repertoire uses: a texture that drops a part a bar, one that puts everything on the last chord, one that does neither, and one that thins and then arrives full. Both gestures that add parts to the last bar arrive below the running impression rather than above it — and the one that adds nothing, and takes away until a single line is left, arrives furthest below it of all.

The reading, and the two things that are backwards

The quantity to read is not the loudness at the end. It is what is sounding at the arrival, against what the listener has been hearing — the fast smoother over the slow one, which is the third rung of the loudness ladder’s whole apparatus and is what makes a subito effect an effect.

Read that way at the moment the final bar arrives:

  • doing nothing at all gives a ratio of 1.000, by construction
  • a full final chord, three parts thicker than the rest, gives 0.952
  • thinning and then arriving full, two parts thicker, gives 0.929
  • thinning to a single line gives 0.764

The first thing backwards is the one this essay is named for. The two endings that add parts arrive quieter relative to what preceded them, and they arrive quieter by about the same amount whichever way the parts are added. That is not a rounding error and it is not a bug. It is the previous rung’s finding turning up where it is least expected: adding parts to a chord does not make it louder, because the extra parts are doublings and a doubling lands in a critical band that is already occupied. Nine parts against six is 1.8 decibels of power and about a third of a decibel of loudness, and against that the fuller chord’s own crowding of the bands costs slightly more than it gains — its occupied-band count falls from 4.7 to 3.7 on the last chord.

The second is larger and it is in the other column. The gesture with the biggest effect on the reading is the one that adds nothing at all. A texture that drops a part a bar arrives at 0.764 — a quarter below what the listener has been hearing, where the loudest-looking gesture in the repertoire manages a twentieth. In the phon that the loudness ladder converts its sones into, that is 3.9 against 0.7, and it is the only number on this page large enough for anybody to hear as a change of dynamic rather than as a shading.

Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 4 parts to 9 adds 3.5 decibels of power and -0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 5.3 to 3.6 as the parts crowd into the same register.
Fig. 2 The mechanism, from the previous essay: parts against loudness, with the count of occupied critical bands beside each point. Between six parts and eight the band count falls and the loudness barely moves. Every number in the hero figure is a walk along a few steps of this axis.

Why the thinning is a cliff and not a slope

A gesture that drops a part a bar looks like a gradual thing, and the reading it produces is not gradual at all. Bar by bar through the closing six, the loudness of the sounding texture runs 16.1, 15.1, 15.5, 15.5, 16.7 and then 12.3 sones. Five parts, four, three and two are all within a phon of six; the two-part bar is the loudest of the six, louder than the full body that preceded it; and the entire fall is the last step, where two parts become one.

The reason is the same one the thickening result rests on, running the other way. Loudness here is the count of occupied critical bands more than it is the count of parts, and thinning a texture spaces it out. The ensemble in this model covers a fixed compass, so removing a part moves the survivors apart: the occupied-band count climbs from 4.8 at six parts to 7.9 at two, which is very nearly enough to pay for the three parts that left. A five-part chord loses a voice and opens a band, and the two effects meet in the middle.

Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 1 parts to 6 adds 7.8 decibels of power and 4.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 7.0 to 4.7 as the parts crowd into the same register.
Fig. 3 The same axis as the figure above, run downward from six parts to one instead of upward from four to nine. Six parts against one is 7.8 decibels of power and 4.1 of loudness — and essentially all of the loudness is the single step between one part and two, which is 4.4 phon on its own, while the four steps from two up to six wobble within a third of a phon and end below where they started. The occupied band count runs the other way: 4.7 bands at six parts against 7.9 at two, because the survivors of a thinning have room to put their partials where nothing else is.

At one part there is nothing left to open. A single line occupies the bands one note’s partials fall in and no others, and no amount of spacing can be recovered from a texture with nothing to space. So the gesture’s whole dynamic effect is concentrated in the bar where the last inner voice leaves, and it is worth 4.4 phon in that one bar — the sounding loudness falls from 16.7 sones to 12.3, which is a quarter of it gone in two seconds.

That has a consequence for how the gesture is written and it is not the obvious one. A thinning ending delivers its dynamic event at the moment it reaches one line, wherever in the gesture that is. Six bars of thinning and two bars of thinning produce the same reading if both end on a single voice, and a thinning that stops at three parts produces essentially none — the model says such an ending is, dynamically, an unchanged texture with fewer players. Which is a testable claim about a repertoire that writes both.

What actually makes a final chord loud

If not thickness, then what? The model answers, and the answer is the one the notation makes explicit rather than the one the scoring implies.

Effort. A dynamic mark is not a level but an instruction about how hard to play, and the mark that is not a level established that it brings a whole spectrum with it — a harder-driven instrument is a brighter one, and brightness is energy moved into bands higher up, where it adds rather than crowds. Twenty decibels of effort dwarfs anything a part count can do.

Contrast. The running impression takes a second or two to build and several to come down, so a chord arriving after a diminuendo is measured against a low reference and a chord arriving after a tutti is measured against a high one. The fourth line of the hero figure — thin, then full — is that gesture, and it is the one composers actually write.

That it still comes out at 0.929, barely below the full chord that has no diminuendo in front of it, says the model thinks that gesture is not doing what it looks like. Part of the reason is the time constants: six bars at two seconds each is twelve seconds, which is several times the slow smoother’s release, so the running impression has fully followed whatever the closing bars did and then had two seconds to start following the final chord up. Whatever contrast the gesture builds, the smoother spends.

The rest of the reason is worse for the gesture, and it is the subject of the next section: the diminuendo it is named for is not a diminuendo. The bar immediately before the final chord holds two parts and is the loudest bar of the whole closing passage — 16.7 sones against the six-part body’s 16.1. A thinning that is read as a loss of parts is, on this model, a gain of space, and the two nearly cancel.

That is a claim about the length of the closing gesture rather than about its shape, and it makes a prediction the repertoire can be read against. A diminuendo of twelve seconds gives the impression time to follow it all the way down and then to start back up; one of two or three seconds does not, and the final chord arrives against a reference that is still high — which is worse — while one of five or six arrives against a reference that has fallen and not yet recovered. The gesture works over a window set by the smoother’s release and not over a window set by the phrase, which is why a subito effect has to be sudden and why a long tapering close is a different device rather than a bigger version of the same one.

Final chord, and what the impression doesAn ending that is louder than everything before it, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.92 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.1.92×02468101205101520secondsloudness, sonessounding nowshort-term — theloudness of a notelong-term — theloudness of a passage
Fig. 4 The gesture drawn as a level rather than as a texture: a passage that ends with a chord much louder than everything before it, with the running impression underneath. This is what the notation asks for and it is enormous — sixteen decibels — against the third of a decibel a part count buys. The hero figure and this one are the same ending written two different ways, and only one of them works.

Where this leaves the closure vector

The ladder’s first rung established that an ending is not one thing but five components with no total: harmonic arrival, metrical placement, melodic descent, a slowing, and a silence. The fifth rung added the silence’s own bounds and named the dynamics as the missing sixth.

This rung supplies the sixth, and what is useful about it is that its size depends entirely on which way the gesture runs. Adding parts is worth about a phon however many are added and however they are arranged; taking them away until one line is left is worth four, and would be worth more if the model let a texture fall below one part. The dynamic component of an ending, read off a page’s texture, is therefore not one number but a strongly asymmetric pair — and either way it is worth twenty decibels less than a dynamic mark on the same page. So the sixth component is three quantities rather than one:

  • thickening, which is in the notes and is nearly nothing;
  • thinning, which is in the notes and takes a quarter of what a listener has been hearing;
  • dynamic marking, which is not in the notes at all and is worth twenty decibels of it.

The closure vector can carry the first. The second is a mark on the page, it is not derivable from the pitches, and every essay in this ladder that reads a scheme rather than a score has been silently discarding it — as has the repetition ladder, whose whole matrix is built from bars of pitch classes with no dynamics in them at all.

The vector as the first rung built it had components with no defensible way to total them. This rung adds one more component and finds it small, which does not make the totalling problem easier — it makes the list longer.

The gesture the model does like

Every one of the four endings comes out at or below the running impression, and none of them is a loud arrival — so it is worth asking what the model would call one, because the answer is a fifth gesture the figures do not draw and the repertoire uses constantly.

Loudness rises when the number of occupied bands rises, and the number of occupied bands rises when a texture opens — the same parts moved further apart, or a new part added well outside the register already used. A final chord scored with the bass an octave lower and the top an octave higher than anything before it is a chord that occupies bands nothing has been in, and the band model says that is worth several decibels where two extra inner parts are worth a third of one.

So the model’s advice, if a model may be said to give any, is: an ending is made loud by spacing rather than by count — and the table below says the “rather than” is not a comparison of two effects but a statement that one of them is zero. Which is what a scored tutti actually does — the bottom of the orchestra and the top of it arrive at the last chord together, and the middle was already there.

That also gives the ladder a testable difference between two things that look alike on a page, and it is worth running rather than asserting. A final chord with two extra inner parts and a final chord with two extra outer ones are the same number of noteheads and the same power per part:

the last chord occupied bands loudness against the six-part body
six parts, MIDI 40–80 4.6 ×1.00
eight parts, same span 3.9 ×0.95
eight parts, span 28–92 5.6 ×1.17
six parts, span 28–92 6.3 ×1.17

The outward chord is 1.23 times the inward one, three phon, against the thickening gesture’s own −0.7 — so the difference between the two ways of writing the same number of notes is larger than the gesture that adds them, and it runs in the opposite direction.

The last row is the one to keep. Spreading the existing six parts outward, adding nothing at all, gives the same loudness as adding two outer parts to eight — 19.01 sones against 18.96, which is a difference of a quarter of a per cent. The extra parts contribute nothing measurable; the entire gain is the register they were placed in. So the model’s advice is not that spacing matters more than count. It is that count does not matter and spacing is the whole of it.

Which sharpens what a scored tutti is doing. A final chord with the bass an octave down and the top an octave up is not loud because there are more players; it is loud because the chord occupies bands nothing has been in, and it would be equally loud with the same six players moved there. That the tutti also has more players is a fact about who is available and about spectrum, not about level.

It also makes the difference between the two writings a real test rather than a modelling artefact, because they are distinguishable by ear in a way the numbers say they should be: three phon is a clear difference in loudness and the two pages are the same count of noteheads at the same dynamic. Nothing in this collection can run that test, and it is the cheapest listening experiment any rung of this ladder has proposed.

4 tones, one power, and the interval between them. 4 tones of fixed total power, spread symmetrically about 262 hertz, drawn against the interval between neighbours. Piled on one pitch they are one sound of that power; separated by more than a critical band — 7.0 semitones here — they are 4 sounds whose loudnesses add, and the same power reaches 2.11 times the loudness at 13 semitones. The two lines are two models of the same rule and they disagree about how abrupt the change is, not about where it goes.
Fig. 5 The quantity that gesture moves: four tones of fixed total power, spread apart by a growing interval, against the loudness that produces. Nothing about the power changes across this figure and the loudness roughly doubles. An ending scored outward rather than inward is a walk to the right along this axis, and it is worth several times what adding parts in the middle is worth.

The parts-added case is worth drawing on its own axis, because the number it gives is smaller than almost anybody expects.

A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance.
Fig. 6 Parts added to a final chord one at a time, each at the same level, against the loudness that results. Going from one part to eight moves the total by 1.8 decibels — and not monotonically: four parts are louder than five and than eight.

Non-monotonic, which is the detail that kills the intuition entirely. Adding a part can make a final chord quieter, because where the new part falls decides whether it shares a critical band with something already there. A page has two decibels of dynamic range in it and a player has sixty, and this is the two.

Which computation produced the numbers

The scheme is a thirty-two bar AABA at six parts throughout, with the last six bars given one of four texture instructions. Each instruction is a rule for how many parts sound in each closing bar: dropping one a bar to six, five, four, three, two and one; adding three on the last; doing nothing; or dropping to two and then arriving on eight.

Each bar is realised the way the previous essay realises one — every complete uncrossed voicing enumerated and up to ninety-six sampled, each given a string spectrum at a fixed level per part, every partial placed in its critical band, and the loudnesses summed — and the resulting level is held for two seconds and run through the fast and slow smoothers.

A bar of one or two parts cannot state a triad, which is the one place the realisation had to be given a rule rather than an enumeration. A three-note chord in two voices is a reduction and which two notes survive it is a decision: the rule here keeps the root first and then the fifth, which is what a closing duet plays and what a single closing line plays. Keeping the root and the third instead is defensible and would draw a different figure.

The reading is taken at the arrival of the final bar rather than at the last sample, and that detail is load-bearing. The slow smoother’s release is two seconds and a bar is two seconds, so by the end of the last bar the running impression has caught up with whatever is sounding and every ending scores a ratio of exactly one. A figure read there would say all four endings are identical. What the rung is about is what a listener has when the ending arrives, which is a fifth of a bar in.

The control is the same scheme with no gesture at all, which reads 1.000 by construction and is what the other three are differences from.

And the dynamic curve stops before the ending does. The silence after the last chord has to reach a length of its own before it is heard as an ending rather than as a gap, which is a duration nobody writes down and which the fifth rung of the form ladder priced separately. An ending is therefore two gestures, and only one of them is in the notation.

Where the model stops

The one thing that matters most is not in it. Every part is played at a fixed level in every bar, so the model has no dynamic marks. That is not a limitation of the machinery — the loudness ladder’s third rung has levels and time constants and the shapes to drive them — it is that a scheme has no dynamic marks in it, and reading them would need a score rather than a chord chart.

The tempo is asserted. Two seconds a bar is a moderate tempo and every conclusion about the smoothers depends on it, because a smoother with a two-second release meets a two-second bar and the two are comparable. At a fast tempo the closing bars go by inside one release and the contrast survives; at a slow one it is spent. That is the most consequential parameter here and it is not varied, which is a real shortfall rather than a caveat: the same four endings at three different tempi would very likely rank differently.

The four gestures are inventions. They are plausible textures written as rules — drop one a bar, add two at the end — rather than textures read out of any score. A corpus of real closing bars with their part counts would be a small object and would very likely contain gestures none of these four resembles, particularly the ones that change register rather than count.

A thinning here is also a re-spacing, and that is a choice. The ensemble covers one fixed compass however many parts it has, so two voices are an octave and a half apart where six are a third apart. That is right for the question the ensemble family was built to answer and it is not what a score does when it thins: a real thinning stops players and leaves the others where they were sitting. The cliff above is a result about a thinning that spreads, and a thinning that merely subtracts would fall earlier and further. Which of the two the repertoire writes is a question about scores rather than about arithmetic, and the difference between them is larger than three of the four gestures on this page.

A scheme is not a piece. These are Roman numerals with a texture attached, held for a bar each, with no melody, no articulation, no rest and no decay. A real ending has a final chord that is struck and then decays, and a decaying chord against a running impression is a completely different shape from a sustained one.

And the three other closure components are not in the same units. This rung produces a ratio; the harmonic component produces a count of resolved tendencies; the ritardando produces a curve parameter. Putting them together is the commensuration problem this collection keeps meeting and has never solved.

Whose endings, and when

The four gestures are drawn from what the repertoire does, and they are not equally distributed across it.

Thinning to nothing is the characteristic ending of a great deal of nineteenth-century song and of most twentieth-century popular music — the texture drops away, the last thing sounding is a single line, and the piece stops. The model says this is the loudest gesture on the page in the only sense that matters to a listener, and says so for a reason that fits how those endings are described: the event is not the thinning, it is the arrival at one line, and everything before it is preparation that costs nothing. That is the same shape as a deceleration, which is not experienced as a slowing so much as an arrival.

The full final chord is baroque and classical, and it is almost always written with a dynamic mark. A Handel cadence with everybody playing is not loud because there are more players; it is loud because it is marked forte and every player is working. The scoring and the marking arrive together and the model here says only the second is doing the work.

Thin then full is the romantic apotheosis and it is the one this figure is most confident is not doing what it looks like — it stops thinning one bar short of the only bar where thinning is worth anything, so the reference the final chord is measured against is the loudest bar in the passage rather than the quietest.

What the picture cannot show

Whether a listener’s running impression is this one. The two time constants come from a published time-varying loudness model, they were fitted to laboratory stimuli, and applying them to a musical form is an extrapolation across four orders of magnitude in duration. A listener at the end of a movement has a sense of how loud the movement has been that no one-pole filter with a two-second release is going to capture.

And it cannot show the silence. The fifth rung found that a silence becomes an ending somewhere between three and a half and five and a half seconds, with a hall’s reverberation eating the first two. The dynamic curve computed here stops when the notes stop, and the most important part of an ending’s loudness contour is what happens in the seconds after that — which is the room’s, not the score’s.

Where this ladder goes next

Six rungs. An ending is five components with no total; half the cadences in the repertoire withhold some of them; the performance slows on a curve with one free parameter; nothing in a piece’s own statistics announces that it is about to stop; a silence is an ending after three and a half seconds; and now the dynamic component, which turns out to be strongly asymmetric — about a phon for any way of adding parts, four for the one bar in which a texture reaches a single line.

What is owed is the tempo. Every number here rests on a bar being two seconds against a smoother whose release is two seconds, and those two being comparable is the whole reason a gesture’s contrast gets spent. A sweep over tempo would say at what speed a diminuendo into a final chord starts working, it needs nothing this ladder does not have, and it is the kind of parameter that has embarrassed this collection before by turning out to decide the answer.

And a second thing is owed, which the cliff put there. Every thinning drawn here re-spaces the voices that remain, so the loss of a part and the opening of a band cancel until there is nothing left to open. A thinning that keeps the surviving voices where they were sitting is the other half of the same instruction, it is the same computation with the ranges taken from the full texture instead of from the thinned one, and the prediction is sharp: the fall would begin at the first bar rather than the last, and the running impression would have the whole gesture in which to follow it down. If that is right, then the two ways of scoring the same instruction differ by more than any two of the four gestures here differ from each other, and the one this page draws is the one that keeps its contrast.

Part 6 of 13

One essay in the series on closure. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CadenceClosureDynamicsExpectationFormLoudnessOrchestrationTexture