Perception and the listener

A page has two decibels

The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.

Assumes: The dynamics are in the score already · A chord is not as loud as its notes

The dynamics are in the score already computed a passage’s loudness from its texture alone — how many parts and in what registers — with no performance in it at all. It ended by naming the rung immediately available:

Every ending in every repertoire is a dynamic event, and the closure ladder had no model of dynamics in a form. It has one now, and applying it is the next essay rather than a further rung of this one.

The closure ladder has spent eight rungs on what makes an ending an ending, and one of its five components is the dynamic. It has been calling that component a corpus debt.

This essay is that model applied, and the answer is that the debt was a category error.

A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance.
Fig. 1 Parts added to a final chord one at a time, each at the same level, against the same chord played by people from forty to a hundred decibels.

Eight parts and one, and almost nothing between them

Take a final chord and build it up: a single note, then two, then three, out to eight, each new part added at the same level as the others and never taking one away. That is the nested version of the question, so what is being compared is a texture growing rather than eight different chords.

The loudness goes 12.3, 17.1, 15.4, 18.2, 12.4, 13.8, 10.7, 10.9 sones, for one part up to eight.

That is a span of 1.8 decibels from end to end, and it is not monotone: four parts are louder than five, and than eight. A naive power sum — eight sources at one level, adding as energy — predicts nine decibels of growth.

A score can produce about two decibels of dynamic by changing how many parts are sounding. A player has sixty-two.

Why, and both reasons are already on this ladder

Neither half of that is new here; what is new is putting them together on an ending.

Notes in one critical band do not add. A chord is not as loud as its notes is the second rung, and it is the reason a triad is not three times a note: the ear sums specific loudness across critical bands and compresses within them, so two parts a third apart are barely louder than one.

And a part’s contribution depends on where it is. The equal-loudness contours make a note at 55 hertz far quieter than one at 550 at the same sound pressure level, so a texture that fills out downward adds almost nothing to the total while filling out upward adds more. That is why the curve above is not monotone: the fifth part is the one that goes into the bass.

Put together they say that the number of parts is nearly orthogonal to the loudness, which is a strange thing to be true of an orchestra and is what the first two rungs of this ladder have been saying separately.

90 players, and how much louder than one. Independent players on one note add in power, so 90 of them are 19.5 decibels above one — and loudness in sones goes as roughly the 0.3 power of intensity, so the section is 3.9 times as loud rather than 90 times. The horizontal axis is doublings of the section, which is why the line is nearly straight: every doubling is 3 dB and about a quarter more loudness.
Fig. 2 The same effect at the scale of a section rather than a chord: ninety unison players are not ninety times one.

The naive sum is the number everybody has in their head

It is worth saying what the faint line on the figure is, because it is the intuition this rung is against.

Eight independent sources at one level, adding as energy, are nine decibels louder than one. That is the arithmetic of a microphone and it is what “eight times as many players” sounds like it should mean. It is also roughly right for eight players on different notes spread across the spectrum, which is why the intuition survives.

What it is wrong about is a chord. A chord puts its parts into a handful of critical bands on purpose — that is what a chord is — and within a band the ear takes something much closer to the largest component than to the sum. So the very thing that makes a chord a chord is the thing that stops its parts adding.

The closer the voicing, the less each new part is worth, and a final chord in a tutti is not usually a wide one.

Which refuses the debt rather than paying it

The closure ladder’s position was that it could not price its dynamic component because it had no account of dynamics that did not require a performance, and that a corpus of scores would supply one.

It would not. A corpus of scores contains markings and textures, and neither of them is a loudness. A marking is an instruction to a player, which this collection has already argued is an ordinal scale laid over a continuous one; a texture is worth two decibels.

So the dynamic component of an ending is a performance quantity, entirely, and the reason the closure ladder could not compute it is not that it lacked data. It is that the thing it wanted is not in a score.

That matters for how the debt should have been read. A corpus debt says the argument exists and the material to run it is missing. This one says the argument was pointed at the wrong object, and the way to find that out was to build the model and watch it return two decibels.

Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness.
Fig. 3 The dynamic component of closure, which is what this essay was asked to supply from the page. Three of the four gestures sit inside the two decibels a part count can buy, and the fourth does not — because it does not stop at a part count.

The one gesture that escapes, and why it is not a counterexample

One of the four closing textures on that figure arrives 3.9 phon below the listener’s running impression, which is twice what this essay has just said a page can produce. It is worth saying exactly where the extra comes from, because it looks like a contradiction and is the same finding read from the other side.

That gesture ends on a single line. Every other reading here compares chords — six parts against eight, four against five — and a chord of any size puts its notes into a handful of bands on purpose. One note does not. Its partials are spread by the harmonic series across seven critical bands where a six-part chord’s crowd into fewer than five, so a single line is not simply a smaller chord and the axis from one part to two is worth more than the axis from two to nine put together.

Which sharpens the claim rather than denting it. A part count is worth two decibels; the difference between a chord and a line is worth four, and the second is not a count at all. It is the same statement as the ordering result below — what a part is worth depends entirely on which bands it lands in — arriving at the one texture where the answer is large.

What the page can do, which is not nothing

Two decibels is not zero and the non-monotonicity is not noise, and both are useful once they are read as what they are.

A texture change is a change of colour at constant loudness. Thinning from eight parts to four leaves the loudness where it was and removes most of the spectrum’s density: fewer partials in each critical band, a narrower spread, a different count of what survives masking. A listener hears that as a change and would not call it a diminuendo.

And the ordering is a compositional control. The part that adds most is the one that goes into an unoccupied critical band in the region the ear is most sensitive — around two to four kilohertz — and the part that adds least is the one doubling something already there or going into the bass. That is why a piccolo entering is an event and a third bassoon is not, and it is a fact about the ear rather than about the instruments.

And it is why orchestration is not addition. An orchestrator asked for more sound does not add parts; they change register, spacing and instrument. The two-decibel figure is why, and the same ladder’s own joint scoring problem is the machinery that does the changing — which takes levels as a continuous variable precisely because the discrete one is worth so little.

There is a corollary about rehearsal. If a passage is not loud enough, adding players to it is close to useless and asking the players for more level works. That is what conductors do, and it has usually been explained by saying an ensemble is more than the sum of its parts. On this arithmetic it is a great deal less than the sum of its parts, and the conclusion is the same for the opposite reason.

The same power, divided among more of the spectrum. One fixed total power — 80 decibels — delivered as 1, 2, … 12 tones, each a band and a half from the next so that none of them shares a critical band with another. Every extra tone takes power away from the others and the sonority gets louder anyway: 12 tones are 5.3 times the loudness of one carrying all of it. The exponent is about 0.7, which is one minus the 0.3 that relates loudness to intensity.
Fig. 4 The variable that does move the loudness: how much of the spectrum the same power is divided among. Parts spread across critical bands are much louder than parts inside one.

What a corpus would and would not have supplied

It is worth being precise about the refusal, because “a corpus debt was a category error” is a strong thing to say about a rung on another ladder.

A corpus of scored endings would have supplied three things. How many parts a final chord typically has; how those parts are usually spaced; and what markings sit over them. The first two are the inputs to the model here and they are what makes the answer two decibels rather than one or three.

What no corpus supplies is the level. A score has no decibels in it anywhere. A marking is a request, and what a player produces in response to it varies by instrument, by period, by hall and by the passage before it — which is the whole of the marking rung’s argument and is why this collection treats a dynamic as ordinal.

So the closure ladder’s debt could have been discharged in part and never in whole, and the part it could have had is the part this essay computed without a corpus at all. The missing ingredient was never in the repertoire; it was in the performance, and the only corpus that would have supplied it is a corpus of recordings.

That is a different and much larger thing to ask for, and it is worth the closure ladder recording it as such.

The ending’s own case, and the arithmetic of a full close

The place this is most surprising is the one the closure ladder cares about, which is a full close in a tutti.

A cadential arrival typically adds parts — the texture fills, the bass enters, the whole ensemble plays. Every account of why a final chord sounds final mentions it. This model says that operation is worth about a decibel, and that if the players do not also play louder the arrival is a change of colour and not of level.

Which is exactly what performance practice supplies. A final chord is not made loud by adding to it is the closure ladder’s own way of saying so, and this rung’s contribution is a number for it: the addition supplies one decibel of the twelve or fifteen a forte arrival actually has.

The score writes the chord and the players write the ending, and the split is not a matter of interpretation. It is a property of how loudness sums.

The same chord is louder in the treble, at the same power. A 3-note chord spanning 7 semitones, drawn at 5 registers with the total power held constant, and scored against one note carrying all of it. Low down the whole chord fits inside one critical band — 35 semitones wide at C2 — so it is analysed as one thing and buys 0.99 times the loudness. At C6 the band is 2.8 semitones, the notes are in separate bands, and the same power is 2.24 times as loud.
Fig. 5 The same chord at several registers, at one power. Where the parts sit is worth several times what how many of them there are is worth.
Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 3 parts to 8 adds 4.3 decibels of power and 0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 6.0 to 3.9 as the parts crowd into the same register.
Fig. 6 The earlier curve over a whole form: adding parts adds power and very little loudness, bar by bar. This essay is the ending of that curve, priced.

Those bar-by-bar totals are the same summation as the chord above it, run over a whole form rather than over one arrival, and they have the same shape for the same reason. What follows is why the summation behaves that way at all.

3 notes, 6 critical bands. The excitation 3 notes of a string spectrum cast along the Bark axis, and the critical-band groups their components fall into. Components inside one group add their intensities and are converted to loudness once; the groups' loudnesses then add. This sonority occupies 6 groups and comes to 48.6 sones, against 121.0 for the same components counted one at a time.
Fig. 7 Why the sum behaves as it does: three notes and the six critical bands their partials fall into. The total is what the bands hold, and adding a part inside an occupied band adds very little to it.

The two decibels are at the limit of noticing, which is the useful way to say it

A difference limen for loudness is about a decibel. So the entire range a texture commands — from one part to eight, at a fixed level — is one or two just-noticeable steps.

That is the sentence to carry away, because decibels are easy to misjudge and a just-noticeable difference is not. A composer who thins from eight parts to four has changed something a listener will certainly notice, and it is not the loudness; the loudness has moved by about as much as the difference between two notes that are as nearly equally loud as a listener can tell.

It also gives the comparison its right size. The player’s sixty decibels is sixty limens. A page holds one and a performance holds sixty, and the ratio is not a fact about music at all — it is the shape of the loudness function.

Which computation produced the numbers

The voicings are nested: one part is middle C, and each next one adds a note without removing any, filling upward and then downward in the order an ensemble does. That is a choice and it is stated, because a non-nested comparison would be a comparison of eight chords rather than of a texture growing.

Every part is a string-like spectrum at 62 decibels sound pressure level. The total loudness is this ladder’s own summation: specific loudness per critical band from the equal-loudness contours, summed across bands and compressed within them, which is the second rung’s machinery unchanged.

The conversion from sones to decibels uses the standard doubling — a doubling of sones is ten phons, and near this level a phon is a decibel — so a span of 1.14 in sones is 1.8 decibels.

The player’s range is the same four-part chord at levels from forty to a hundred, which is a generous estimate of what an ensemble can produce and is bounded below by the room rather than by the players.

Where the model stops

The parts are all one timbre. A real ending adds a trombone rather than a fifth violin, and a different spectrum contributes to different bands. That would raise the two decibels and not by very much, because the constraint is the band structure rather than the source.

And they are all at one level. An entering part in a real tutti is not playing at the same level as the ones already there; it is playing forte because the passage is. That is a performance change and is exactly the thing this essay is separating out.

The contours are for steady tones. A final chord is a decaying event in a hall, and loudness summation over a decaying complex is a running quantity rather than a static one.

And two decibels is at the edge of audibility. The difference limen for loudness is around one decibel, so the whole range a texture commands is one or two just-noticeable steps. That is the cleanest statement of the result, it is the reason the non-monotonicity does not matter, and it means the model would have to be wrong by a factor of ten before the conclusion changed.

What the picture cannot show

It cannot show what a texture change does to anything else. Roughness, brightness and the sense of a full ensemble all move a great deal when parts are added, and none of them is loudness.

Nor can it show what an ending is for. The dynamic component is one of five on the closure ladder, and the four this essay has not touched — the cadence, the deceleration, the silence and the metrical position — are all things a score does contain.

Nor can it show the marking. A score with fortissimo over the final chord is not silent about dynamics; it is delegating them. The model here is about what the notes alone determine.

It cannot show the room. A tutti chord excites a hall more than a solo line does, and the reverberant field adds to the direct sound in a way that is closer to a power sum than to a loudness sum.

It cannot show a decaying instrument. A piano’s final chord is loudest at its onset and inaudible eight seconds later, and every number here is a steady state. The texture question and the decay question are separable and neither has been asked of the other.

And it cannot show the ending. Whether a chord sounds final is five components and no total, and this essay has priced one of them and found it small.

Whose music, and when

The chord is a generic triad and the levels are generic. Nothing here is a measurement of a repertoire.

The observation with a period in it is about markings. Scores before about 1750 carry almost no dynamic markings and scores after 1850 carry a great many; the texture, meanwhile, is written down in both. If a texture is worth two decibels then the earlier repertoire is not less dynamic — it is dynamic in a way that is not written down at all, and the growth of markings is the growth of a notation for something that was always being done.

That is a familiar claim and this figure is a small argument for it: the alternative reading, that the dynamics were in the scoring, does not survive the arithmetic.

Where this ladder goes next

Five rungs. What one tone’s loudness is; what more than one tone’s is, which the critical band decides; what a passage’s is, which two constants decide; what a page’s is, which needs no performance; and now what a page’s is worth, which is two decibels.

What is owed after this is the entry. Every rung above adds parts that are already sounding by the time the loudness is measured, and the moment an orchestrator actually controls is the entry — a part beginning, into a texture that was already there. That is a loudness change with a rise time in it, and this collection has both halves: the running impression with its ninety-nine-millisecond attack, and the onset ladder’s account of how quickly each instrument gets to its level. What comes out is whether an entering part is heard as an increase in loudness or as a new object, which is a question the auditory-scene ladder would recognise and this one has the units for.

Part 5 of 8

One essay in the series on loudness. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

ClosureCritical bandwidthDynamicsLoudnessNotationOrchestrationTexture