The dynamics are in the score already
Assumes: Loud is relative, and it comes down slowly · A chord is not as loud as its notes
The third rung of this ladder built a model of loudness in time — two asymmetric smoothers, one fast and one slow, so that what a listener has is both what is sounding now and a running impression that took a second or two to build and takes several to come down. Its last paragraph named what the ladder could then do and had not:
A texture that thins is quieter for two reasons at once — fewer sources, and fewer bands — and the band model of the second rung says the second is much the larger. So the dynamic curve of a piece with a changing texture is computable from a score without any performance data at all.
Both halves are on this site. A chord is not as loud as its notes put every partial of every note into its critical band and summed; loud is relative put the result through time. Nothing else is needed.
Why a fuller texture is barely louder
The reason is the second rung’s, and it is worth restating because it is the least intuitive thing in this field.
Loudness is not summed over components. It is summed over critical bands: components inside one band compete for the same place on the basilar membrane and add roughly as power, which is to say hardly at all; components in different bands each contribute their own loudness and add nearly linearly. So the loudness of a chord depends on how many bands it occupies, and the number of bands is set by its spacing rather than by its note count.
Now put more parts into a fixed register. There are only so many bands between the bottom and the top of an orchestral texture, and a chord has only so many pitch classes — so a fifth part, a sixth and a seventh are doublings, at the octave or the unison, and a doubling lands in a band that is already occupied.
The figure says the count of occupied bands actually falls as parts are added, from 6.0 at three parts to 3.9 at eight. More voices, fewer bands. The extra voices are not adding places on the membrane, they are crowding into the places already used.
Which is a consequence of the register being held, and the register is not something an orchestrator holds. Every figure here gives the parts ranges spanning MIDI 40 to 80 whatever their number, so an eight-part texture is eight voices squeezed into the forty semitones three voices had — the density rises because the span was fixed. Letting the span widen with the part count instead, which is what happens when an orchestrator adds a piccolo and a contrabassoon rather than a fifth viola:
| parts | bands, fixed span | bands, widening | loudness, widening |
|---|---|---|---|
| 3 | 6.0 | 6.0 | ×1.00 |
| 5 | 4.9 | 7.3 | ×1.40 |
| 8 | 3.9 | 7.5 | ×1.74 |
The band count goes the other way and the loudness goes with it. Three parts to eight is a factor of 1.01 in loudness with the span held — a hundredth, which is nothing — and a factor of 1.74, about eight phon, with the span allowed to grow. So the essay’s headline is not a fact about adding parts. It is a fact about adding parts into a fixed register, and the two are different orchestral decisions with opposite consequences.
That widening model is illustrative — six extra semitones of span per added part is a choice, not a measurement — and the direction is what matters rather than the size. What the two columns establish between them is that the whole of the effect is register and none of it is the part count: eight voices in forty semitones occupy fewer bands than three, and eight voices in seventy occupy more, and the loudness follows the bands in both directions exactly as the second rung’s model says it must.
So the finding survives in a narrower and more useful form. Doubling the parts without widening the register buys about a decibel, which is a real and counterintuitive statement about a real orchestral option — the one where a section is thickened rather than the orchestra extended. And the reason an orchestral crescendo built by adding instruments works at all is that the instruments added are not in the register already sounding: a tutti is not a denser chord, it is a wider one.
What the curve looks like on a real scheme
Applying it bar by bar to a thirty-two bar song with a texture that thins in the bridge gives a dynamic reading of the piece, and the reading has a shape that no count of parts would produce.
The A sections at five parts and the B section at three differ by two decibels of power. What the band model gives is under one, and what the running impression gives is less still — because the third rung’s slow smoother has a two-second release, and a bridge eight bars long at a moderate tempo is sixteen seconds, which is long enough for the impression to follow the texture down and sit there.
So the computed curve is not a step. It is a shallow dip that arrives late and recovers late, and its depth is a fraction of what the arithmetic of sources alone predicts.
Which is a claim about what a composer is doing when they thin a texture, and the claim is that they are not making it quieter. They are making it thinner — changing which bands are occupied and therefore the timbre of the whole — while leaving the loudness nearly alone. Those are separable and the notation has one word for both.
The claim holds only for a thinning that keeps the outer voices, and that is the ordinary case: a bridge scored for three parts of a five-part texture normally drops inner voices rather than the bass or the top line, so the span survives and the density falls. A thinning that also narrows the register — dropping the bass, or the top — is a different event and the arithmetic says it is quieter, because it takes bands away rather than merely emptying them. So the notation’s one word covers two operations, and it is which voices go rather than how many that decides which one has happened.
The realisation problem, and how it was solved
There is a decision inside this computation that had to be made twice, and the first answer was wrong in a way worth recording.
A bar of a scheme is a Roman numeral, which is a set of pitch classes. To get frequencies out of it, the parts have to be given actual notes — a voicing — and a triad in five parts has 667 complete uncrossed voicings within the ranges. Which one?
The first version took the middle one of the enumeration, and the curve it produced was noise: a three-part bar came out louder than a four-part one because the particular voicing the index landed on happened to put two of its notes in one band. That is a fact about an arrangement and the rung is about a texture.
So the answer is to average over the realisations, sampled evenly across the enumeration, which is ordered by register — so the sample is a sample of registers rather than of one corner of the space. The curve is then smooth, monotone where it should be, and is a property of the number of parts rather than of a choice nobody made.
That is a small thing and it is the difference between a figure and a plausible-looking artefact.
What this does to the two marks a page carries
There is a consequence for notation, and it sharpens something the dynamic-mark rung found.
A page carries two independent instructions about loudness. One is the dynamic mark, which that rung established is not a level at all but an instruction about effort — and which brings a timbre along with it, because a harder-driven instrument is a brighter one. The other is the number of staves that have notes in them, which is not usually thought of as a dynamic instruction and which this figure says is the weaker of the two by a wide margin.
So the two things a score says about loudness are of very unequal strength, and the one that looks quantitative is not. Eight parts against three is a decibel; pp against ff is twenty or more. A texture change is a colour change with a loudness change attached as an afterthought, and a dynamic mark is the reverse.
That is worth stating carefully, because the usual account runs the other way round and the section above says why it is not simply wrong. Orchestration treatises talk about the number of instruments as a means of building a crescendo, and this arithmetic says that adding instruments inside the register already sounding barely does it — what such an addition contributes is spectrum rather than level. Adding instruments that extend the register does raise the loudness, by rather more than a decibel, so the treatises are right about the practice and this figure is right about the special case where the register is held.
Which computation produced the numbers
The scheme is one of this site’s own — a thirty-two bar AABA — and the texture is a part count per bar, which is the one thing that comes from the page. Everything else is machinery that already existed.
Each bar’s Roman numeral gives pitch classes; the parts are given ranges spanning MIDI 40 to 80 whatever their number, so that the span is held and only the density varies — which is the one modelling choice in the chain that decides the headline, and the section on the band count says by how much; every complete uncrossed voicing is enumerated and up to ninety-six are sampled evenly. Each voicing’s notes are given a string spectrum at a fixed level per part, every partial is placed in its critical band by the second rung’s grouping, and the loudness in sones is the sum over bands.
Sones are converted back to a level so that the time-domain smoothers see the units they were built for, and the result is sampled at two seconds a bar and run through the fast and slow smoothers with their published time constants.
Nothing in that chain is fitted to anything. The band grouping is the second rung’s, the time constants are the third’s, the voicing enumeration is the part-writing ladder’s, and the only free parameter is the level each part is played at, which is common to every bar and cancels out of every comparison.
Where the model stops
Every part is the same instrument at the same level. That is the largest simplification here by a long way. A real orchestral texture thins by removing the loudest parts or the highest ones, and which parts leave changes the band occupancy in ways a count cannot see. Which player on which note is the ladder about that, and joining it to this one is the obvious next thing and is a substantial piece of work.
The span is held and a real texture’s is not. Thinning usually narrows the range as well as reducing the count, and narrowing the range reduces the bands directly. So a real thinning should be more audible than this model says, and by how much depends on what leaves.
There is no masking between parts. The band grouping treats components in one band as adding in power, which is a crude stand-in for what actually happens — a strong component raises the threshold for a weaker one nearby, which is an upward spreading skirt rather than a box. The second rung’s own alternative, an excitation-pattern integration, exists on this site and disagrees with the band grouping by a few per cent on a single tone.
The span held fixed is a choice with a consequence. Parts are given ranges spanning MIDI 40 to 80 whatever their number, so eight parts are eight voices packed into the span three would have used. That is the right comparison for the question asked — what does adding parts do? — and it is not what an orchestrator does, because an orchestrator adding parts usually also opens the register. Holding the span is what makes the band count fall; letting it grow would make it rise, and the answer would come out closer to the power sum.
And a score is not a performance in a much deeper sense than this figure admits. Every bar here is played at one level for its whole duration, with no attack, no decay and no articulation. What is computed is the loudness of a sequence of sustained chords with the harmonic rhythm of a song, which is a thing no performance of that song resembles.
Whose music, and when
The reading is a reading of a texture, so it applies to any repertoire that writes one down. What it says most about is the one that thinks in part counts.
Four-part writing — the chorale, the string quartet, the SATB choir — is a texture with a fixed count, and this figure says the fixed count is doing much less to the loudness than a musician’s vocabulary suggests. What the eighteenth-century vocabulary for texture actually tracks is spacing: close position and open position, the gaps between the upper voices, whether the bass is far below. Those are the variables that move the band occupancy, and they are the ones the pedagogy is careful about.
The nineteenth-century orchestral crescendo is a different object and the model is much less confident about it, because it is almost always a simultaneous thickening and loudening — more players, playing louder, in a wider register — and this figure varies only the first. Its whole claim about that gesture is negative: the part count is not where the loudness is coming from.
What the picture cannot show
What thinner sounds like. The model produces a number of sones and the interesting thing about a thinned texture is not its loudness but its colour, which is a statement about which bands are occupied rather than about how many. The occupancy is computed here and is thrown away in the summing.
Nor whether any of this is what a composer thinks they are doing. The figures compute what a texture does to a loudness; they say nothing about intention, and the honest position is that a composer thinning a texture is almost certainly after the colour rather than the level, and has been getting the colour with the level attached without having to think about which was which.
And it cannot show what a listener is attending to. A texture of five parts in which one is a melody is not five equal contributors; a listener is following one line and the others are a background, and the loudness of the thing being attended to is a different quantity from the loudness of the whole. Nothing in this collection has a model of that at all.
Where this ladder goes next
Four rungs. What one tone’s loudness is; what more than one tone’s is, which the critical band decides; what a passage’s is, which two constants decide; and now what a page’s is, which needs no performance at all.
The rung immediately available is the one the closure ladder has been calling a corpus debt — every ending in every repertoire is a dynamic event, and the closure ladder had no model of dynamics in a form. It has one now, and applying it is the next essay rather than a further rung of this one.
What this ladder owes after that is the thing the model’s largest simplification names. Every part here is one instrument at one level, and the whole of orchestration is that they are not. The joint problem — which instrument on which note, at which dynamic, to produce which loudness — is a discrete choice and a continuous one at once, the spectrum ladder has the discrete half, this rung has the continuous half, and neither has ever been handed to the other.
Part 4 of 8
One essay in the series on loudness. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 9.
- A final chord is not made loud by adding to it
- A general pause is spent by the note after it
- A part that leaves is not a part that arrives
- The ranking survives the dynamic and the chord does not
- The release is on the wrong side
- Who plays what and how loud is one question
- A staccato is a dynamic mark
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Critical bandwidthDynamicsLoudnessMaskingOrchestrationSoneTextureVoicing
- A part entering is not a change of level critical bandwidth, dynamics, loudness, orchestration, texture
- A bass chord low enough to balance has already hidden its tenor critical bandwidth, loudness, masking, voicing
- A loud chord is a smaller chord critical bandwidth, dynamics, masking, voicing
- A rough arrival is rough because of its spacing critical bandwidth, dynamics, loudness, voicing
- A soft chord has to fade in dynamics, loudness, masking, orchestration
- A subito piano is a rate, not a level dynamics, loudness, orchestration, texture