A subito piano is a rate, not a level
Assumes: A dynamic mark changes what a note is · Loud is relative, and it comes down slowly
Every scoring in this anchor holds one chord still. Who plays what and how loud is one question solves an assignment and a set of levels together, at an instant; the ranking survives the dynamic varies the level and keeps the instant; a dynamic mark changes what a note is lets the spectra move and keeps it still.
Its last paragraph named the variable none of the three had touched:
Every scoring here is one chord held steadily, and an orchestration is a succession — so a balance that is correct at an instant is not necessarily correct in a passage, and a part that enters is not the same as a part that was already there.
Adding time turns the objective into a functional and produces a result about notation rather than about balance.
A written dynamic is an instruction to the impression
The premise of the rung is a reading of what a dynamic marking is for, and it is worth arguing rather than assuming.
Nobody writing a crescendo means “let the sound pressure be this at every instant”. A marking is an instruction about what the passage should seem to do, and seeming is a property of the listener’s running impression rather than of the instantaneous sound. Loud is relative, and it comes down slowly is this collection’s account of that impression: two smoothers, of which the long-term one has an attack of a tenth of a second and a release of two seconds.
Two seconds is a long time in music. A chord lasting one second and a chord lasting four are on opposite sides of it, and the same written marking is a different instruction in each case.
That is the whole of this rung. Solve for the levels that make the running impression arrive at the written dynamic, rather than the levels that make the instantaneous loudness equal it, and the two answers differ by an amount that depends on how fast the music moves.
Where the two policies agree, and where one fails
The passage used is a five-chord phrase with a written shape: a moderate opening, a growth over three chords into a six-part fortissimo, and then a two-part subito piano.
For the first four chords the two policies agree exactly. The written shape grows steadily, and the running impression follows a growth without difficulty — the long-term smoother’s attack is a tenth of a second, so an increase is tracked almost immediately. Playing each chord at its written loudness gives the impression the written value.
The fifth chord fails. Its written value is 2.0 sones and the impression, playing literally, sits at 4.02 — twice what was asked for. The previous chord was 5.6 sones and the release is two seconds and the chord lasts 1.2, so most of the previous chord is still in the listener when the new one arrives.
And the correction does not exist. Solving for the level that would bring the impression down to 2.0 drives the players toward silence, and at silence the impression is still 3.14 — more than half again the marking, with nothing sounding at all.
So the fifth chord’s dynamic is not producible. Not difficult; not a matter of players not being quiet enough; not producible, because the quantity being asked for is a property of the listener’s memory and the players have no access to it.
Which is to say a subito piano is a rate
That is a strange conclusion and it has a natural restatement.
A subito piano is not a level. It is a rate of change, and the rate is what the smoother refuses. The listener’s impression can fall at a certain speed — an exponential with a two-second time constant — and a marking asking for a fall faster than that is asking for something the ear does not do.
What a listener actually gets from a subito piano is therefore not the written dynamic. It is the largest fall the release allows, plus the contrast between what is sounding and what is remembered — which is precisely the ratio the closure ladder measures at an ending, and which is the quantity a sudden change is actually about.
So the marking works, and it works by a different mechanism from the one it names. It does not make the passage quiet; it makes the passage quiet relative to a memory that has not caught up, and that relative quantity is much larger than the absolute one.
How fast a passage has to move
Sweeping the pace says where the effect begins and it is not at an exotic tempo.
At 4.5 seconds a chord — a slow harmonic rhythm, one chord every two bars at a moderate tempo — the impression misses the written dynamic by 0.39 sones on the worst chord, which is small.
At 2 seconds, one release, it misses by 1.35.
At 1.2 seconds, which is a chord a bar at a moderate tempo, it misses by 2.02.
At 0.3 seconds, a chord a beat, by 3.09.
The transition is smooth and it is centred on the release, as it must be. What matters musically is that the range where the effect is large — one to two seconds a chord — is exactly the range most music lives in. How often the chord changes puts the harmonic rhythm of most tonal repertoire between one and four chords a bar, which at ordinary tempi is between half a second and two seconds a chord.
So the discrepancy is not a corner case. It is the normal condition of music that changes its dynamics at the rate music changes its dynamics.
The asymmetry, which is the mechanism
The reason growth is free and collapse is not is the single most important number in the loudness ladder and it is easy to walk past.
The long-term smoother’s attack is 99 milliseconds and its release is 2 seconds. They differ by a factor of twenty.
So an increase in level is followed almost at once — a crescendo that takes half a second is tracked by the impression with a tenth of a second’s lag — and a decrease is followed slowly. The impression rises with the sound and comes down on its own schedule.
That asymmetry is why the first four chords of the passage are free and the fifth is not, and it is why the effect this rung finds has a direction. Every marking asking for a fall is compromised and no marking asking for a rise is. A subito forte is producible and a subito piano is not; a crescendo is heard as written and a diminuendo is heard as slower than written.
There is a musical reading of that which is worth putting plainly. A composer writing a fast diminuendo is writing something the listener smooths, and a composer writing a fast crescendo is not — so the two are not mirror images, and any account of dynamics that treats them symmetrically is wrong in a direction this figure can price.
It also predicts an asymmetry in notation practice: markings for sudden loudness should be common and markings for sudden quiet should need help. That is roughly what scores show — sforzando and its relatives are a whole family of marks with no quiet counterpart, and the subito piano is usually written with an explicit word rather than a symbol.
What this does to the anchor’s own objective
The first three rungs pose a constrained optimisation: hold each part’s share of the loudness at a target, hold the whole scoring at one total loudness, and minimise the roughness. Adding time changes what the constraint means.
At an instant, the loudness constraint is a statement about the sound.
Over time, it has to be a statement about the impression, because that is what a marking is an instruction about — and the impression at a moment is a weighted history rather than a value.
So the objective becomes a functional: a scoring is not a set of levels but a set of level trajectories, judged by what the impression does across the passage rather than by what the sound does at each point.
That is a much larger optimisation and this rung does not solve it. What it does is show that the two formulations differ, by how much, and where — and identify the one place where the difference is not a refinement but an impossibility.
There is a practical corollary for the earlier rungs. A balance solved at an instant is correct for a chord that has been sounding for several seconds and is not correct for a chord that has just arrived after a louder one. A part entering into a loud texture has to be louder than the same part entering into silence, to produce the same impression, and how much louder is the same exponential.
What a player would have to do instead
The correction the figure computes is unavailable, and the corrections that are available are worth listing because they are what performers actually do.
Take time before the change. Two seconds of anything — a caesura, a lengthened final note of the phrase, a breath — lets the impression fall, and the new chord then arrives against a reference that has come down. This is the commonest solution and it is a change to the tempo rather than to the dynamic.
Make the previous chord shorter. A loud chord that lasts a quarter of a second has put less into the impression than one that lasts two, because the attack is fast but the energy accumulating is proportional to duration. So a short fortissimo costs the following piano less than a long one.
Change the texture rather than the level. A part dropping out changes the spectrum as well as the loudness, and the ear builds objects is the ladder about how much a listener attends to a change in what is sounding rather than in how much. A subito piano achieved by removing instruments is heard as a change even where the loudness has not fallen as far as the marking says.
Or accept the contrast. What the marking reliably produces is a large ratio between what is sounding and what is remembered, and that is a real perceptual event even though it is not the written level.
All four of these are things conductors and players do, none of them is in the score, and the model says why each one works.
The same passage answers one more question, and it is the one that makes the rate the point rather than a detail.
Two quantities of the same passage, through two windows nearly two orders of magnitude apart. A dynamic is smeared by the ear before it is judged and a dissonance is not, which is why a subito piano can be written and a subito consonance cannot: the second arrives intact and the first arrives as a slope.
Which computation produced the numbers
Each chord’s loudness is the collection’s excitation-pattern model: partial levels from a stated timbre for every note, summed across critical bands, giving a total in sones. That is the second rung of the loudness ladder’s machinery and it is why a six-part chord is not six times a one-part one.
The level track is that loudness converted back to a level and run through the loudness ladder’s two smoothers — a short-term one with a 22-millisecond attack and 50-millisecond release, and a long-term one with a 99-millisecond attack and a 2-second release.
The literal policy sets each chord’s level so that its instantaneous loudness equals its written value, by bisection. The over-time policy iterates the same bisection on the smoothed reading at the end of each chord, with the correction clamped at 45 decibels below the literal level — which is what makes an unreachable target report as unreachable rather than as a correction of seventy decibels.
Where the model stops
A written dynamic is not a number of sones. The mark that is not a level is the essay about that: forte is an ordinal instruction whose realisation depends on the instrument, the register, the ensemble and the hall. Assigning sones to the markings is a modelling choice and the passage’s written shape is invented.
The smoothers are a meter’s. They come from a published loudness standard designed for broadcast, and whether a listener’s impression has those constants is an assumption this collection makes throughout and cannot check.
What can be checked is how much the impossibility depends on the one constant that produces it. The floor the fifth chord cannot get below is the previous chord’s 5.6 sones decaying for the 1.2 seconds the new chord lasts, so it is 5.6·e^(−1.2/τ), and it crosses the written 2.0 at τ = 1.17 seconds:
| release constant | floor at silence |
|---|---|
| 1.0 s | 1.69 sones — producible |
| 1.17 | 2.00 — exactly the marking |
| 1.5 | 2.52 |
| 2.0, the published value | 3.07 |
| 3.0 | 3.75 |
So the result survives the published constant being wrong by 40 per cent in either direction and does not survive its being halved. A listener whose impression released in a second rather than two could hear this subito piano as written. The crossover at 1.17 seconds against a chord lasting 1.2 is also the tidiest statement of the whole rung: the marking becomes producible at about the point where the listener’s release gets shorter than the chord, which is the same sentence the essay makes about tempo, arrived at from the constant instead.
Chords are rectangles. Every chord here begins and ends instantly at a fixed level, and real chords have attacks, decays and internal shaping, all of which the smoothers would follow.
And the clamp is a modelling device. Forty-five decibels below the literal level stands for silence; a real orchestra’s floor is set by the hall and by how quietly its quietest instrument can play, and it is higher than that.
What the picture cannot show
It cannot show the conductor. A conductor faced with an unproducible subito piano does not attempt it literally; they lengthen the preceding chord, or shorten it, or take time before the change — all of which give the release room and none of which is in the score. That is the standard solution and it is a change to the tempo rather than to the dynamic.
Nor can it show the roughness. The anchor’s objective has two halves and this rung has only used the loudness one. A scoring over time should minimise the roughness trajectory as well, and a part entering has a roughness relation to what is already sounding that the instantaneous calculation does not have.
And it cannot show that a listener has one impression. A listener following a line in a texture is not integrating the whole sound; they are attending to a part, and the impression that matters may be that part’s rather than the ensemble’s.
Whose music, and when
The passage’s shape — a growth to a full tutti and then a sudden thinning — is a classical and romantic device, and the subito piano after a fortissimo is one of the most reliably marked things in nineteenth-century orchestral scores. Beethoven is the composer most associated with it and it is everywhere in the repertoire after him.
That the device is common is what makes the result interesting rather than a modelling error. If the marking were unproducible in the naive sense, composers would have stopped writing it. What the arithmetic says is that it produces something other than what it says — a contrast against a memory rather than a quiet — and that the something is large and reliable.
There is a second historical reading available. The dynamic markings in scores got finer and more numerous through the nineteenth century, at exactly the period when orchestras got larger and harmonic rhythm slowed. Both of those changes move music toward the slow end of the sweep, where the impression follows the marking — and away from the range where a marking and its effect part company.
Where this ladder goes next
Four rungs. The two halves of the scoring problem joined; the ranking survives a change of dynamic and the chord does not; a dynamic mark changes what a note is; and now a dynamic mark is an instruction to a listener’s memory, which sometimes cannot be followed.
What is owed next is the roughness over time. This rung has made the loudness half of the objective a functional and left the other half at an instant, and the two are not symmetric: loudness integrates and roughness may not. Whether a listener’s impression of harshness has a release like the loudness one is a question with a partial answer already in this collection — a dissonance has to last prices how long a roughness must persist to register — and joining those two would give the anchor a complete objective in time rather than half of one.
Part 4 of 14
One essay in the series on orchestration. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
BalanceDynamicsLoudnessNotationOptimisationOrchestrationTemporal integrationTexture
- A page has two decibels dynamics, loudness, notation, orchestration, texture
- A final chord is not made loud by adding to it dynamics, loudness, orchestration, texture
- A final chord stands out for a twentieth of a second dynamics, loudness, notation, temporal integration
- A part entering is not a change of level dynamics, loudness, orchestration, texture
- A staccato is a dynamic mark dynamics, loudness, notation, temporal integration
- The dynamics are in the score already dynamics, loudness, orchestration, texture