Form and structure

Loud is relative, and it comes down slowly

The account of loudness had a model of a moment and the account of closure asked it for a model of a form. The published one exists and its content is a pair of numbers that are not the same: a listener's running impression of how loud the music is rises to meet a step in a fifth of a second and takes seven seconds to come back down. A twenty-decibel crescendo spread over eight seconds therefore buys almost no contrast at all, and the same twenty decibels taken as a step buys a factor of two.

Assumes: A chord is not as loud as its notes · A silence long enough to be an ending

The closure ladder recorded a debt in the plainest possible terms: every ending in every repertoire is a dynamic event, and this collection had no model of dynamics in a form at all. The loudness ladder’s second rung supplied a model of a moment and recorded, from the other side, that the missing piece was no longer the loudness of a moment but the loudness of a moment against the moments around it.

The model for that exists, has existed since 2002, and its content is two time constants that are not the same.

Crescendo, and what the impression doesA crescendo of 20 dB over 8 seconds, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.02 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.1.02×024681012051015secondsloudness, sonessounding nowshort-term — theloudness of a notelong-term — theloudness of a passage
Fig. 1 A twenty-decibel crescendo spread over eight seconds, with three loudnesses drawn on it. The pale line is what is physically sounding. The dashed line is the short-term loudness — the loudness of the moment. The heavy line is the long-term loudness, which is what a listener reports as the loudness of the passage, and it follows the crescendo so closely that the two are hard to tell apart. The widest gap between the two anywhere in the passage is a factor of 1.02.

A crescendo of twenty decibels, taken slowly, is heard against a reference that has moved with it. Not entirely — the absolute loudness has quadrupled and nothing takes that away — but the contrast, which is what a crescendo is written for, is almost completely spent on moving the listener’s own reference.

The two constants, and why only one of them is interesting

The model is two one-pole smoothers in series applied to the instantaneous loudness. The first is short-term and gives the loudness of a note; the second is long-term and gives the loudness of a passage. Each has a different time constant depending on whether the loudness is rising or falling:

rising falling
short-term 22 ms 50 ms
long-term 99 ms 2 s

Three of those four numbers are on the same order and the fourth is twenty times larger than its own partner. The whole musical content of the model is in that one asymmetry.

That is nearly true and the section on peak spacing below finds the exception, which is worth flagging here because it is the other end of the same chain. The short-term attack of twenty-two milliseconds does nothing at all to a dynamic device — every device in this essay lasts tenths of seconds at least — and it does a great deal to a texture, because a passage of short loud events shorter than about that is a passage whose peaks the model never reaches. So the twenty-two milliseconds is inert for everything the notation writes and decisive for everything the orchestration produces, which is a fair description of the boundary between the two halves of this collection.

A step up is absorbed in a fifth of a second and a step down takes seven. How long the running impression of loudness takes to come within a tenth of a step of a stated size, drawn separately for a step up and a step down. The two curves are the same model with the same input and differ only in which of two published time constants applies — 99 milliseconds while the loudness is rising and 2 seconds while it is falling. A 25-decibel step is absorbed in 0.23 s going up and 7.7 s coming down, a ratio of 34 to one.
Fig. 2 How long the running impression takes to come within a tenth of a step, drawn separately for a step up and a step down. Twenty decibels is absorbed in 0.22 seconds going up and takes 6.85 seconds coming down — a ratio of thirty-one to one, on the same input, from the same model, with nothing changed but which of the two constants applies. The buttons play a step in each direction.

The asymmetry has an obvious purpose outside music. A listener needs to know at once that something has got loud, and needs to keep a memory of how loud things have been so that a lull is recognised as a lull rather than as silence. It is the same shape as the asymmetry in forward and backward masking, where the ear’s forgetting is slow in one direction and nearly instant in the other, and the two are almost certainly the same machinery.

Inside music, it has a consequence nobody would write down as a rule.

Why the passage’s loudness is nearer its maximum than its mean

Before the devices, one consequence of the asymmetry that is not about any of them.

If the impression rises in a tenth of a second and falls over two, then a passage that is loud a tenth of the time and quiet the rest is not heard as a tenth as loud as a continuously loud one. The impression jumps to the loud moments and slides back only part of the way before the next one arrives. The loudness of a passage is a maximum-weighted statistic, not an average, and how far toward the maximum it sits depends entirely on how the loud moments are spaced.

That is worth stating because it is the arithmetic behind a complaint every recording engineer has heard. A dense, busy passage with frequent peaks reads as loud even when its average power is modest; a sparse passage with the same peaks and long gaps does not, because the impression has time to fall between them.

Running it rather than reasoning about it says the effect is real and that “entirely” is too strong. Taking a passage that is twenty-five decibels louder for a tenth of its time, and asking where the running impression sits between the passage’s mean and its maximum:

peaks every how far toward the maximum
0.25 s 40%
0.5 s 51%
1 s 57%
2 s 56%
4 s 44%
8 s 25%
16 s 10%

The curve has a maximum and it is at about one second. Sparse peaks fall off as the reasoning above predicts — at sixteen seconds the impression is a tenth of the way up and the passage is heard at its mean. But closely spaced peaks fall off too, and that direction is not in the reasoning at all.

The reason is the other constant. At a quarter-second spacing and a tenth duty the loud moments are twenty-five milliseconds long, which is about the short-term smoother’s own attack time — so the impression never reaches the peak before the peak has gone. Widening the peaks to a quarter of the cycle instead puts the half-second spacing at 75 per cent rather than 51, ahead of everything in the table above.

So there are two quantities and the essay’s sentence names one. The impression sits near the maximum when the loud moments are long enough for the fast smoother to reach them and close enough for the slow one not to fall between them, which is a window rather than a direction — and the window’s best spacing, one to two seconds, is the tempo range this collection has spent a whole ladder on. A passage whose peaks arrive at about a beat is the loudest arrangement of a fixed amount of energy that the model allows.

The device that gets the whole two seconds

Subito piano, and what the impression doesA step down of 20 dB, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.00 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.1.00×024681012051015secondsloudness, sonessounding nowshort-term — theloudness of a notelong-term — theloudness of a passage
Fig. 3 The same twenty decibels, taken downward as a step. The moment drops instantly and the impression does not: for the next several seconds the passage is quieter than what a listener is still carrying, by up to a factor of 3.6. The gap here is more than three times the largest gap the crescendo above could produce with the same twenty decibels.

A subito piano buys three and a half times the contrast a slow crescendo of the same size buys. That is not a claim about taste; it is the ratio of two numbers computed from one model with one input and one pair of published constants.

What each dynamic device buys, against the running impression. Six dynamic devices, each scored twice: how much louder the loudest moment is than the running impression at that moment, and how much quieter the quietest is. A crescendo of twenty decibels spread over eight seconds buys a contrast of 1.02 — the impression tracks it almost exactly — while the same twenty decibels taken as a step buys 2.22. The largest number in the figure is a step down rather than any step up, which is the asymmetry in the two time constants read as a piece of orchestration advice.
Fig. 4 Seven dynamic devices scored against the running impression: the upper bar is how much louder the loudest moment is than the impression at that moment, the lower bar how much quieter the quietest is. Everything with a step in it scores; the eight-second crescendo scores 1.02, which is to say nothing. A sforzando scores on both bars at once, because it is a step up followed immediately by a step down.

The sforzando row is the one worth reading twice. It scores 2.22 on the upper bar and 3.58 on the lower, from a single loud note four tenths of a second long in a passage that is otherwise steady — which is more total contrast than any other device in the list, bought with less than half a second of anybody’s effort. The reason is that it is two steps rather than one, and the second of them gets the two-second constant.

The ordering is the finding. A crescendo is the least efficient way to buy loudness contrast, per decibel spent, of every device in the list — and it is the one that costs the most effort to execute, because it has to be sustained by everybody for eight seconds.

The sforzando’s efficiency has a limit the table does not draw, and it is the short constant again. A sforzando four tenths of a second long clears the twenty-two millisecond attack by a factor of eighteen, so the impression reaches it completely; shorten it toward a tenth of that and the upper bar collapses while the lower one survives, because the release constant does not care how brief the loud moment was. A very short sforzando is all aftermath, which is a description of an accent rather than of a dynamic, and the boundary between the two is a duration this model puts at a few tens of milliseconds.

What a crescendo is actually for

The obvious objection is that composers write crescendos anyway and are not stupid, so the model must be measuring the wrong thing.

It is measuring one thing, and the answer is that a crescendo is not bought for contrast.

The absolute loudness at the end of the crescendo is four times what it was at the start, and it stays there. What a crescendo buys is a new plateau, arrived at without a seam — a level the music can then sit on, with the impression having come along quietly rather than being startled. What a step buys is the seam. They are different products and the model says so: the crescendo’s long-term loudness ends at exactly the same place as the step’s, and only its middle is different.

That reading also explains the one place the model would otherwise look wrong, which is the long orchestral crescendo of the eighteenth century onward. The Mannheim crescendo is not remembered for a contrast at its peak; it is remembered for the arrival, and the arrival is the plateau.

Twice as loud is ten decibels, not twice the pressure. Loudness in sones against loudness level in phons, from Stevens's power law: above 40 phons the sone value doubles for every ten phons. Ten equal sources are ten times the power and about ten decibels, so they sound roughly twice as loud as one — which is why a section of ten violins is not ten violins loud.
Fig. 5 The conversion underneath every number above: loudness in sones against loudness level in phons, with equal-source doublings marked. Above 40 phons the curve doubles every 10 decibels — which is why a twenty-decibel crescendo is a factor of four in loudness and not a factor of a hundred in anything a listener would report.

The ending the closure ladder asked about

The debt that produced this rung was about endings, and endings are where the asymmetry is worth the most.

Final diminuendo, and what the impression doesAn ending that fades over 6 seconds, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.00 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.1.00×0246810120246810121416secondsloudness, sonessounding nowshort-term — theloudness of a notelong-term — theloudness of a passage
Fig. 6 An ending that fades over six seconds. The impression cannot follow it down — the fall is faster than the two-second constant allows — so the last four seconds of the piece are heard against a reference set by the music that has already finished. The ratio reaches 2.5 at the quietest point. A fade is an ending, in this model, because the reference is still up there.

That is a different mechanism from the four the closure ladder already has. An ending is five harmonic components with no total; an ending is a deceleration with one free parameter; an ending is a silence between three and a half and five and a half seconds long. Each of those is a cue that has to be recognised. A dynamic ending is not recognised; it is a state a listener is left in, and the state lasts about as long as the silence that follows it.

The two numbers are worth putting side by side: the loudness impression takes 6.9 seconds to come down from a twenty-decibel drop, and the silence rung found that a gap is heard as an ending somewhere between 3.5 and 5.5 seconds. The silence that ends a piece is shorter than the time the loudness impression takes to decay through it, which means the listener is still holding the loud music when the silence has already become an ending.

The consequence for how an ending is written is that the two cues are not interchangeable. A ritardando can be recognised as an ending as soon as enough of it has happened, which the deceleration rung puts at a handful of notes; a diminuendo cannot be recognised at all, and works by leaving a gap between what is sounding and what is remembered that persists for as long as the constant says.

Three ways to arrive at the same final tempoTempo against position in the closing passage, ending at 35 per cent of the opening tempo, for curvature exponents 1, 2, 3. All three begin and end at the same tempo, so what separates them is the middle: at the halfway point they read 68 per cent for linear in score position, 75 per cent for constant deceleration, 80 per cent for q = 3. The straight line is the one nobody plays. Measured ritardandos fit the decelerating curves, which is the whole of Kronman and Sundberg's argument: a closing gesture has the shape of a body stopping rather than of a dial being turned, and the parameter that varies between performances is the final tempo rather than the shape.the final tempo — 35%linear in score positionhalfway: 68%constant decelerationhalfway: 75%q = 3halfway: 80%all three endat the same tempo0.000.200.400.600.801.0000.20.40.60.81position in the closing passagetempo, as a fraction of the opening
Fig. 7 The other continuous ending cue, for comparison: the tempo curve of a final ritardando, with one free parameter for its shape. A deceleration is a cue about rate and is recognised from the pattern of onsets; the dynamic cue is about level and is not recognised at all. A real ending has both at once, and each of them can now be computed, which was not true before.

Which computation produced the numbers

The instantaneous loudness at each sample is the site’s own conversion — a level to a loudness level through ISO 226, and a loudness level to sones through Stevens’s power law, both of them the functions the ladder’s first rung built.

That series is then passed through two one-pole smoothers, each of which chooses its coefficient by comparing the incoming value to its own current output: rising takes the attack constant, falling takes the release. Nothing is fitted; the four constants are Glasberg and Moore’s, quoted in the source as per-millisecond coefficients and converted here to time constants, which for coefficients this small is a division to well under one per cent.

The “contrast” of a passage is then the largest ratio anywhere in it of the short-term loudness to the long-term loudness, and the “relief” is the largest ratio the other way. Those two are the essay’s own definitions rather than published measures, and they are chosen to be the smallest thing that could express what a dynamic device is for.

Where the model stops

The long-term loudness is a model of a reported number, not of an experience. It was fitted to listeners saying how loud a time-varying sound was overall, on stimuli measured in seconds. A symphonic movement is measured in minutes and nobody has asked whether the same integrator describes what a listener carries across four of them.

Nothing here knows about anything but level. A texture thinning from twelve parts to two is a dynamic event in every practical sense and this model sees only whatever level change comes with it. The loudness ladder’s own second rung says that is a serious omission: the same power spread over more of the spectrum is a great deal louder, so a passage can get quieter with no reduction in power at all, by closing up.

And the corpus question is exactly where it was. Everything above says what a stated dynamic shape does. Which shapes are actually used, how long a real crescendo lasts, whether the subito piano is more common at the ends of phrases than in the middles — all of that needs a corpus of performances with dynamics extracted, which this collection does not have and which the closure ladder recorded as owed for the same reason.

Whose music, and where the arithmetic is already known

The claim about the ear is not about anybody’s music. The claim about devices is about a repertoire, and it is a narrow one: the crescendo as a sustained orchestral device belongs to European art music from the middle of the eighteenth century, and the subito markings are common in the same tradition from Haydn onward.

What the arithmetic here says about that repertoire is not new to the people who wrote it. It is present as craft in exactly the form the model predicts: a subito piano is written as a step and never as a fast diminuendo, and the marking that means “get loud suddenly” barely exists — a composer wanting a sudden loud writes an accent or a sforzando on a single chord, which the figure above scores on both bars because it is a step up and immediately a step down.

That asymmetry in the notation is what the asymmetry in the two time constants predicts. A sudden quiet is a device and a sudden loud is an accent, and they are different words because they last for different lengths of time.

The place the reading fails is the tradition it was not built from. A gamelan’s dynamic shaping is done by the number and kind of instruments playing rather than by everybody playing harder, so the level and the spectrum move together and this model cannot separate them. That is not a failure of the ear’s constants, which are the same everywhere; it is a failure of a model that reads only one axis.

What the picture cannot show

Whether the reference is loudness at all. The model integrates loudness, which is one candidate for what a listener holds. A listener might instead hold a level, or a loudness per band, or the loudness of the loudest recent event rather than a running average — and the last of those would behave very differently, because a single sforzando would raise the reference for the whole passage.

Nor whether it is one reference or several. Loudness adds across critical bands and not within them, so there is a serious possibility that the running impression is per-band as well — in which case a passage could get quieter in one part of the spectrum and louder in another with no net change in the number this model computes, and a listener would report a change anyway.

And the figures are of one dynamic axis with nothing else moving. In a real crescendo the tempo moves, the texture thickens and the register spreads, and each of those is a loudness change by the band arithmetic of the previous rung rather than by anybody playing harder.

Where this ladder goes next

Three rungs. What one tone’s loudness is; what more than one tone’s is, which turned out to be decided by the critical band; and now what a passage’s is, which turns out to be decided by two constants that differ by a factor of twenty in the one direction that matters.

The rung after it is the one the last omission names, and it is a computation this collection now has both halves of. A texture that thins is quieter for two reasons at once — fewer sources, and fewer bands — and the band model of the second rung says the second is much the larger. So the dynamic curve of a piece with a changing texture is not the curve of anybody’s effort, and it is computable from a score without any performance data at all: count the parts, place them in bands, sum the loudnesses, and run the result through the two smoothers above. That would give a dynamic reading of a written score, which is a thing this collection has never produced.

Part 3 of 8

One essay in the series on loudness. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 14.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AdaptationClosureDynamicsIntegration windowLoudnessMemory decayOrchestrationSone