Perception and the listener

A subito piano is four seconds longer in the bass

Every loudness figure with time in it converts level to loudness at one kilohertz, and the equal-loudness contours say that no other frequency works that way. Joining the two sorts the published numbers into those that were about the treble and those that were not. Three move a great deal — a twenty-decibel crescendo is worth 27 phons on a bass note and 20 on a high one, and the seven seconds a subito piano takes becomes eleven and a third. Three do not move at all, and the reason they do not is the same reason in every case.

Assumes: A staccato is a dynamic mark · The quietest thing audible, and why the volume knob is a tone control

This ladder has two halves and they have never been introduced.

The first rung is the equal-loudness contours, and their entire content is that the map from a level to a loudness is a different map at every frequency — the curves are not parallel, they converge at the bottom of the spectrum, and a volume knob is therefore a tone control. Everything on the ladder above it that has time in it — the running impression, the two decibels a page commands, the one decibel an entering part is worth, the eight phons an articulation mark commands — converts level to loudness at one kilohertz, because that is the convention Glasberg and Moore’s model is published in.

So a dynamic mark on a double bass part and the same mark on a piccolo part have been treated as the same event, on the same page, in the same ink, by every figure on this ladder.

A 20-decibel crescendo is 27 phons on a bass note and 20 on a high one. The same change of level, from 60 to 80 decibels, converted to loudness at each register through ISO 226's equal-loudness contours rather than at one kilohertz. The heavy curve gives each note a string spectrum, so its partials are converted in their own bands and summed; the pale one is the fundamental alone. On the spectrum-aware curve the crescendo is worth 27.3 phons at C1 and 20.3 at C7. On the fundamental alone it is 77 at C1, which is not a finding but an artefact: a 60-decibel tone at 33 hertz sits 1.8 decibels above the threshold of hearing and is very nearly nothing. The honest correction is the smaller one, and it is still a difference of 7.1 phons across the compass for a mark written in the same ink.
Fig. 1 What twenty decibels of crescendo is worth in phons, register by register, computed through ISO 226 rather than at one kilohertz. The heavy curve gives each note its own partials; the pale one takes the fundamental alone, and below about a hundred hertz that is a tone very near the threshold of hearing.

The join is one function and nothing in the published model has to be rewritten. Convert the note’s level to a loudness in the bands its own partials occupy, ask what level at one kilohertz would be that loud, and hand that to the integrator. The integrator is untouched; it is simply being fed a level that has been through the contours instead of around them.

The crescendo, which is larger in the bass and not by as much as it looks

Take a crescendo from sixty decibels to eighty — mezzo-forte to fortissimo, roughly, and the step every settling figure on this ladder uses.

At a kilohertz that is twenty phons by definition, and a factor of four in sones. On a note at C2 with a string spectrum it is 23.2 phons and a factor of five; at C1 it is 27.3 phons and a factor of 6.65. The same mark, the same player, the same twenty decibels, and a loudness event two thirds again as large at the bottom of the compass as at the top.

That is the honest number, and it is worth saying loudly that the naive version of this computation gives a far more exciting one and is wrong. Taking each note as its fundamental alone, the twenty decibels is worth 77 phons at C1 — a factor of two hundred in sones. It is an artefact. A pure tone at 32.7 hertz and sixty decibels sits 1.8 decibels above the threshold of hearing, so the quiet end of that crescendo is very nearly silence and any rise from it is enormous by ratio. Nothing in music is that tone. A double bass playing its bottom C radiates a spectrum whose partials run up through the register where the ear is at its best, and the rung that counted which partials survive a chord’s own masking is the same observation from the other side: at the bottom of the compass a note is carried by everything except its fundamental.

So the register correction is real and it is a small fraction of the correction a pure-tone calculation would have made. Measured against the twenty phons a kilohertz gives, the naive computation adds 57 phons at C1 and the spectrum-aware one adds 7.3 — a factor of eight in phons and thirty-two in sones. That is the reason this rung has to be computed with a spectrum in it rather than as a substitution of one frequency for another.

It is worth pausing on what the surviving number means musically, because a factor in sones is easier to feel than a count of phons. Sones are a scale on which loudness doubles every ten phons, so a crescendo of 23.2 phons on the C two octaves below middle C makes the note five times as loud, and the same crescendo on a high one makes it four times. Five against four is not a subtlety. It is about the difference between a mezzo-forte and a forte in every published account of what those marks are worth, arriving free, because the low instrument was written low.

And the direction is the one an orchestrator would guess wrong. The bass of a texture is the part conventionally described as needing help to be heard — it is the part the masking ladder finds hides itself — and it is also the part whose dynamic range is largest once the contours are taken into account. Both are true at once and they are about different things: how much of the note arrives, and how much a change in it is worth.

Every point on one curve sounds equally loud. The equal-loudness contours of ISO 226:2003, evaluated from the standard's own parameters. The lowest curve is the threshold of hearing. Because the curves are not parallel — they crowd together in the bass and spread apart in the middle — the same change in decibels is a different change in loudness at every frequency, and a spectrum that was balanced at one level is not balanced at another.
Fig. 2 ISO 226’s equal-loudness contours, which everything here begins from, drawn at the five loudnesses a passage of music actually spends its time between. Every point on one curve sounds equally loud. The curves are further apart at the bottom of the spectrum than in the middle, and the gap closing as the level rises is the whole mechanism of this essay.

The mechanism is the contours converging, and it is one number

Why the bass and not the treble is a question about the shape of the family above, and it has a one-number answer: how many phons one decibel buys.

One decibel is one phon only at one kilohertz. The slope of ISO 226's map from level to loudness — how many phons one decibel buys — against the level it is taken at, for 7 registers. At a kilohertz it is exactly one at every level, which is the definition of the phon and not a measurement. At C1 it runs from 0.00 at 40 decibels to 1.87 at 100, because the contours down there converge as the level rises. Every running-loudness figure has been drawn along the flat line, and a note in the bass is on one of the steep ones.
Fig. 3 The slope of ISO 226’s map from level to loudness at each register and each level. The flat line at one is a kilohertz, where it is one by the definition of the phon and not by measurement. Every running-loudness figure so far was computed along that line.

At a kilohertz the answer is one at every level, and it is one because the phon is defined as the level of an equally loud kilohertz tone. That row of the figure is arithmetic and not a measurement, which is what makes it the right baseline.

Everywhere else it is not one. At C3 a decibel buys 1.00 phons at forty and 1.38 at a hundred; at C2 it runs from 0.57 to 1.61; at C1 from nothing at all — because forty decibels there is under the threshold — to 1.87. The contours are converging as the level rises, so the same decibel is worth more the louder the passage already is, and worth more the lower the note is.

The treble does the opposite by a small amount: at C7 a decibel buys 0.97 phons at every level, so a piccolo’s dynamics are slightly compressed relative to the kilohertz convention. The whole compass therefore spans about a factor of two in what a decibel is worth, and the ladder has been working at one end of it.

And the subito piano, which is where it costs the most

The result that moves furthest is not a size. It is a time.

The rung that established the running impression found the asymmetry this ladder is built on: the impression rises with the music in about a fifth of a second and takes seven to come down, because the published model’s release constant is two seconds against an attack of ninety-nine milliseconds. Seven seconds is a number about a kilohertz.

A subito piano takes 11.3 seconds to arrive in the bass and 7.3 in the treble. How long the running impression of loudness takes to come within a tenth of a 20-decibel step, register by register, with each note's level converted to loudness in its own bands and its own partials rather than at one kilohertz. Going up the answer barely moves: 0.23, 0.23, 0.23, 0.22, 0.22, 0.22, 0.22 seconds, because the attack constant decides it and the attack constant has no frequency in it. Coming down it moves a great deal — 11.30 seconds at C1 against 7.30 at C7 — because the criterion is a fraction of where the loudness ends up, and in the bass the same decibels take the loudness much further down. The asymmetry published as thirty-one to one is 49 to one at C1.
Fig. 4 How long the running impression takes to come within a tenth of a twenty-decibel step, register by register, with each note’s level taken through the contours in its own bands and with its own partials. The rise barely moves. The fall moves by four seconds.

Going up, the answer is 0.220 seconds at C7 and 0.230 at C1 — four per cent across six octaves, and the reason is that the attack constant decides it and the attack constant has no frequency in it.

Coming down, it is 7.30 seconds at C7 and 11.30 at C1. The published asymmetry of about thirty-one to one is 33 in the treble and 49 at the bottom of the compass.

The mechanism is worth spelling out because it is not the same as the crescendo’s. The criterion for “the impression has caught up” is a fraction of where the loudness ends up — within a tenth of the final value — and the release is exponential in sones, so the time taken is the release constant times the logarithm of the ratio the loudness has to fall through. In the bass that ratio is much larger, because the same twenty decibels takes the loudness much further down: the diminuendo from sixty to forty decibels is worth 35.2 equivalent decibels at C1 against 22.5 at C7. More decades to fall through, at the same rate, is more time.

There is a second consequence of the same arithmetic and it runs the other way. The step up is barely register-dependent, so the contrast a sforzando buys — how much louder the moment is than the running impression at that moment — is nearly the same everywhere. What differs is only how long the ear takes to give the contrast back. A ladder that had measured only the rises would have found the register question to be nothing, and it is a real risk of the way these figures are usually drawn: the attack is the visible part of a settling curve and the release is the part that runs off the right-hand edge.

So a subito piano on the double basses is still arriving when a subito piano on the flutes has finished — by four seconds, which at a slow tempo is two bars. That is a claim about scoring rather than about physics, and the repertoire it is a claim about is the one where the device is a structural event rather than a colour: the Classical and early Romantic symphonic literature, where a subito piano is written into a full orchestral tutti and the passage after it is expected to sound quiet. If the device is placed in the bass alone the relief takes half again as long to arrive.

Two of this ladder’s numbers do not move at all

Which makes the rest of the sorting interesting, because the obvious conclusion — that every number on the ladder is a number about the treble — is false, and the two that stand are the two that looked most vulnerable.

A note reaches the passage's loudness at 250 milliseconds in the treble and 255 in the bass. What a note of a stated length delivers to the loudness of the passage it is in, as a fraction of what the same level reaches when it is held, computed separately at 7 registers. The curves are indistinguishable. Nine tenths is reached at 255, 255, 250, 250, 250, 250, 250 milliseconds from C1 up to C7 — a spread of 5 milliseconds, which is one step of the search that found them. The reason is the same as for the articulation: the quantity is a ratio to the same level held at the same pitch, so the map from level to loudness cancels and only the integrator's own two constants are left, and those have no frequency in them.
Fig. 5 What a note of a stated length delivers to the loudness of the passage it is in, computed separately at seven registers. Seven curves are drawn and one is visible. Nine tenths is reached at 250 milliseconds in the treble and 255 in the bass.

The note-length curve — the finding that a note needs 245 milliseconds to deliver its contribution to a passage’s loudness, which is what makes an articulation mark a dynamic — is the same curve at every register to within one step of the search that found it. And the articulation dynamic itself is 8.46 phons at C1 against 8.22 at C7, a difference of a quarter of a phon over six octaves.

An articulation is worth the same 8.5 phons wherever the note is. What an articulation mark commands, in phons, against tempo, computed separately for 7 registers with each note's level taken through the contours in its own bands. The curves lie on top of one another: at the fastest tempo drawn the span from a full legato to a quarter articulation is 8.22 phons in the treble and 8.46 in the bass, a difference of 0.26 phons over six octaves. This is the half of the account that does not move, and the reason is that the quantity is a ratio to the same level held at the same pitch — the contour is in the numerator and the denominator and cancels. What is left is its curvature, which is why the bass rows sit a little above the treble ones rather than exactly on them.
Fig. 6 What an articulation mark commands, against tempo, computed separately for seven registers. The curves lie on top of one another: a quarter of a phon separates the top of the compass from the bottom.

The reason is the same in both cases and it is a single sentence. Both quantities are ratios to the same level held at the same pitch. The map from level to loudness appears in the numerator and the denominator and cancels; what survives is only its curvature, which is why the bass rows sit a hair above the treble ones rather than exactly on them.

That is a general rule for reading this ladder, and it is the useful thing to take from the sorting. A number on this ladder is about the treble if it compares a loudness at one register against a loudness at another, or against an absolute criterion in phons. It is register-free if it compares a loudness against the same loudness with one parameter changed. The crescendo’s worth and the settling time are of the first kind; the note-length curve and the articulation span are of the second.

The test also explains why the two register-free results are the two an orchestrator would have expected to be fragile. An articulation mark is a relative instruction — play four fifths of each note rather than all of it — and a relative instruction against a fixed level is exactly the shape that cancels. A dynamic mark is an absolute one, and absolute instructions are the ones that do not survive the change of frequency.

By that test, the two decibels a page commands and the one decibel an entering part is worth are of the first kind — both are absolute quantities in phons, both were computed at a kilohertz, and both will move. Neither is recomputed here, because both are properties of a whole texture spread across the compass rather than of one note at one register, and a texture needs its own conversion rather than this one.

Which computation produced the numbers

The contours are ISO 226:2003, evaluated at any frequency through its published parametric form rather than interpolated between tabulated ones — a partial is not one of the standard’s twenty-nine frequencies, so a table lookup would have had nothing to say about most of what a note contains.

A note’s loudness is the sum, in sones, of its audible partials, each converted at its own frequency: a string spectrum of eight partials falling as 1/n, with the whole tone’s level defined as the root-sum-square of its components, which is the convention every masking figure on this site uses. A partial below the threshold of hearing at its own frequency contributes nothing.

The equivalent level is the level at one kilohertz of a tone that loud, obtained by inverting Stevens’s sone law and then the contours. The running loudness is then Glasberg and Moore’s 2002 model exactly as this ladder has always used it: two one-pole smoothers in series, 22 and 99 milliseconds attack, 50 milliseconds and 2 seconds release, driven at a kilohertz on the equivalent level. Nothing in the integrator was changed and nothing was fitted.

The settling criterion is within a tenth of the final loudness, which is the criterion the third rung used, and the step is twenty decibels from a base of sixty in the register concerned — so the up-step and the down-step are not the same number of equivalent decibels, and the figure reports both.

Where the model stops

The register of a note is not the frequency of its fundamental. Everything here places a note by its written pitch and gives it one spectrum, and the whole point of the previous section is that this is a spectral question rather than a pitch one. A bassoon and a double bass at the same written C are two different distributions across the ear’s bands and this arithmetic gives them the same answer. The right version of this rung takes a measured spectrum per instrument, which this collection carries for several instruments and has never assembled into a loudness figure.

The contours are for pure tones and this is applied to partials. ISO 226 is a standard about a single sinusoid presented frontally in a free field, and summing sones partial by partial ignores that components inside one critical band do not add as separate loudnesses — which is exactly what the rung about a chord not being as loud as its notes is about. In the bass, where a note’s low partials share a band, this over-counts; the direction of the error is known and its size is not computed here.

What 20 decibels off the volume takes away. The perceived loudness lost, in phons, when 20 dB is removed from a tone that was 80 phons loud, computed frequency by frequency from ISO 226:2003. The line is not flat, so turning a piece of music down does not turn all of it down equally: the bottom of the spectrum loses about 42 phons where the middle loses 20.
Fig. 7 What twenty decibels off the volume takes away, phons lost against frequency, from a tone that was eighty phons loud. An equal loss would be the flat line. This is the earliest result and it is the same asymmetry this whole essay is about, read at one instant rather than through an integrator.

A crescendo is not a level change on a fixed spectrum. A player producing twenty more decibels produces a brighter note as well as a louder one, and the brightening moves energy upward into bands where the contours are flatter — which would reduce the register effect further. This computation holds the spectrum fixed and varies only the level, so it is an upper bound on the correction rather than an estimate of it.

And the threshold of hearing is a young adult’s. The standard’s threshold curve is the one that makes a sixty-decibel tone at C1 marginal, and it rises with age fastest at the top of the spectrum and least at the bottom. An older listener has a flatter set of contours in the region that matters here, so every number in this essay is a number about a young ear.

Where this ladder goes next

Eight rungs. What one tone’s loudness is; what more than one tone’s is; what a passage’s is, which two constants decide; what a page’s is, from the texture alone; what a page’s is worth, which is two decibels; what an entry is worth, which is one; what the note lengths are worth, which is two and sometimes eight; and now which of those numbers were about a kilohertz, which is three of them and not the three that looked most exposed.

What is owed now is the balance, and it is the one the sorting above has already named and refused to do. Two of this ladder’s headline numbers — the two decibels a page commands and the one an entering part is worth — are absolute quantities in phons computed at a kilohertz on a texture, and a texture is spread over the whole compass rather than sitting at one register. Converting them properly means giving every part its own range and its own spectrum, summing the loudnesses across bands rather than at one frequency, and asking again what a change of scoring is worth; and the answer will not be a simple scaling, because the parts that dominate a texture’s loudness are not the parts that dominate its energy. This collection already carries the ranges, the spectra, the band-summing machinery and the scoring schemes those two rungs were computed from. It needs no corpus and no listener, and what would come out is whether the two decibels a whole page of orchestration commands survives being computed in the registers the orchestration is actually written in — a rung that could raise that number as easily as lower it, since the parts it under-weighted are the bass ones and the bass is where the contours do the most.

Part 8 of 8

One essay in the series on loudness. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DynamicsEqual-loudness contourIntegration windowLoudnessPhonRegisterSoneTemporal integration