Perception and the listener

A chord is not as loud as its notes

The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.

Assumes: The quietest thing audible, and why the volume knob is a tone control · A third is rougher in the bass

The first rung of this ladder worked out what a tone’s loudness is, ended by saying what it had not done, and named the debt precisely: the contours say what a tone’s loudness is; they do not say what a chord’s loudness is, or a room’s, or an orchestra’s. Two years of orchestration advice rest on the answer and this collection had no model of it at all.

The rule turns out to be short, and it is not the one anybody guesses.

3 tones, one power, and the interval between them. 3 tones of fixed total power, spread symmetrically about 440 hertz, drawn against the interval between neighbours. Piled on one pitch they are one sound of that power; separated by more than a critical band — 4.5 semitones here — they are 3 sounds whose loudnesses add, and the same power reaches 2.08 times the loudness at 5 semitones. The two lines are two models of the same rule and they disagree about how abrupt the change is, not about where it goes.
Fig. 1 Three tones carrying one fixed total power, spread symmetrically about A4, drawn against the interval between neighbours. Nothing about the power changes across the whole picture. Piled on one pitch the three are one sound; pulled apart by more than a critical band they are three sounds whose loudnesses add, and the same wattage is twice as loud. The two curves are two ways of computing the same rule and they disagree about how abruptly the change happens, which is the honest state of the question.

Intensity inside a band, loudness across them

The rule has two clauses and the whole of the subject is in the difference between them.

Components that fall inside one critical band add their intensities. Two tones a few hertz apart are not analysed separately by the ear — they excite the same patch of the basilar membrane and arrive at the loudness machinery as one thing with one level. Twice the intensity is three decibels, and three decibels is about a quarter more loudness. Two such tones are 1.23 times as loud as one of them.

Components in different bands add their loudnesses. Each band is converted to sones on its own and the sones are summed. Two tones a band apart, at the same total power as the pair above, come to 1.62 times one tone.

Same power. A ratio of 1.62 to 1.23, which is a third more loudness bought with no more energy, no more players and no more effort — bought entirely by where in the spectrum the power was put.

3 notes, 5 critical bands. The excitation 3 notes of a string spectrum cast along the Bark axis, and the critical-band groups their components fall into. Components inside one group add their intensities and are converted to loudness once; the groups' loudnesses then add. This sonority occupies 5 groups and comes to 40.0 sones, against 109.5 for the same components counted one at a time.
Fig. 2 One sonority as the loudness model sees it: a C major triad on C4 with a string spectrum, its components’ excitation drawn along the Bark axis, and the critical-band groups they fall into shaded underneath with each group’s own loudness printed on it. The total is the sum of those numbers. The figure in the margin is what the same components come to when each one is counted as a separate loudness, which is what this site’s published function has been doing.

Neither clause is a new discovery and neither is controversial; both have been in the psychoacoustics literature since Zwicker, Flottorp and Stevens measured them in the nineteen-fifties. What is new here is that nothing in this collection had ever applied them, and that applying them changes several results the collection had already published.

The experiment that isolates it

The hero figure is the whole argument in one variable, and the variable was chosen to make it hard to cheat.

Three tones. Their total power is held fixed, so every point on the horizontal axis is the same wattage arriving at the ear. Their geometric centre is held fixed at A4, because moving the centre moves the threshold of hearing under the whole calculation and would swamp the effect being measured. The only thing that changes is the interval between neighbouring tones.

At zero the three are one tone at 80 decibels. At twelve semitones apart they are three tones at 75.2 decibels each, which is the same total. And the ear reports the second as twice as loud as the first.

The solid line is the grouping rule stated above, applied literally: components within one critical band of each other are pooled, the pools are converted, the results are added. It has a step in it, at about five semitones at this pitch, because the rule has an edge in it — either two components share a band or they do not.

The dashed line is the same physics without the edge. Every component casts a skirt of excitation along the frequency axis rather than sitting in a box; the skirts add; the specific loudness at each point on the axis is computed from the total excitation there and the whole pattern is integrated. That version rises smoothly and reaches the same place.

The two models disagree about the sharpness and not about the destination, and neither is preferred here, because the published measurements sit between them: the loudness of a multi-tone complex at constant power is roughly flat while the spread is inside a band and rises after, with a knee that is softer than a step and harder than the excitation model’s slow curve.

The chord that is one sound

Now put a real chord into it, and the result stops being about acoustics and becomes about register.

A major triad spans seven semitones. Whether those seven semitones are wider or narrower than a critical band is not a property of the triad — it is a property of where the triad is.

The same chord is louder in the treble, at the same power. A 3-note chord spanning 7 semitones, drawn at 5 registers with the total power held constant, and scored against one note carrying all of it. Low down the whole chord fits inside one critical band — 35 semitones wide at C2 — so it is analysed as one thing and buys 0.99 times the loudness. At C6 the band is 2.8 semitones, the notes are in separate bands, and the same power is 2.24 times as loud.
Fig. 3 The same major triad at five registers, each at the same total power, scored against one note carrying all of that power. At C2 the critical band is thirty-five semitones wide, the whole triad is inside it, and the chord is not louder than the note at all. At C6 the band is under three semitones, the three notes are three separate sounds, and the chord is 2.2 times as loud. The dashed line is the smooth model, which finds a weaker effect for the same reason it found a gentler knee.

A major triad at the bottom of a piano is not louder than one note of the same total power. The same triad two octaves up is more than twice as loud. The notes have not changed, the power has not changed, and the ratio has moved by a factor of two.

That is worth putting beside the result the triad ladder already had. A chord is a register found that the same three pitch classes are about five times rougher at the bottom of a piano than in the middle, and the mechanism was the critical band: notes inside one band beat against each other, and roughness is what that beating sounds like. Here is the same band, doing something else entirely to the same chord — and the same width that sorts every interval into resolved and unresolved is doing both jobs.

Every voicing of a major triad, least rough first. All 27 arrangements of the same three pitch classes within 3 octaves from 65 Hz, scored for roughness. The best is spaced 28 then 3 semitones — wide below, close above — and the worst is the chord in close position at the bottom of the range, 4.7 times rougher with exactly the same notes in it.
Fig. 4 The register figure, drawn an octave lower than its own: every arrangement of one triad within three octaves of C2, ranked by roughness. The worst is close position at the bottom, and close position at the bottom is also the arrangement the figure above scores at 0.99 for loudness. The two rankings are not the same ranking — the smoothest voicing here is wide, and a wide voicing is also the loud one — but they are driven by the same width.

A low triad is rough and quiet, and both for the same reason. It is rough because its notes are inside one band, and it is quiet — relative to the power spent on it — because its notes are inside one band. The orchestration rule about not scoring close thirds in the bass has always been justified by the muddiness. Half of the case is that the muddiness is not even loud.

The critical band, measured in semitones. The width of the ear's frequency-analysis band at each pitch, converted from hertz into semitones. Two intervals drawn as horizontal lines cross the curves: below the crossing the interval fits inside one band and its notes are not resolved from each other, and above it they are.
Fig. 5 The width of the band, in semitones, at each pitch, which is the quantity the whole result turns on. A major third fits inside a band below about 500 hertz and is resolved above it. That crossing is the one the roughness result is built on; it is also, exactly, the register at which a triad starts being three loudnesses instead of one. Two models of the width are drawn, because they differ by a factor of three in the bass and the bass is where the question is.

Ninety players, and four times as loud

The other half of the debt was the orchestra, and it is the same arithmetic read in the direction that pays worst.

Independent players do not add their pressures — their phases are unrelated, so what adds is power. Ten players are ten times the intensity and ten decibels; ninety are 19.5 decibels. Then loudness goes as roughly the 0.3 power of intensity, which is what a sone is: ten decibels doubles it.

The conversion the section arithmetic runs on is the one the ladder’s first rung drew and did not spend: loudness in sones against loudness level in phons. Above 40 phons the curve doubles every ten phons, so a tenfold increase in power is a doubling in sones and no more — which is the whole reason a chord’s notes do not add up to a chord.

90 players, and how much louder than one. Independent players on one note add in power, so 90 of them are 19.5 decibels above one — and loudness in sones goes as roughly the 0.3 power of intensity, so the section is 3.9 times as loud rather than 90 times. The horizontal axis is doublings of the section, which is why the line is nearly straight: every doubling is 3 dB and about a quarter more loudness.
Fig. 6 Players on one note against how loud the section is, with the horizontal axis in doublings. Each doubling of the section is three decibels and about a quarter more loudness, so ninety violins are 19.5 decibels above one violin and just under four times as loud. Nothing in the picture is about the players; it is entirely a consequence of incoherent sources adding in power and loudness going as a fractional power of that.

Ninety violins are 3.9 times as loud as one violin. Not ninety times, and not the twenty-odd times a decibel count might suggest to somebody reading decibels as loudness.

This is where the choir rung and the soloist rung meet the arithmetic. A trained soloist heard over a full orchestra is not louder than it and could not be — the section arithmetic above says what a single voice would have to do to compete on power, and the answer is impossible. What the voice does instead — and what the dynamic marking on the page cannot say — is move to a part of the spectrum the orchestra has vacated, which in the language of this rung is: occupy a band nobody else is in, where the loudness adds rather than the intensity.

And what a choir does that a soloist cannot is now a second thing. Sixteen singers on one note are 12 decibels and 2.3 times the loudness of one, which is a poor return; the essay’s finding was that what a choir buys is not level at all but the disappearance of a beat rate. The loudness arithmetic here says why the section was never going to be the way to buy volume.

The cheap way to be loud

The expensive way to get louder is more power in the same place. The cheap way falls straight out of the second clause of the rule.

The same power, divided among more of the spectrum. One fixed total power — 80 decibels — delivered as 1, 2, … 12 tones, each a band and a half from the next so that none of them shares a critical band with another. Every extra tone takes power away from the others and the sonority gets louder anyway: 12 tones are 5.3 times the loudness of one carrying all of it. The exponent is about 0.7, which is one minus the 0.3 that relates loudness to intensity.
Fig. 7 One fixed total power delivered as one tone, then two, then twelve, each a band and a half from the last so that no two share a critical band. Every extra tone takes power away from every other one and the sonority gets louder anyway: twelve tones sharing 80 decibels are 5.3 times the loudness of one tone carrying all 80. The dashed line is n to the power 0.7 — which is 1 minus the 0.3 that turns intensity into loudness — and the arithmetic is that simple.

Twelve tones sharing one tone’s power are 5.3 times as loud as that tone. Every one of them is 10.8 decibels quieter than the single tone was, and the total is five times the loudness.

That is the tutti, and it explains a thing about orchestration that is usually taught as taste. The way to make an orchestra louder is not to make the strings play harder; it is to add instruments in registers and colours that are not already occupied. A doubling of the whole string section buys 3 decibels and a quarter more loudness. Adding a piccolo, three trombones and a triangle buys four new bands, and four new bands at the same total power would be more than twice the loudness — the trombones being much more than a band away from the violins is the entire mechanism.

The corollary is the one every recording engineer knows in the other direction: a mix in which everything occupies the same two octaves is quiet for its power, and the remedy is not the fader.

Which computation produced the numbers

Two models, both of them named, both of them running on functions this site already had.

The grouping model sorts the components by frequency, walks the list, and starts a new group whenever the gap to the previous component exceeds one Bark. Inside a group the intensities are summed and the group’s level is converted once, at the group’s intensity-weighted mean frequency, through ISO 226 to phons and then through Stevens’s power law to sones. Across groups the sones are added. A single component returns exactly what this site’s existing loudness function returns for it, so the scale is the sone by construction and nothing is calibrated.

The excitation model replaces the group with a skirt. Each component spreads along the Bark axis at 27 decibels per Bark downward and a level-dependent 6 to 16 decibels per Bark upward — which is not a new function either, it is exactly the masking spread this collection has been drawing since the masking ladder opened. The excitations are added as intensities at every point on the axis, converted to a specific loudness by Zwicker’s compressive law, and integrated. One constant is fixed by the definition of the sone at 1 kilohertz and 40 decibels, and it is the only fitted number in either model.

Both models borrow the same spreading function, written for a different job: the level a probe needs in order to be heard beside a masker. Its downward slope is fixed and its upward one flattens with level, so a loud low note reaches further up the spectrum than a quiet one does — and that asymmetry is doing work here that nobody put in on purpose.

The two disagree in a way that is worth stating rather than hiding. The grouping model reproduces the sone scale exactly — 10 decibels doubles the loudness, at every level — because that scale is inside it. The excitation model does not: it grows by a factor of 2.3 per 10 decibels rather than 2.0, because its upper skirt flattens as the level rises, so a louder tone occupies more of the axis and gains loudness twice over. Every number quoted above as a ratio between two spectra at the same total power is one where that error is common to both sides and cancels to about a per cent. No number here is an absolute loudness, and that is why.

What this collection had been getting wrong

There is a version of this that is not a new rung but a correction, and it is in the docstring of the site’s own loudness function, which has said so since the day it was written: the loudness of a complex tone summed over its partials is an upper bound, because partials inside a single critical band do not add as separate loudnesses.

What counting the partials one at a time overstates. One note of a string spectrum at 70 dB, its loudness computed three ways at 5 registers. The tall bar counts each partial as a separate loudness, which is what the published function does and says it does; the short one groups the partials that share a critical band. The overstatement is worst at C2, where it is a factor of 5.6, because a low note's partials are packed inside a band that is wide in hertz.
Fig. 8 One note of a string spectrum at 70 decibels, its loudness computed three ways at five registers. The tall bar counts every partial as a separate loudness, which is what this site’s published function does; the short bar groups the partials that share a band; the dot is the excitation model. The overstatement is worst in the bass, where a note’s first eight partials are packed inside a band that is very wide in hertz, and it is nearly a factor of six at C2.

The error is small in the middle of the range and enormous at the bottom. In the treble a note’s partials are spread over many bands and counting them separately is nearly right; at C2 the first eight partials of a string are all inside the same wide band, the grouping model pools every one of them, and the published upper bound overstates the note by a factor of five and a half.

No figure on this site was wrong, because the function was only ever used for its slope against level, which the docstring says and which the grouping does not change. But the number it returns was being read as a loudness in the one register where it is furthest from one.

Whose music, and when

The claims here divide into two kinds and the division is worth marking, because this site’s rule is to say whose music a claim is about.

The band arithmetic is not about anybody’s music. It is a property of a listener, measured in the nineteen-thirties and refined ever since, and it applies to a Balinese gong-chime and a Mahler tutti in the same way.

The orchestration reading is a claim about a repertoire. The specific practice — a tutti built by adding families rather than by adding desks, low chords voiced open, the melody given to whatever instrument owns a vacant band — is European orchestral practice from roughly Berlioz onward, and it was arrived at by ear. Berlioz’s treatise of 1843 gives the rule about spacing low chords and gives no reason for it beyond the sound. What the arithmetic here adds is not the rule but a second reason for it, and the second reason is a different one from the roughness the triad ladder found.

The one place the two traditions visibly disagree is the register question. Gamelan writing puts dense, low, closely-spaced strikes exactly where European orchestration says not to, and the arithmetic here says those strikes are getting very little loudness for their power. That is a claim about volume rather than about value; the point of a low unison in that repertoire is not to be loud.

What the picture cannot show

The grouping is single-linkage, and that is a real defect rather than a simplification. A chain of components each nine-tenths of a band from the next is pooled into one group however long the chain, which is wrong for a dense cluster spanning several bands. The excitation model has no such defect, which is why both are drawn on every figure that turns on the transition.

Nothing here is heard. Loudness is a report from a listener, and the sone scale is a scale of reports. Every number above is what a model of those reports predicts, and the model was fitted to listeners doing a magnitude-estimation task that some psychophysicists regard as measuring the task rather than the ear.

The components are steady and simultaneous. Loudness needs time to build — a very short note is quieter than a long one at the same amplitude — and none of the figures above has a time axis at all. A chord struck and released inside a fifth of a second is not the chord these numbers describe, and neither is a chord whose notes do not start together, where the simultaneous window is a fraction of each note.

And the register result is stated at constant total power, which no performance holds constant. A pianist playing a low triad does not play it at the same power as a single note; they play it with three fingers. At constant power per note the low triad is louder than one note, by about the amount three notes ought to be — the finding here is about what the spending buys, not about what happens when more is spent.

Where this ladder goes next

Two rungs. The first says what one tone’s loudness is and that the map from decibels to it bends with frequency. This one says what happens when there is more than one tone, and finds that the answer is decided by a quantity that was already on this site for a completely different purpose.

The rung after it is the one this makes unavoidable, and it is the one the closure ladder asked for from the other end. An ending is a dynamic event — a diminuendo, a final chord louder than everything before it, a texture thinning — and the closure rung recorded that this collection has no model of dynamics in a form. It now has one, and the missing piece is no longer the loudness of a moment but the loudness of a moment against the moments around it: a fortissimo is a fortissimo relative to what a listener has been hearing for the last twenty seconds, and there is a published adaptation time constant for exactly that. Whether a crescendo of a stated size is heard as one depends on that constant, and it is a computation rather than a corpus.

Part 2 of 8

One essay in the series on loudness. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 17.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Critical bandwidthDecibelEqual-loudness contourLoudnessMaskingOrchestrationPhonSone