A chord is not as loud as its notes
Assumes: The quietest thing audible, and why the volume knob is a tone control · A third is rougher in the bass
The first rung of this ladder worked out what a tone’s loudness is, ended by saying what it had not done, and named the debt precisely: the contours say what a tone’s loudness is; they do not say what a chord’s loudness is, or a room’s, or an orchestra’s. Two years of orchestration advice rest on the answer and this collection had no model of it at all.
The rule turns out to be short, and it is not the one anybody guesses.
Intensity inside a band, loudness across them
The rule has two clauses and the whole of the subject is in the difference between them.
Components that fall inside one critical band add their intensities. Two tones a few hertz apart are not analysed separately by the ear — they excite the same patch of the basilar membrane and arrive at the loudness machinery as one thing with one level. Twice the intensity is three decibels, and three decibels is about a quarter more loudness. Two such tones are 1.23 times as loud as one of them.
Components in different bands add their loudnesses. Each band is converted to sones on its own and the sones are summed. Two tones a band apart, at the same total power as the pair above, come to 1.62 times one tone.
Same power. A ratio of 1.62 to 1.23, which is a third more loudness bought with no more energy, no more players and no more effort — bought entirely by where in the spectrum the power was put.
Neither clause is a new discovery and neither is controversial; both have been in the psychoacoustics literature since Zwicker, Flottorp and Stevens measured them in the nineteen-fifties. What is new here is that nothing in this collection had ever applied them, and that applying them changes several results the collection had already published.
The experiment that isolates it
The hero figure is the whole argument in one variable, and the variable was chosen to make it hard to cheat.
Three tones. Their total power is held fixed, so every point on the horizontal axis is the same wattage arriving at the ear. Their geometric centre is held fixed at A4, because moving the centre moves the threshold of hearing under the whole calculation and would swamp the effect being measured. The only thing that changes is the interval between neighbouring tones.
At zero the three are one tone at 80 decibels. At twelve semitones apart they are three tones at 75.2 decibels each, which is the same total. And the ear reports the second as twice as loud as the first.
The solid line is the grouping rule stated above, applied literally: components within one critical band of each other are pooled, the pools are converted, the results are added. It has a step in it, at about five semitones at this pitch, because the rule has an edge in it — either two components share a band or they do not.
The dashed line is the same physics without the edge. Every component casts a skirt of excitation along the frequency axis rather than sitting in a box; the skirts add; the specific loudness at each point on the axis is computed from the total excitation there and the whole pattern is integrated. That version rises smoothly and reaches the same place.
The two models disagree about the sharpness and not about the destination, and neither is preferred here, because the published measurements sit between them: the loudness of a multi-tone complex at constant power is roughly flat while the spread is inside a band and rises after, with a knee that is softer than a step and harder than the excitation model’s slow curve.
The chord that is one sound
Now put a real chord into it, and the result stops being about acoustics and becomes about register.
A major triad spans seven semitones. Whether those seven semitones are wider or narrower than a critical band is not a property of the triad — it is a property of where the triad is.
A major triad at the bottom of a piano is not louder than one note of the same total power. The same triad two octaves up is more than twice as loud. The notes have not changed, the power has not changed, and the ratio has moved by a factor of two.
That is worth putting beside the result the triad ladder already had. A chord is a register found that the same three pitch classes are about five times rougher at the bottom of a piano than in the middle, and the mechanism was the critical band: notes inside one band beat against each other, and roughness is what that beating sounds like. Here is the same band, doing something else entirely to the same chord — and the same width that sorts every interval into resolved and unresolved is doing both jobs.
A low triad is rough and quiet, and both for the same reason. It is rough because its notes are inside one band, and it is quiet — relative to the power spent on it — because its notes are inside one band. The orchestration rule about not scoring close thirds in the bass has always been justified by the muddiness. Half of the case is that the muddiness is not even loud.
Ninety players, and four times as loud
The other half of the debt was the orchestra, and it is the same arithmetic read in the direction that pays worst.
Independent players do not add their pressures — their phases are unrelated, so what adds is power. Ten players are ten times the intensity and ten decibels; ninety are 19.5 decibels. Then loudness goes as roughly the 0.3 power of intensity, which is what a sone is: ten decibels doubles it.
The conversion the section arithmetic runs on is the one the ladder’s first rung drew and did not spend: loudness in sones against loudness level in phons. Above 40 phons the curve doubles every ten phons, so a tenfold increase in power is a doubling in sones and no more — which is the whole reason a chord’s notes do not add up to a chord.
Ninety violins are 3.9 times as loud as one violin. Not ninety times, and not the twenty-odd times a decibel count might suggest to somebody reading decibels as loudness.
This is where the choir rung and the soloist rung meet the arithmetic. A trained soloist heard over a full orchestra is not louder than it and could not be — the section arithmetic above says what a single voice would have to do to compete on power, and the answer is impossible. What the voice does instead — and what the dynamic marking on the page cannot say — is move to a part of the spectrum the orchestra has vacated, which in the language of this rung is: occupy a band nobody else is in, where the loudness adds rather than the intensity.
And what a choir does that a soloist cannot is now a second thing. Sixteen singers on one note are 12 decibels and 2.3 times the loudness of one, which is a poor return; the essay’s finding was that what a choir buys is not level at all but the disappearance of a beat rate. The loudness arithmetic here says why the section was never going to be the way to buy volume.
The cheap way to be loud
The expensive way to get louder is more power in the same place. The cheap way falls straight out of the second clause of the rule.
Twelve tones sharing one tone’s power are 5.3 times as loud as that tone. Every one of them is 10.8 decibels quieter than the single tone was, and the total is five times the loudness.
That is the tutti, and it explains a thing about orchestration that is usually taught as taste. The way to make an orchestra louder is not to make the strings play harder; it is to add instruments in registers and colours that are not already occupied. A doubling of the whole string section buys 3 decibels and a quarter more loudness. Adding a piccolo, three trombones and a triangle buys four new bands, and four new bands at the same total power would be more than twice the loudness — the trombones being much more than a band away from the violins is the entire mechanism.
The corollary is the one every recording engineer knows in the other direction: a mix in which everything occupies the same two octaves is quiet for its power, and the remedy is not the fader.
Which computation produced the numbers
Two models, both of them named, both of them running on functions this site already had.
The grouping model sorts the components by frequency, walks the list, and starts a new group whenever the gap to the previous component exceeds one Bark. Inside a group the intensities are summed and the group’s level is converted once, at the group’s intensity-weighted mean frequency, through ISO 226 to phons and then through Stevens’s power law to sones. Across groups the sones are added. A single component returns exactly what this site’s existing loudness function returns for it, so the scale is the sone by construction and nothing is calibrated.
The excitation model replaces the group with a skirt. Each component spreads along the Bark axis at 27 decibels per Bark downward and a level-dependent 6 to 16 decibels per Bark upward — which is not a new function either, it is exactly the masking spread this collection has been drawing since the masking ladder opened. The excitations are added as intensities at every point on the axis, converted to a specific loudness by Zwicker’s compressive law, and integrated. One constant is fixed by the definition of the sone at 1 kilohertz and 40 decibels, and it is the only fitted number in either model.
Both models borrow the same spreading function, written for a different job: the level a probe needs in order to be heard beside a masker. Its downward slope is fixed and its upward one flattens with level, so a loud low note reaches further up the spectrum than a quiet one does — and that asymmetry is doing work here that nobody put in on purpose.
The two disagree in a way that is worth stating rather than hiding. The grouping model reproduces the sone scale exactly — 10 decibels doubles the loudness, at every level — because that scale is inside it. The excitation model does not: it grows by a factor of 2.3 per 10 decibels rather than 2.0, because its upper skirt flattens as the level rises, so a louder tone occupies more of the axis and gains loudness twice over. Every number quoted above as a ratio between two spectra at the same total power is one where that error is common to both sides and cancels to about a per cent. No number here is an absolute loudness, and that is why.
What this collection had been getting wrong
There is a version of this that is not a new rung but a correction, and it is in the docstring of the site’s own loudness function, which has said so since the day it was written: the loudness of a complex tone summed over its partials is an upper bound, because partials inside a single critical band do not add as separate loudnesses.
The error is small in the middle of the range and enormous at the bottom. In the treble a note’s partials are spread over many bands and counting them separately is nearly right; at C2 the first eight partials of a string are all inside the same wide band, the grouping model pools every one of them, and the published upper bound overstates the note by a factor of five and a half.
No figure on this site was wrong, because the function was only ever used for its slope against level, which the docstring says and which the grouping does not change. But the number it returns was being read as a loudness in the one register where it is furthest from one.
Whose music, and when
The claims here divide into two kinds and the division is worth marking, because this site’s rule is to say whose music a claim is about.
The band arithmetic is not about anybody’s music. It is a property of a listener, measured in the nineteen-thirties and refined ever since, and it applies to a Balinese gong-chime and a Mahler tutti in the same way.
The orchestration reading is a claim about a repertoire. The specific practice — a tutti built by adding families rather than by adding desks, low chords voiced open, the melody given to whatever instrument owns a vacant band — is European orchestral practice from roughly Berlioz onward, and it was arrived at by ear. Berlioz’s treatise of 1843 gives the rule about spacing low chords and gives no reason for it beyond the sound. What the arithmetic here adds is not the rule but a second reason for it, and the second reason is a different one from the roughness the triad ladder found.
The one place the two traditions visibly disagree is the register question. Gamelan writing puts dense, low, closely-spaced strikes exactly where European orchestration says not to, and the arithmetic here says those strikes are getting very little loudness for their power. That is a claim about volume rather than about value; the point of a low unison in that repertoire is not to be loud.
What the picture cannot show
The grouping is single-linkage, and that is a real defect rather than a simplification. A chain of components each nine-tenths of a band from the next is pooled into one group however long the chain, which is wrong for a dense cluster spanning several bands. The excitation model has no such defect, which is why both are drawn on every figure that turns on the transition.
Nothing here is heard. Loudness is a report from a listener, and the sone scale is a scale of reports. Every number above is what a model of those reports predicts, and the model was fitted to listeners doing a magnitude-estimation task that some psychophysicists regard as measuring the task rather than the ear.
The components are steady and simultaneous. Loudness needs time to build — a very short note is quieter than a long one at the same amplitude — and none of the figures above has a time axis at all. A chord struck and released inside a fifth of a second is not the chord these numbers describe, and neither is a chord whose notes do not start together, where the simultaneous window is a fraction of each note.
And the register result is stated at constant total power, which no performance holds constant. A pianist playing a low triad does not play it at the same power as a single note; they play it with three fingers. At constant power per note the low triad is louder than one note, by about the amount three notes ought to be — the finding here is about what the spending buys, not about what happens when more is spent.
Where this ladder goes next
Two rungs. The first says what one tone’s loudness is and that the map from decibels to it bends with frequency. This one says what happens when there is more than one tone, and finds that the answer is decided by a quantity that was already on this site for a completely different purpose.
The rung after it is the one this makes unavoidable, and it is the one the closure ladder asked for from the other end. An ending is a dynamic event — a diminuendo, a final chord louder than everything before it, a texture thinning — and the closure rung recorded that this collection has no model of dynamics in a form. It now has one, and the missing piece is no longer the loudness of a moment but the loudness of a moment against the moments around it: a fortissimo is a fortissimo relative to what a listener has been hearing for the last twenty seconds, and there is a published adaptation time constant for exactly that. Whether a crescendo of a stated size is heard as one depends on that constant, and it is a computation rather than a corpus.
Part 2 of 8
One essay in the series on loudness. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 17.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Critical bandwidthDecibelEqual-loudness contourLoudnessMaskingOrchestrationPhonSone
- An entrance is a change of colour critical bandwidth, loudness, masking, orchestration
- The same chord is harsher when it is louder critical bandwidth, equal-loudness contour, loudness, sone
- A bass chord low enough to balance has already hidden its tenor critical bandwidth, loudness, masking
- A part that leaves is not a part that arrives loudness, masking, orchestration
- A soft chord has to fade in loudness, masking, orchestration
- Room is used up by whoever enters first critical bandwidth, masking, orchestration