Harmony and voice leading

The ranking survives the dynamic and the chord does not

Roughness is quadratic in pressure and loudness is compressive, so sixty decibels multiply a chord's roughness by a million and its loudness by ninety-six. Roughness per sone therefore rises ten thousandfold between a pianissimo and a fortissimo of the same four notes — and yet the ranking of which doubling is smoothest, over four hundred and eighty voicings, does not move by a single place.

Assumes: Who plays what and how loud is one question · The same chord is harsher when it is louder

The first rung of this anchor held every scoring at one total loudness on purpose, because roughness is quadratic in pressure and a comparison that let one candidate be quieter would have found the quiet one and stopped. That was a device for making a comparison honest. It also postponed the question it was hiding, which is what happens when the loudness itself moves.

The answer has two halves and they point in opposite directions, which is why this is a rung rather than a footnote. One half is enormous and the other is exactly zero.

Roughness and loudness do not rise together. One four-note chord, played at levels from 35 to 95 decibels, with both quantities drawn as multiples of what they are at the quietest. Roughness is quadratic in pressure, so 60 decibels multiply it by 1.0e+6. Loudness is compressive — about ten phons to a doubling of sones — so the same range multiplies it by 96. The gap between the two lines is the quantity: roughness per sone rises by a factor of 1.0e+4 between a pianissimo and a fortissimo of the same chord.
Fig. 1 One four-note chord over sixty decibels, with roughness and loudness each drawn as a multiple of what it is at the quietest. The two lines are both straight and they have different slopes, which is the whole of the arithmetic: roughness goes as pressure squared and loudness goes as roughly the six-tenths power of it. Sixty decibels multiply the first by a million and the second by ninety-six.

The half that is enormous

A decibel is a ratio of pressures. Plomp and Levelt’s roughness is a sum over pairs of partials of the product of their two amplitudes, so it is quadratic: raise every partial by twenty decibels and the roughness goes up by a factor of a hundred. Over the sixty decibels between a soft chamber dynamic and a loud orchestral one, that is a factor of a million, and the model says so exactly — the figure’s straight line has slope two on a decibel axis and nothing in it is fitted.

Loudness does not do that. It is compressive, in the way the quietest thing audible sets out: ten phons buys a doubling of sones, so ten decibels roughly doubles the loudness where it multiplies the roughness by ten. Over the same sixty decibels the chord goes from just over one sone to about a hundred, a factor of ninety-six.

90 players, and how much louder than one. Independent players on one note add in power, so 90 of them are 19.5 decibels above one — and loudness in sones goes as roughly the 0.3 power of intensity, so the section is 3.9 times as loud rather than 90 times. The horizontal axis is doublings of the section, which is why the line is nearly straight: every doubling is 3 dB and about a quarter more loudness.
Fig. 2 Why the second line is not the first. Loudness sums compressively over sources as well as over levels — ninety players are not ninety times one player — and the same compression is what flattens the loudness curve against level. Roughness has no such term anywhere in it: every pair of partials contributes its own product and nothing groups them.

Divide the two and the quantity that comes out is roughness per sone, and it rises by a factor of ten thousand between the same chord played very quietly and played very loudly.

That is a large enough number to be worth restating in words. A chord is not a fixed amount of harsh that gets louder. Its harshness relative to how loud it is — which is the only version of harshness a listener has any access to — is four orders of magnitude greater at the top of a dynamic range than at the bottom. The same chord is harsher when it is louder is the essay that established the direction; this is the size of it, in the units a listener is actually in.

It is also the arithmetic behind a very ordinary piece of practice. A close-spaced chord low in the piano is a routine sound at pianissimo and an unusable one at fortissimo, and no rule of harmony distinguishes the two cases because harmony has no dynamic in it at all.

The half that is exactly zero

The obvious next step is to ask whether the ranking moves — whether a doubling that is smoothest when quiet stops being smoothest when loud. There is a mechanism that could do it, and it is not a small one: a partial below the threshold of hearing contributes nothing to roughness, not a little, and which partials are below threshold depends on the level. Two voicings whose roughness comes from different parts of the spectrum could plausibly change places as the quiet one’s contributors switch off.

They do not. Every four-part voicing of a major triad inside the four standard compasses — four hundred and eighty of them — grouped by which chord member is doubled and ranked by mean roughness, gives the same ranking at every level from thirty-five decibels to ninety-five. Doubled root, then doubled fifth, then doubled third, separated by one and a half per cent and nine per cent, at each of seven levels. There is not one inversion.

The doubling ranking, asked again at every dynamic. Every four-part voicing of the chord inside the four standard compasses — 480 of them — grouped by which chord member is doubled, and each group's mean roughness drawn as a multiple of the smoothest group's at that level. The absolute roughness rises by a factor of 1.0e+6 across the range drawn and the ranking does not move at all: root 1.000, fifth 1.013, third 1.090 at the quietest, and 1.000, 1.013, 1.090 at the loudest. There are no inversions: not one pair of doublings changes places.
Fig. 3 The doubling ranking asked again at every dynamic. The vertical axis is each group’s mean roughness as a multiple of the smoothest group’s at that level, so the millionfold scaling has been divided out and only the ranking is left. The three lines are flat and they do not cross. The rule that says double the root is a rule about a ratio, and a ratio is what survives a scaling.

This is a negative result and it is the useful one. It says why part-writing can be taught without dynamics at all, which is a fact about the pedagogy that has never had a reason attached to it. Why the exercise is in four parts is about one arbitrary-looking convention having a computable justification; this is another. A doubling rule is a ranking over arrangements of one chord, and a ranking is invariant under a monotone transformation of the quantity it ranks. Multiply every roughness by the same million and nothing moves.

The argument is good and it proves less than it looks, which a later section takes apart: a ranking is invariant under a monotone rescaling, but the threshold cutoff is not a rescaling — it removes terms — so the invariance found here is an empirical result about one spectrum rather than a consequence of the algebra.

An orchestration rule is not like that. A balance is a magnitude — so many decibels for this part — and a magnitude has no such protection. That is the difference between the two disciplines stated in one line, and it falls out of the units rather than out of anybody’s taste.

Where the level does change something

The mechanism that failed to reorder the doublings has not gone away; it was looked for in the wrong place. Partials do drop below threshold, and the level at which each one goes is not what the obvious picture suggests.

In the treble, the highest partials go first. That is expected: a 1/n spectrum puts less into partial twelve than into partial two, so partial twelve is the first to run out.

In the bass it is the other way round, and it is the other way round by a wide margin. The threshold of hearing rises about fifty decibels between five hundred hertz and thirty, which is far faster than a 1/n spectrum falls. A cello’s bottom C at sixty-five hertz needs thirty-nine decibels before its fundamental is audible at all, and only twenty-four before its eighth partial is. Fifteen decibels of the quietest part of the dynamic range are spent on a note whose fundamental is inaudible and whose upper partials are not.

The level at which each partial becomes audible at all. Each point is one partial of one note, drawn at the whole-note level below which that partial is under the threshold of hearing and contributes nothing — not a small amount, nothing. C2 needs 39 dB before its FUNDAMENTAL is audible and 24 before its eighth partial is; E4 needs 10 dB before its FUNDAMENTAL is audible and 15 before its eighth partial is. In the treble the highest partials go first, which is expected. In the bass they do not: the threshold of hearing rises about fifty decibels between 500 and 30 hertz, far faster than a 1/n spectrum falls, so a quiet bass note loses its fundamental and keeps the partials that were supplying the roughness. At 20 dB the bass has 0 of 16 partials and the treble note has 12.
Fig. 4 The level at which each partial of two notes crosses the threshold of hearing, so that below the line it contributes nothing. The treble note’s curve rises with partial number, which is the expected shape. The bass note’s falls for the first several partials and then rises, so its fundamental is the last thing to become audible rather than the first — and a pianissimo bass is a set of upper partials with a hole where its note should be.

The chord whose root is a residue

The consequence of that is not about roughness at all, and it is the surprising connection this rung has.

A bass note played softly enough is a spectrum with its lowest components missing. The pitch a listener assigns to such a thing is the subject of a whole ladder here: it is a residue pitch, supplied by the ear from the spacing of the partials that are present, and the note that is not there is its first rung. Nothing in the missing-fundamental ladder ever mentioned a dynamic, because every stimulus in it has its fundamental removed by construction.

Here the fundamental is removed by playing quietly. The bass of a soft low chord is being heard the way an organ’s resultant bass or a small loudspeaker’s bottom octave is heard — as a pitch computed from a pattern rather than received as a tone. And the partials the ear is using to compute it are exactly the partials that were supplying the chord’s roughness, since which harmonics carry the pitch puts the dominance region between the third and the fifth partial, and the cutoff figure says those are the ones that survive.

So a pianissimo bass line is doing two jobs with one set of components, and the two jobs are the ones this collection has been treating as belonging to different fields.

What 4 partials imply, and how many answers there are. The partials are at 196.2, 261.6, 327.0, 392.4 Hz. Each row is a harmonic series they are consistent with to within 30 cents: the filled dots are the partials in their assigned slots and the open dots are slots the series predicts that nothing occupies. There are 2 such series with harmonic numbers up to 16, the best fitting them to 0.0 cents with a fundamental of 65.40 Hz and no unoccupied slot. The rest sit at one half of it and their arithmetic is exactly as exact; what separates them is the count of empty slots, which is why a residue pitch built from few partials is reported an octave out by a minority of listeners and not by the rest.
Fig. 5 What the ear has left to work with when the bottom of a low note has gone under the threshold: partials three, four, five and six of a sixty-five hertz fundamental, and the fundamentals that would fit them. The best fit is the note that is not sounding, and the fit is exact — which is why a soft bass line does not sound as though it has lost anything.

The check that the invariance is about the level rather than about one chord is to ask the same question of a minor triad, whose third sits a semitone lower and whose partials therefore collide differently.

The doubling ranking, asked again at every dynamic. Every four-part voicing of the chord inside the four standard compasses — 392 of them — grouped by which chord member is doubled, and each group's mean roughness drawn as a multiple of the smoothest group's at that level. The absolute roughness rises by a factor of 1.0e+6 across the range drawn and the ranking does not move at all: fifth 1.000, root 1.037, third 1.110 at the quietest, and 1.000, 1.037, 1.110 at the loudest. There are no inversions: not one pair of doublings changes places.
Fig. 6 The doubling ranking for a minor triad. The gaps between the three groups are not the same gaps as the major triad’s — a minor third puts a different pair of partials near each other, and the ordering it produces is its own. What is the same is that the three lines are flat: over sixty decibels, and a millionfold change in the quantity being ranked, nothing changes places.

The two chords disagree about which doubling is smoothest and agree that the answer does not depend on the dynamic. That is the shape a genuine invariance has, and it is a stronger statement than either chord alone: the ranking is a property of the chord, the scaling is a property of the level, and the two do not interact.

Tempo is the axis this essay has not varied, and it is the one that breaks the symmetry between the two quantities.

One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics.
Fig. 7 How much of each quantity’s contrast between chords a listener still has by the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not.

One contrast survives every tempo anybody plays and the other does not. The roughness ranking is intact until chords are arriving faster than twenty a second; the loudness contrast is being eroded at every tempo, because its window is two seconds and no chord is. That is the same asymmetry the rest of the essay found in level and chord, arriving through the clock.

Where the invariance does break

A negative result is worth exactly as much as the attempts made to break it, and the caveat at the foot of this essay names three ways it might: more partials, a different spectrum, a chord spread wider than a choir. All three are one argument away, so all three were tried.

condition ranking at 35 dB ranking at 95 dB inversions
string, 8 partials (as shipped) root 1.000, fifth 1.013, third 1.090 identical 0
24 partials at 1/n root 1.000, fifth 1.013, third 1.081 root, fifth 1.014, third 1.082 0
clarinet, weak even partials root 1.000, fifth 1.000, third 1.064 root, fifth 1.002, third 1.067 0
bell, no even partials fifth 1.000, root 1.005, third 1.067 identical 0
compasses widened an octave each way, 2,007 voicings root 1.000, fifth 1.000, third 1.052 root, fifth 1.001, third 1.055 0
16 partials, amplitude the inverse square root of partial number root 1.000, fifth 1.001, third 1.064 fifth 1.000, root 1.001, third 1.066 1

Five of the six hold and the sixth does not, and the sixth is the informative one. A spectrum bright enough that its sixteenth partial is a quarter of the amplitude of its first — rather than a sixteenth, as a 1/n spectrum gives — brings doubled root and doubled fifth to within a tenth of a per cent of each other, and at that separation the two change places between the quietest level and the loudest.

So the exact statement is narrower than “the ranking survives”, and it is the statement worth carrying. Doubled third is last in every one of the six conditions, by between five and nine per cent, and never moves. That part is unconditional. The gap between root and fifth is 1.3 per cent for a string spectrum, 0.5 for a bell, and 0.1 for a bright one, and at a tenth of a per cent it is not a ranking at all — it is a tie, and a tie is decided by whichever way the level-dependent threshold effects happen to lean.

The bell row makes the same point from the other side. Its ranking is different — doubled fifth is smoothest, not doubled root — and it is just as invariant. So the ranking is a property of the spectrum and the invariance is a property of the arithmetic, and the two are independent: which doubling wins depends on the timbre, and that it keeps winning as the level rises does not.

Widening the compasses by an octave in each direction, which quadruples the voicing count to 2,007, changes nothing at all except to compress root and fifth toward each other. That is the third stressor the caveat named, and it is the one that mattered least.

Which computation produced the numbers

The roughness at a level is dissonanceAtLevel, which the consonance ladder’s tenth rung introduced and which differs from the ordinary roughness sum in exactly two ways: the partial amplitudes are pressures in pascals rather than a spectrum normalised to one, so the sum is quadratic in level as Plomp and Levelt’s product of two amplitudes says it must be; and a partial under the threshold of hearing is dropped rather than made small. A chord’s roughness is that sum over every pair of its notes.

The loudness is bandLoudness — every partial of every note placed on the Bark axis, components within one critical band summed as a single band rather than as separate loudnesses, each band converted to sones through the equal-loudness contours. It is the same function the dynamics are in the score already reads a page with.

The voicings are the site’s own chordVoicings inside the four SATB compasses, which is the enumeration where to put the third and the voice-leading ladder both use. Grouping is by which chord member appears twice, and the number reported per group is the mean over the group rather than its best member — a property of the texture rather than of one arrangement of it, for the reason the loudness ladder’s fourth rung had to discover the hard way.

The threshold curve is hearingThreshold, the standard minimum audible field, and the level at which partial n crosses it is computed rather than read off a figure: a whole-note level L puts partial n at L plus twenty times the log of its amplitude over the spectrum’s root-mean-square, so the crossing level is the threshold at that frequency minus that offset.

Where the model stops

One million is the model’s number and not the world’s. The roughness sum has no saturation in it. Real sensory dissonance does not grow without limit as a sound gets louder — nothing perceptual does — and every published roughness model is calibrated over a modest range of levels. What the figure asserts safely is the ordering and the shape: harshness grows much faster than loudness, which is a statement about two exponents, and the exponents are the ones the models state.

The threshold used is the minimum audible field, which is a threshold for a tone in silence. In a chord, every component is partly masked by the others, so the level at which a partial stops contributing is higher than the figure says — one sound hides another is the ladder about it. The bass-note inversion is not an artefact of that, because masking would raise both ends of the curve together, but the absolute levels are lower bounds.

And nothing here is about a room. Every level is at the listener. A fortissimo in a hall arrives with its reverberant tail, which adds components at every frequency and cannot be a simple gain on the direct sound.

What the picture cannot show

It cannot show that ten thousand is audible. The ratio of roughness to loudness is a ratio of two model outputs, and no experiment in this collection’s reach says a listener is sensitive to it, or to what power of it. What can be said is that the two quantities separate, by a lot, and that any account of a chord that reports one number for it is reporting a number that depends on a variable the account does not contain.

Nor can the ranking figure prove a negative, and one of the ways it might fail turns out to be real: the section above finds a spectrum under which doubled root and doubled fifth change places, and the reason is that they were tied to a tenth of a per cent before the level moved. What survives every condition tried is the last place rather than the first. An inharmonic spectrum is the one stressor not tested, because the roughness machinery here places partials at whole multiples of a fundamental and cannot express one.

And the cutoff figure is about one note at a time. A partial’s audibility in a chord depends on the whole chord, and the curves are computed for each note alone.

Whose music, and when

The doubling ranking is a claim about four-part writing in the European common practice, which is the repertoire the SATB compasses and the doubling rules both come from. Doubled root, then fifth, then third, is the ordering the textbooks give, and the arithmetic reproduces it — which is the useful direction of agreement, since the rule was arrived at by ear and the computation was not.

The dynamic claim is a claim about scoring, and it belongs to the repertoire in which a dynamic is written down and expected to be large. That is roughly Beethoven onward: an orchestral fortissimo is a nineteenth-century sound, and the range between the marks widens through the century as the instruments and the halls make it possible. A ratio that moves by four orders of magnitude across a dynamic range only matters where the dynamic range is four orders of magnitude wide, and for a good deal of the music the four-part rules were written for, it is not.

Where this ladder goes next

Two rungs. The problem is a constrained optimisation whose halves do not separate, and now the level has been let move: it changes the magnitude of everything and the order of nothing.

The rung after it is the one the second half of this rung’s own arithmetic names. Every number above treats a player’s spectrum as a fixed shape with a level in front of it, so that raising a part by ten decibels raises every one of its partials by ten. That is false for every instrument in an orchestra except one, and the exception is instructive: an organ flue pipe really does have a shape and a level, because its dynamic was fixed at the voicing bench and the player has no dial to turn. Everybody else changes what they contribute as well as how much, and the direction is computable.

Part 2 of 14

One essay in the series on orchestration. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DoublingLoudnessOrchestrationResidue pitchRoughnessSensory dissonanceThreshold of hearingVoicing