Harmony and voice leading

A loud chord is a smaller chord

Two earlier essays hold the level fixed, and the level decides how much of a chord a listener is given. At thirty decibels twenty of a triad's twenty-four partials stand above what the rest of it masks; at a hundred, nine do. Every roughness figure until now counts partials that are in the score, and a partial the chord masks is not a partial the listener has.

Assumes: One sound hides another, and it hides upward · The same chord is harsher when it is louder

One sound hides another draws the masking pattern for a single tone and finds it asymmetric: a masker hides much more above itself than below, and the asymmetry gets worse as the masker gets louder. A sound hides what came before it does the same on the time axis.

Both are drawn at a stated level, and the level is the parameter neither varies. That is the one to vary, and the reason is not local to this ladder.

Every roughness figure on this site counts partials that are present in the arithmetic. A chord’s roughness is computed by summing a kernel over every pair of partials in the score, and a partial that lies below the threshold the rest of the chord puts up is present in the arithmetic and absent from the ear.

A loud chord is a smaller chord. The share of a voicing's partials that stand above what the rest of it masks, and the share of its computed roughness that is between partials a listener actually has, from 30 decibels to 100. Both fall: 83 per cent of the partials survive at 30 decibels and 38 at 100, and the roughness share goes from 88 per cent to 67. The direction is the upward spread of masking, which grows faster than linearly with level: a loud partial masks a band above itself much wider than a quiet one does, so the chord's own top disappears into its own bottom. Two earlier essays are drawn at one level, and this is what they were holding.
Fig. 1 The share of a triad’s partials that stand above what the rest of the chord masks, and the share of its computed roughness that lies between partials a listener actually has. Both fall as the chord gets louder.

What a chord does to itself

Take a C major triad in close position, give every note ten partials falling as 1/n, and ask of each partial whether it stands above the summed masked threshold of all the others.

At thirty decibels a note, twenty of the twenty-four survive. The four that do not are the highest partials of the lowest note, which fall below the absolute threshold of hearing rather than below anything the chord does.

At seventy, nineteen survive.

At ninety, twelve.

At a hundred, nine — just over a third of the score.

The direction is the upward spread of masking, and the reason it accelerates is that the spread grows faster than linearly with level. A quiet partial masks a narrow band just above itself; a loud one masks a wide band, and each note’s own fundamental is loud enough at high levels to take out most of the partials of everything above it.

So a fortissimo chord delivers a smaller set of components than a pianissimo one. That is a genuinely strange sentence and it is the well-established behaviour of the ear.

What that does to the roughness

The count is interesting and the consequence for the roughness is the reason this rung reaches outside its own ladder.

Recomputing the roughness over only the audible partials gives, as a share of the roughness computed over all of them:

  • 88 per cent at thirty decibels,
  • 87 at seventy,
  • 67 at ninety and above.

So at a loud dynamic a third of the roughness this collection computes for a triad is between pairs of partials the listener does not receive.

That is a systematic overestimate, in a known direction, of a size that depends on the dynamic — which means it is not a constant that cancels when two chords are compared. Two voicings at the same level lose different amounts, and the same voicing at two levels loses different amounts.

The same chord is harsher when it is louder is the rung on the consonance ladder that first put a level into a roughness calculation, and it found the roughness rising with level because amplitude enters the kernel quadratically. This rung finds a term pulling the other way and it does not cancel: the masking correction is at most a factor of two thirds and the amplitude term is a factor of the amplitude squared, so the loud chord is still very much rougher. What changes is by how much, and that is one multiplication:

level partials audible roughness kept rise from 30 dB, uncorrected corrected
30 dB 20 of 24 88% ×1 ×1
70 19 87% ×10⁴ ×9,840
90 12 67% ×10⁶ ×758,000
100 9 67% ×10⁷ ×7.58 million
110 2 of 24 26% ×10⁸ ×29.4 million

The correction bends the curve rather than scaling it. Over the ordinary range from thirty decibels to a hundred the roughness rises by a factor of ten million uncorrected and seven and a half million corrected — a quarter less — which in exponent terms takes the level dependence from exactly 2.00 to 1.97. Push to a hundred and ten and it falls to 1.87.

So the answer to “by how much” is: very little, until it is a great deal. Across every dynamic an ensemble ordinarily uses the correction is a few per cent of an effect that spans seven orders of magnitude, and the collection’s level-aware roughness figures are near enough right. Above about ninety-five decibels the masking starts removing partials faster than the amplitude adds roughness, and the two terms are within a factor of four of each other by a hundred and ten.

The last row is worth reading on its own. At a hundred and ten decibels a close-position triad with ten partials a note delivers two audible components. That is a chord which has, as far as the ear is concerned, almost no interior at all — one loud thing and one other — and it is the loudest a listener is ever asked to hear a chord at. Whatever a fortissimo tutti is doing to a listener, it is not delivering a spectrum, and no roughness computed over a score can be a description of it.

The partials a chord delivers and the partials it hides. Every partial of a 3-note voicing at 90 decibels, against the threshold the rest of the chord puts up at that frequency. A partial below the line is in the arithmetic and not in the ear. 12 of 24 stand above it, and the roughness computed over the ones that survive is 67 per cent of the roughness computed over all of them. That second number is the one that reaches furthest: every roughness figure here counts partials that are present in the score, and a partial the chord masks is not a partial the listener has.
Fig. 2 One chord at ninety decibels, partial by partial, against the threshold the rest of it puts up. The partials below the line are in the score and not in the ear, and the roughness computed over them is roughness nobody hears.

Spacing decides more than level

Holding the level and varying the voicing gives a larger effect than varying the level did, and it lands on a rule every orchestration manual states.

At seventy decibels a note:

voicing audible roughness delivered
close, low (C3 E3 G3) 16 of 24 72%
close, middle (C4 E4 G4) 19 of 24 87%
close, high (C5 E5 G5) 17 of 24 67%
open, over two octaves (C3 E4 G5) 16 of 24 40%
the orchestral spacing (E2 E4 G4 C5) 20 of 32 51%

The open voicing delivers forty per cent of its computed roughness. A chord spread over two octaves has a bass whose upper partials sit exactly where the middle voice’s fundamental is, and the bass at seventy decibels masks a great deal of what the middle and upper voices contribute.

That is the ordinary rule about spacing — put the chord’s notes further apart at the bottom — arriving from a direction it is never derived from. The usual account is that a wide voicing is smoother because its notes are further apart in critical-band terms, and where to put the third is this collection’s version of that argument.

This is a second mechanism on top of it. A spread voicing is smoother partly because its pairs are further apart and partly because its own bass has removed the partials that would have been rough.

The orchestral spacing — a low bass, a gap, and a close upper triad — delivers 51 per cent, which is the second lowest here. That is the spacing an orchestrator is taught, and it turns out to be a spacing that hides half its own roughness from the listener.

How much of a voicing's roughness the voicing delivers. Five spacings of a chord at 70 decibels, with the share of each one's computed roughness that lies between partials the listener is actually given. The close middle voicing delivers 87 per cent; the widest delivers 40. That is the ordinary orchestration rule arriving from a direction it is never derived from — a spread voicing is not smoother because its partials are further apart, it is smoother because its own bass masks the roughness its top would otherwise have. The best-delivering spacing here is close, middle at 87 per cent.
Fig. 3 Five spacings at one level, with the share of each one’s computed roughness that reaches a listener. The widest delivers forty per cent, and the rule about spacing acquires a second mechanism.

Which partials go, and in what order

The count is a summary and which partials it removes is the part with the musical content.

The survivors are always the fundamentals and the low partials of every note. A fundamental is the loudest component of its own note and it is below everything the other notes have much energy at, so nothing masks it — the asymmetry of masking protects the bottom of the spectrum comprehensively.

What goes first is the upper partials of the lowest note. At quiet levels they go because they are below the absolute threshold; at loud levels they go because the middle voice’s fundamental has arrived and is masking them. Either way the bass’s own colour is the first thing a chord loses.

Then, as the level rises, the upper partials of the middle voice, masked by the top voice’s fundamental and lower partials.

So a loud chord is not a chord with its top removed, which is what a low-pass would do. It is a chord in which each note’s upper spectrum is taken away by the note above it, leaving something much closer to a set of pure tones than the score contains — which is a specific and rather striking prediction about what a fortissimo chord is.

It also explains why the effect is so much larger for the open voicing. A chord spread over two octaves has each note’s fundamental sitting in the middle of the note below’s partial series, which is exactly where a masker does the most damage.

The same machinery answers a question this essay has so far only asked of chords, which is what happens when the notes are not simultaneous at all.

Nothing at all until fifteen decibels, and then it depends on the tempo. The fraction of a line's partials that stay above threshold, over how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement — forward masking leaves a threshold at most ten decibels below the masker, and a note's own partials mask each other from the same components at full level. The boundary is between twelve and eighteen decibels, and beyond it the loss grows with the tempo: at 280 to the crotchet and 36 decibels of contrast, 26 per cent of the line's partials are gone. Fifteen decibels is about the gap between a forte and a piano.
Fig. 4 The fraction of a line’s partials that stay above threshold, against how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement.

The white region is the useful half of it. Masking between successive notes needs a level difference, not merely a fast tempo — so the effect this essay is about belongs to simultaneity and to dynamics, and a fast line at an even dynamic loses nothing to it. That is a sharper boundary than the chord figures could draw, because they hold everything at one instant and cannot show what the axis does.

a major triad, spread by 40 milliseconds. Three notes of a major triad, each lasting 600 ms, with their onsets 40 ms apart. The shaded band is the window in which all three are sounding — 520 ms, which is 87 per cent of a note's length. The chord exists as a simultaneity only inside that band; before it and after it the passage is a melody. At 40 ms the notes are further apart than the asynchrony at which a mistimed partial leaves its note.
Fig. 5 Every voicing of one triad within a span, scored for roughness. This essay says that the roughness computed here is delivered to a listener in different proportions by different voicings, so the ordering is not quite the ordering a listener would produce.

Why the level was held, and what holding it cost

It is worth saying why two rungs of this ladder are drawn at a fixed level, because the reason is a good one.

Masking is usually studied as a relation between two sounds — one masker, one probe — and the classical figures are of a probe’s threshold against a masker of stated frequency and level. That is a two-sound experiment and there is nothing to vary except the masker’s level, which the first rung does: it draws the pattern at forty, sixty and eighty decibels and the asymmetry growing between them.

What the first two rungs do not do is put a chord into the pattern. A chord is many maskers at once, each masking the others, and the level then decides not how far one pattern spreads but how much of the whole set survives its own company. That is a different question and it needs the level to be a variable rather than a case.

A sound hides what came before it has the same shape on the time axis and the same gap, and the equivalent calculation there — how much of a passage a listener receives, given that each event masks the ones near it — is a rung this ladder still owes.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward.
Fig. 6 The earliest figure of all: the classical pattern for one masker at three levels, with the upward spread growing as the masker gets louder. Every partial in this essay is a masker of that shape, and there are twenty-four of them at once.

What this asks of the rest of the collection

The correction is a third of the roughness at a loud dynamic, which is large enough to matter to several results elsewhere and not large enough to overturn any of them.

Roughness can be computed and everything built on it produce orderings — which interval is rougher, which voicing is smoothest — and an ordering survives a multiplicative correction that is roughly common to the things being ordered. Where the correction differs between them, as it does between voicings, the ordering can move.

The place it is most likely to move is where to put the third, which ranks voicings, and the voicing table above says the ranking’s inputs differ by a factor of two in how much of themselves they deliver. That is a real threat to a result rather than a footnote, and the honest statement is that the ranking has not been recomputed with the correction applied and should be.

The places it is least likely to move are the ones comparing intervals at a fixed spacing and level, which is most of the consonance ladder. Two notes a semitone apart and two a fifth apart at the same register mask each other similarly, so their computed roughnesses are reduced by similar amounts.

That is the general shape: a correction that varies with the thing being varied is dangerous, and one that does not is a scale factor.

Which computation produced the numbers

Each note contributes ten partials at levels falling as 1/n from the stated level. For each partial, the masked threshold is the energy sum of the thresholds every other partial imposes at that frequency, from the spreading function the first rung uses — the standard asymmetric pattern with a level-dependent upper slope.

A partial is audible if its own level exceeds both that masked threshold and the absolute threshold of hearing at its frequency, which is what removes the top partials of a quiet bass note.

The roughness is the Plomp–Levelt kernel summed over pairs of partials belonging to different notes, weighted by both partials’ amplitudes, with the kernel written out rather than borrowed because every other roughness function in this collection takes a spectrum and this takes a list of partials that already have their own levels.

Where the model stops

A partial is either audible or not. Real partial masking is graded — a partial a few decibels above its masked threshold contributes less than one well above it — so a binary count is a caricature of a continuous reduction, and it overstates the effect at the boundary and understates it above.

Masking is applied within one instant. The spreading function is a simultaneous-masking pattern, and a chord’s partials all begin at different times and decay at different rates, so a real chord’s masking pattern is a moving thing. The partials that do not die together is the essay about that variation.

The spectra are 1/n. A real instrument’s spectrum is not, and the amount of self-masking depends heavily on it: a spectrum with strong high partials masks more of the chord above it and delivers less.

And the roughness kernel and the masking model are from different literatures. Combining a sensory-dissonance model with a spreading function is a reasonable thing to do and it is not something either was fitted for. Whether the correct account of a masked partial’s contribution to roughness is zero is not established.

What the picture cannot show

It cannot show what a listener judges. Every number here is a share of a computed quantity, and whether a chord whose roughness is two thirds hidden is heard as two thirds as rough is a listening question this collection cannot ask.

Nor can it show the room. A hall puts a comb on every spectrum and adds a reverberant tail whose own energy masks — so a chord in a hall is masked by itself and by its own past, and the second is not modelled here at all.

It cannot show the ear’s own distortion. At the levels where this effect is largest, the ear makes its own sound — combination tones generated inside the cochlea that are not in the signal at all — so a loud chord loses partials it had and gains ones it did not, and only the first is modelled.

1000 and 1200 hertz, and what the ear addsTwo tones presented to a listener, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions: they are products of the pair, at f₂ − f₁ = 200 Hz, 2f₁ − f₂ = 800 Hz, 3f₁ − 2f₂ = 600 Hz, 2f₂ − f₁ = 1400 Hz. The cubic difference tone sits just below the lower primary and is audible at modest levels; the quadratic one is far below both and needs a loud pair.f₂ − f₁2003f₁ − 2f₂6002f₁ − f₂800f₁1000f₂12002f₂ − f₁1400hertzin the air:two tonesin the ear:4 moref₂/f₁ = 1.20the cubic product islargest near 1.22
Fig. 7 The components a loud pair adds that were never played. This essay counts what a loud chord takes away and this is what it puts back, and the two have never been computed together.

And it cannot show attention. Masking measured with a listener attending to a probe is not the same as masking in a musical texture, where a listener following a line hears things a threshold measurement says are inaudible. That is the standard caution about applying threshold data to music and it applies with full force.

Whose chords, and when

The voicings are the ones a European tonal practice uses: close position in three registers, an open spacing over two octaves, and the standard orchestral layout with a low bass under a close upper triad. The last of these is what four-part orchestral writing has looked like since the classical period, and the rule that produces it — spread at the bottom, close at the top — is in every treatise.

The dynamics span thirty to a hundred decibels a note, which is roughly a soloist’s pianissimo to an orchestral fortissimo at the listener’s ear. The interesting part of the range for this argument is the top of it, and that is where the least of the score is delivered.

There is a reading of that which is a claim about orchestration rather than about hearing, and it is offered as a suggestion rather than a finding. A loud tutti chord is heard as massive rather than as harsh, and one reason may be that a third of its computed roughness never arrives — the ear having removed exactly the high-partial pairs that a roughness model finds most objectionable.

Where this ladder goes next

Three rungs. One sound hides another, and it hides upward; a sound hides what came before it; and now a chord hides itself, by an amount that depends on how loud it is and how it is spaced.

What the ladder owes next is the passage. Both of the earlier rungs and this one are about one instant, and forward masking has a time constant of a couple of hundred milliseconds — long enough that a note masks its successors as well as its neighbours. A melodic line at a moderate tempo is a sequence of events well inside each other’s forward-masking windows, and how long a note has to be is the collection’s account of the durations involved. Putting the two together would say how much of a fast passage a listener receives, which is a question about texture that this ladder has the arithmetic for and has never posed.

Part 3 of 8

One essay in the series on masking. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Critical bandwidthDynamicsMaskingPartialRoughnessSpreading functionThreshold of hearingVoicing