Perception and the listener

The listener is given the top voice, and the bass as a sine

Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.

Assumes: An equal note cannot be masked · A loud chord is a smaller chord

An equal note cannot be masked closed by naming what all four rungs of this ladder have in common:

Every rung above puts the masker and the probe in the same voice — the same chord, or the same melodic line — and the case that made this anchor worth opening is the other one: a quiet inner part under a loud outer one, sounding at the same time, where the spreading function and the forward decay both apply and the levels are set by the scoring rather than by a marking.

That is the case a listener meets all the time and this ladder had never computed. It needs one change to the arithmetic: keep the count of surviving partials per part rather than for the sonority as a whole, so the question stops being how much of a chord arrives and starts being how much of each voice does.

The top voice arrives whole and the bottom one arrives as a sine. 4 parts sounding together, each of 8 partials, with every partial tested against the summed masked threshold of every component in the texture. A filled mark is a partial the listener receives and an open one is a partial the part would have had alone and does not have here. The bass at C3 keeps 1 of 8, the tenor at C4 keeps 2 of 8, the alto at E4 keeps 5 of 8, the soprano at G5 keeps 8 of 8, every one of them at 70 decibels. Every part is at the same level and the difference is entirely where each one sits: masking spreads upward, so the part at the top of the texture has nothing above it to be masked by and the part at the bottom has everything.
Fig. 1 Four parts sounding together at seventy decibels each, with every partial tested against the summed masked threshold of every component in the texture. Filled: the listener has it. Open: the part would have had it alone.

Four parts, one level, and four completely different spectra

Take a four-part texture in the spacing four-part writing has used for three centuries — a bass on C3, a tenor on C4, an alto on E4 and a soprano on G5 — hold every one of them at seventy decibels, and count what each part is left with.

The soprano keeps all eight of its partials. The alto keeps five. The tenor keeps two. The bass keeps one.

The bass arrives at the listener as its fundamental and nothing else, which is to say as a sine tone. Every player is at the same level, every part has the same spectrum in the score, and the four of them are delivered as four different objects.

That is a larger effect than anything else on this ladder. A loud chord is a smaller chord found a triad losing about a fifth of its partials at seventy decibels and half at a hundred, which is a real number and a moderate one. Per part, at one moderate dynamic, the range runs from nothing lost to seven eighths lost.

And the mechanism is not the one the word “inner” suggests

The obvious reading is that the inner parts are covered by the outer ones, and it is wrong in a specific and useful way.

Take the alto out and put it under each of the other three in turn. Under the bass alone it keeps eight of eight. Under the tenor alone, eight of eight. Under the soprano alone it keeps five — exactly what it keeps in the full texture. The soprano accounts for the whole of the alto’s loss and the two parts below it account for none of it.

So a part is masked by the parts above it and by nothing else, and the reason is the one the first rung of this ladder established. One sound hides another, and it hides upward: the masked threshold falls at twenty-seven decibels per Bark below a masker and at six to sixteen above it, depending on level, so a masker reaches a long way up the spectrum and hardly at all down it.

Now put that together with what a note is. A part’s own upper partials are not in its own register; they are in the register of the parts above it. The alto’s third partial is at 990 hertz, its eighth at 2,640 — the soprano’s territory — and the soprano’s fundamental is sitting underneath all of them, masking upward into every one.

A part is not masked by the parts above it as objects. Its partials are masked by their fundamentals, and the higher a part is, the fewer fundamentals there are above its own partials to do it.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward.
Fig. 2 The two fundamentals doing the work, established earlier: the bass’s at 131 hertz and the soprano’s at 784, at the same level. The soprano’s curve covers the whole region the alto’s and tenor’s upper partials live in, and the bass’s covers almost nothing.

That reframing matters for what counterpoint can and cannot do about it. Giving the inner voice a more interesting line does not change any of this, and neither does giving it to a soloist. What changes it is the level and the register of everything above it.

Balancing the loudness changes nothing that matters

The debt asked for the case where the levels are set by the scoring, so the next thing is to set them that way and look again.

The orchestration ladder states the constraint a scoring is solved under, and it is a loudness constraint: each part’s share of the total loudness within a tolerance of a stated target, which for an even texture means every part contributing a quarter. Who plays what, and how loud, is one question searches the assignment under that constraint; here the assignment is fixed and only the levels are solved for.

The answer is bass +3.6 decibels, tenor −0.6, alto −0.3, soprano −2.3. That is very close to what a conductor asks for, and it moves the loudness shares to exactly a quarter each.

It moves the spectra to one of eight, four of eight, six of eight and eight of eight.

Equal loudness is not equal spectrum, and cannot be made so. The levels at which every part of a 4-part texture contributes the same share of the total loudness — the balance a conductor asks for and the constraint the scoring problem is solved under — with the offsets from a flat 70 decibels printed under each part. The pale bar is that loudness share, equal by construction at 25 per cent. The dark bar is the share of its own spectrum the part is left with, and it runs from 13 per cent for the bass to 100 for the soprano. Balancing the loudness costs the bass a rise of 3.6 decibels and buys it nothing at all in spectrum.
Fig. 3 The levels at which every part carries the same share of the loudness, with the offsets from a flat seventy decibels printed under each. The pale bars are equal by construction. The dark bars are what each part keeps of its own spectrum.

Equal loudness is not equal spectrum, and the balance that produces the first is not even pointed at the second. The bass gains three and a half decibels, which is a real instruction to a real player, and it buys the bass exactly nothing: still one partial of eight.

This is not a defect in the scoring problem. It is a statement about what a loudness constraint can express. Loudness is a scalar per part and the thing counterpoint cares about — whether the line is there, as a voice with a colour — is a set of partials. There is no assignment of levels that equalises both, because the two quantities move together only for parts in the same register, and parts in the same register are not what a texture is.

What it costs a part to be heard as itself

Since a balance will not do it, the next question is what will, and the answer is a level and it is specific to the part.

Lift one part on its own, leaving everything else exactly where it is, and count again. The alto has its whole spectrum back at +6 decibels. The tenor needs +12. The bass needs +18.

Six decibels is about one dynamic marking on the ordinal scale a marking sets up, so an inner voice marked one step above the texture is being given its spectrum, and that is what marking an inner line espressivo against a mezzo forte accompaniment actually buys. Eighteen decibels is three markings, and no bass line is ever marked three steps above the rest.

It is worth comparing that with what the other kind of change costs. A part entering is not a change of level found the loudness ladder’s own control — adding a part — worth about one decibel in the middle of its range. Six decibels of extra level on one existing part is therefore a far larger intervention than any change of texture, and it is the one that decides whether the part is heard as itself. The two ladders are measuring different things and both of them say the same thing about which lever matters.

The bottom voice pays 18 decibels to be heard as itself. Each part of the texture raised on its own while the others stay where they are, against the count of its own 8 partials that stand above what the rest puts up. The bass has its whole spectrum at 18 decibels, the tenor has its whole spectrum at 12 decibels, the alto has its whole spectrum at 6 decibels, the soprano has its whole spectrum at 0 decibels. The curves are not monotone at the top: a part loud enough masks its own upper partials with its own fundamental, because the spreading function's upward slope flattens as the masker gets louder — so the bass is down to 5 of 8 again at 36 decibels, and there is no level at which everything in a texture is fully audible.
Fig. 4 Each part raised on its own with the others standing still. The soprano starts whole; the alto costs six decibels, the tenor twelve, the bass eighteen — and every curve turns back down at the top.

The curves turn back down, and that is the second result of the figure. At thirty-six decibels above the texture the bass is down to five partials of eight again, and the mechanism is the one the third rung established at the scale of a chord: the spreading function’s upward slope flattens as the masker gets louder, so a loud enough note masks its own upper partials with its own fundamental.

So there is no level at which every part of a texture is fully audible. Raise a part far enough and it starts eating itself, and the best the bass ever does is eight of eight over a narrow window between eighteen and thirty decibels above everything else — a window no music is ever in.

Which answers the debt, and not with the number it expected

The question was how much of an inner voice a listener is actually given, in decibels. The answer is five partials of eight at an even balance and six at a conductor’s, which is more than nothing and much less than the score.

But the finding the debt did not predict is that the inner voice is not the problem. The alto is the second-best-served part in the texture. The part that is not delivered is the bass, which nobody thinks of as vulnerable, which is usually doubled and reinforced precisely because it is thought of as the foundation, and which arrives as a sine tone at every dynamic an ensemble ordinarily uses.

There is a further asymmetry in it that is easy to miss. The forward-masking half of this ladder does not apply here at all: a sound hides what came before it is about events in sequence, and the parts of a texture are simultaneous, so the whole of the effect above is the spreading function acting at one instant. That is why the numbers are so much larger than the ones the fourth rung found — that essay was composing a decayed masker with a full-level self-masker and getting nothing, and this one has every masker at full level.

That has a consequence for how a bass line is heard that this collection can state and not test. A sine tone at 131 hertz is a weak carrier of anything: the ear hears the list of a note’s partials and a partial-less note has no list. So whatever makes a bass line perceptually solid, it is not its own spectrum arriving intact, and the candidates left are its fundamental, the room, and the fact that everything above it is built on its harmonics anyway.

One repair that suggests itself does not work. Doubling the bass at the octave — the commonest orchestral reinforcement there is — leaves the C3 part at one partial of eight and gives the new C4 part two of eight. The doubling adds a second weak object rather than restoring the first, because the octave doubling’s own fundamental is exactly where the original’s second partial was, and masks it.

The partials a chord delivers and the partials it hides. Every partial of a 4-note voicing at 70 decibels, against the threshold the rest of the chord puts up at that frequency. A partial below the line is in the arithmetic and not in the ear. 16 of 32 stand above it, and the roughness computed over the ones that survive is 36 per cent of the roughness computed over all of them. That second number is the one that reaches furthest: every roughness figure here counts partials that are present in the score, and a partial the chord masks is not a partial the listener has.
Fig. 5 The same four parts as one sonority, in an earlier drawing: every partial against the threshold the rest of the texture puts up. The partials below the line are the seventeen that do not arrive.

What the spacing does, which is the part orchestration already knows

The third rung compared five voicings at one level and found the widest delivering the least of its computed roughness. Per part, the same comparison says something more usable.

A texture’s masking is decided almost entirely by how far each part’s fundamental is from the parts below it, in Bark rather than in semitones. Close spacing at the bottom puts several fundamentals into one analysis band and each one masks the others’ upper partials from close range; wide spacing at the bottom does not. The rule about spreading a texture at the bottom and closing it at the top is the arrangement that maximises how much of each part arrives, and it is stated in every treatise as a rule about muddiness.

How much of a voicing's roughness the voicing delivers. Five spacings of a chord at 70 decibels, with the share of each one's computed roughness that lies between partials the listener is actually given. The close middle voicing delivers 87 per cent; the widest delivers 40. That is the ordinary orchestration rule arriving from a direction it is never derived from — a spread voicing is not smoother because its partials are further apart, it is smoother because its own bass masks the roughness its top would otherwise have. The best-delivering spacing here is close, middle at 87 per cent.
Fig. 6 The earlier comparison of five spacings at one level. The orchestral layout — a low bass, a gap, a close upper triad — delivers about half its computed roughness, and this essay’s version of the same fact is that it is the layout that gives each part the most of itself.

That is a second mechanism arriving at the same rule, and it is worth separating from the first. Where to put the third derives the spacing from roughness — a low third beats, and the least rough arrangement of three pitch classes turns out to be the harmonic series’ own spacing. This essay derives the same arrangement from audibility, which is a different quantity computed from a different literature. Two independent reasons for one rule is the sort of thing that usually means the rule is about something structural, and here the structure is the same in both cases: the width of a critical band at the bottom of the range.

The band structure is what makes it work, and it is worth seeing directly, because everything above is a consequence of where the boundaries fall.

4 notes, 9 critical bands. The excitation 4 notes of a string spectrum cast along the Bark axis, and the critical-band groups their components fall into. Components inside one group add their intensities and are converted to loudness once; the groups' loudnesses then add. This sonority occupies 9 groups and comes to 52.4 sones, against 127.0 for the same components counted one at a time.
Fig. 7 The four parts’ excitation along the Bark axis, with the critical-band groups their components fall into. Where two parts share a group they mask each other from inside it, and the loudness of that group is counted once.

Which computation produced the numbers

Each part is a string-like spectrum of eight partials falling as one over n, at the level the figure states, with the whole tone’s level fixed so the partials are scaled to the root-sum-square of their amplitudes.

Every partial of every part is tested twice: against the summed masked threshold of every other component in the whole texture, and against the summed threshold of that part’s own other partials alone. The difference is what the texture costs the part. The spreading function is this collection’s own since the first rung — twenty-seven decibels per Bark downward, and 24 + 0.23/fₖ − 0.2L upward — and the threshold of hearing in quiet is the floor.

The balance is a fixed point rather than a search: each part’s offset is moved half way to what its own loudness share says it needs, sixty times, which converges because a part’s loudness rises monotonically with its own level and depends only weakly on the others. The shares come out equal to sixteen decimal places, so the figure is pricing a balance that was actually achieved.

The spacing is the standard four-part orchestral layout and it is a choice. A close-position chord in one octave gives worse numbers and a widely spread one gives better.

Where the model stops

Every part is the same timbre. A real texture is a bassoon under an oboe under a violin, and two different spectra put their partials in different places, so a part whose partials fall between the ones above it rather than on them would keep more. That is the case an orchestra is in and it is what an orchestrator is choosing between.

There is no room. A hall keeps every part alive for a second or two and adds a reverberant field that is much more diffuse in the spectrum than the direct sound. Whether reverberation masks more or less than the direct sound is a real question and this model has not got it.

Nothing here moves. The parts sound together, forever, at fixed levels. A real inner voice moves against a held outer one, and a partial that is masked at one instant is not masked at the next.

Eight partials is what a string spectrum has here. A real instrument has twenty or thirty, and the ones this model does not have are the ones most likely to be masked, so the fractions would look worse and the counts would look better.

And a masked partial is not established to contribute nothing. The threshold model says a partial below the masked threshold is not detectable on its own. Whether it contributes nothing to the timbre of the note it belongs to is a different claim and a stronger one.

What the picture cannot show

It cannot show whether a voice is followed. A listener tracking a line uses continuity, direction and onset as well as spectrum, and the ear builds objects out of far less than a full spectrum. A bass line delivered as a sine tone is still a line.

Nor can it show what a lost partial costs. A note with one partial and a note with eight are different timbres, and how different is a question about listening rather than about counting.

It cannot show the fundamental’s own strength. Every number here is about partials two and above; the fundamental of every part survives in every arrangement tried, so nothing on this page says a part disappears.

And it cannot show what a scoring is for. The loudness constraint the balance solves is not trying to equalise spectra, and reporting that it fails to is not a criticism of it. Roughness can be computed and loudness can be computed; how much of a part’s spectrum arrives has not been an objective in any scoring problem this collection has posed.

Whose music, and when

The texture is a generic four-part layout and the levels are generic. Nothing here is a measurement of a repertoire.

The observation that fits one is about bass doubling. Orchestral practice from the eighteenth century onward doubles the bass line at the unison across several instruments and, in the nineteenth, at the octave below as well — and the reason given is weight. On this arithmetic the octave doubling does not restore the bass’s own spectrum and the unison doubling adds level, which does. That is consistent with the practice being about level rather than about colour, which is what conductors say about it.

The observation that does not fit is four-part vocal writing, where the bass is a singer and a sung bass is exactly the part a listener reports hearing most clearly. Nothing here explains that, and the most likely candidates are that a voice has far more partials than eight and that a hall does something to a low voice that this model has not got.

Where this ladder goes next

Five rungs. One sound hides another, and it hides upward; a sound hides what came before it; a chord hides itself, by an amount that depends on how loud it is and how it is spaced; a line hides nothing unless there is a dynamic contrast in it; and now a texture, which hides its own bass and delivers its top voice whole.

What the ladder owes now is the register, and the criterion is the one this collection uses to find a rung: what has every figure here held fixed? The answer is where the music is. Every count of audible partials on this ladder — the chord, the voicings, the line, and the texture above — is computed on a chord built on middle C or a line that starts there, and the three registers the third rung compared span C3 to C5, which is a third of the range a chord is written in. That is not a neutral choice, because a critical band is thirty-six semitones wide at C2 and under three at C6: the same close triad is a different object at the bottom of the range and the same in name only. Moving it and counting again needs no corpus and no listener, and what would come out is whether this anchor’s headline number is a fact about chords or a fact about the register somebody happened to draw them in.

Part 5 of 8

One essay in the series on masking. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CounterpointCritical bandwidthMaskingOrchestrationPart-writingPartialSpectral balanceVoicing