Perception and the listener

An equal note cannot be masked

Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.

Assumes: A loud chord is a smaller chord · A sound hides what came before it

A loud chord is a smaller chord counted the partials of a chord that survive what the rest of the chord masks, and closed by naming what all three rungs of this anchor share:

Both of the earlier rungs and this one are about one instant, and forward masking has a time constant of a couple of hundred milliseconds — long enough that a note masks its successors as well as its neighbours.

That looks like a straightforward extension. A sound hides what came before it already has the time course, how long a note has to be already has the durations, and a semiquaver at a moderate tempo is a hundred and twenty-five milliseconds — well inside two hundred. Put the two together and a fast line should be partly hidden from its listener.

It is not, and the reason is a proof rather than a number.

Nothing at all until fifteen decibels, and then it depends on the tempo. The fraction of a line's partials that stay above threshold, over how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement — forward masking leaves a threshold at most ten decibels below the masker, and a note's own partials mask each other from the same components at full level. The boundary is between twelve and eighteen decibels, and beyond it the loss grows with the tempo: at 280 to the crotchet and 36 decibels of contrast, 26 per cent of the line's partials are gone. Fifteen decibels is about the gap between a forte and a piano.
Fig. 1 The fraction of a line’s partials that stay above threshold, over how fast it moves and over how much louder everything before each note is. The whole left-hand side is white.

What was expected, and it is worth writing down first

The expectation this rung was set to test is a reasonable one and it appears in the literature in the form a fast passage is partly masked. It has three plausible parts: forward masking lasts two hundred milliseconds; a semiquaver at 120 is a hundred and twenty-five; therefore each note arrives inside the previous note’s shadow.

All three are true. The conclusion does not follow, and the step that fails is between the second and the third — it assumes that being inside a shadow means being darkened, and the depth of the shadow was never compared with anything.

The proof, and it is one line

Forward masking leaves a threshold that is at most ten decibels below the masker’s own level, at the instant the masker stops, decaying from there.

A note contains the same components as its predecessor — the same partials of a similar fundamental — at full level, and those components are already masking each other. That is the third rung’s whole calculation.

So the predecessor’s contribution to the threshold at any partial is at most ten decibels below the contribution the note’s own partials are already making at that partial. Adding a term ten decibels down to a term that is already there changes a total by 0.4 of a decibel, and never takes a partial from above threshold to below it.

An equally loud predecessor cannot mask anything, at any tempo, ever. Not because two hundred milliseconds is short — it is not — but because a note masks itself harder than anything before it can.

That is a stronger statement than a computation, and it is why the whole left-hand column of the figure above is white at every tempo from forty to two hundred and eighty crotchets a minute.

What does hide a line

The variable the proof leaves open is the one it fixed: the predecessor’s level.

Make everything before a note louder than the note and the residual threshold rises above what the note’s own partials contribute. The boundary is between twelve and eighteen decibels, and past it partials start to go.

louder by 40 to the crotchet 130 280
0–12 dB nothing lost nothing nothing
18 dB nothing 7% 10%
24 dB nothing 7% 15%
36 dB 4% 13% 26%

Fifteen decibels is roughly the gap between a forte and a piano. So the phenomenon this rung is about is not fast music at all — it is dynamic contrast, and specifically the case a composer writes as an echo: a loud phrase answered quietly by the same instrument.

The echo really is spectrally impoverished, and not merely quieter. A piano answer to a forte phrase arrives into a threshold the forte left behind, and loses about a tenth of its partials at an ordinary tempo and a quarter at a fast one.

It is worth being precise about what fifteen decibels means, because a decibel is not a musical unit. On the ordinal scale a dynamic marking sets up, fifteen decibels is about three markings — ff to mp, or f to pp — which is a large gesture rather than an ordinary shading. Anything within a marking or two of its neighbour is on the white part of the map.

And the tempo enters through the fraction, not the depth

The tempo does nothing to how hard a note is masked. It changes how much of the note is inside the shadow.

A note is masked most at its own beginning. How many of a note's ten partials are above threshold at each instant of its own life, with the note before it 24 decibels louder and decaying behind it. A predecessor ends exactly when a legato successor begins, so the threshold it leaves is at its highest at the new note's onset and falls away over about two hundred milliseconds. At 50 to the crotchet the note lasts 300 milliseconds and spends most of it clear; at 200 it lasts 75 and never gets there. The tempo does not change how hard a note is masked — it changes what fraction of the note is inside the shadow.
Fig. 2 How many of a note’s ten partials are above threshold at each instant of its own life, at two tempi, with a predecessor twenty-four decibels louder decaying behind it.

A legato predecessor ends exactly when its successor begins, so the threshold is at its highest at the new note’s onset and falls from there over two hundred milliseconds. Every note is masked most at its own beginning.

At fifty to the crotchet a semiquaver lasts three hundred milliseconds, so the note spends its first fifty in shadow and the remaining two hundred and fifty clear. At two hundred it lasts seventy-five and never gets out.

That is a real difference and it is a difference in fraction, which is why the loss grows with tempo in the table above without the mechanism changing at all.

It also predicts something about articulation. A detached note leaves a gap after its predecessor, which is exactly the interval over which the threshold falls — so staccato buys back the partials that legato spends. Half a semiquaver of silence at two hundred to the crotchet is thirty-seven milliseconds, and over that interval the residual threshold falls from sixty decibels to twenty-seven — a gift of thirty-three decibels for a gap nobody hears as a gap.

Playing a quiet answer detached restores most of what the loud phrase before it took, and that is a thing performers do without a reason. It is also the reverse of the usual advice, which is that a quiet passage should be played more legato to carry — and the two are not in conflict, because carrying is about level and this is about spectrum.

The same arithmetic says a rest is worth a great deal. A quaver rest at two hundred to the crotchet is a hundred and fifty milliseconds, which takes the residual threshold from sixty decibels to under five: an entrance after a rest arrives into essentially nothing, whatever came before it.

Two envelopes. How loudness changes over the life of a note, for legato and detached. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.
Fig. 3 The difference the paragraph above is about: where a note’s energy stops. A legato release hands the next note a fresh threshold and a detached one hands it a decayed one.

The interval matters and not the way it looks

A second variable the third rung did not have is how far the line moves.

Masking spreads upward far more readily than downward — the pattern falls at 27 decibels per Bark below the masker and at six to sixteen above it, depending on level — so a note masks its higher neighbours much more than its lower ones.

For a line that means the direction of a leap should matter. A rising line puts each note into the upward spread of the one before; a falling line puts it into the downward spread, which is three to four times steeper.

Computed on the same line ascending and descending, at the same tempo and the same twenty-four decibels of contrast, the difference is 0.7 per cent — 57.9 partials against 57.5 out of sixty-four — and it is in the wrong direction: the descending line loses very slightly more.

The asymmetry is enormous in the spreading function and absent in the line. The reason is the same one the proof rests on: almost all of the masking a partial receives comes from the note’s own lower partials, at full level, and the previous note is a correction on a correction. A direction effect that is a factor of four in the underlying function comes out as noise once it has been put behind that.

That is worth recording as a refusal rather than a result. The upward-spread asymmetry is one of the most quoted facts in this part of the subject and it is quoted as though it had melodic consequences. On this model it has none.

That is the same reason the proof works, restated: the note’s own spectrum is the dominant masker of the note’s own spectrum, and everything else is a correction.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward.
Fig. 4 The asymmetry the paragraph above is about, established earlier: a masker reaches much further up than down, and further up the louder it is.

The one place the numbers say to look

Putting the two variables together says where in music the effect should be findable, and it is a narrow place.

It needs a level contrast of more than fifteen decibels between successive events, a tempo fast enough that a note lives inside the shadow, legato articulation so there is no gap, and one instrument so the components coincide. That is a subito piano inside a fast legato line on a wind instrument or a voice — and it is a specific enough recipe to be worth stating, because everything looser than it comes out at zero.

Most music fails at least one of the four conditions. A piano fails the first by decaying rather than stopping; an orchestra fails the last by having several instruments whose partials do not coincide; a slow movement fails the second. The anchor’s third rung applies everywhere and its fourth applies almost nowhere, and knowing which is which is what the rung is for.

What this does to the anchor’s own claim

The third rung’s headline is that the count of audible partials — which every roughness figure on this site takes as given — is itself a function of dynamic and voicing. That survives here and gains a boundary.

Within a chord the effect is large, because every component is present at once at full level and the spreading function decides everything.

Between notes the effect is zero at one dynamic and appears only past fifteen decibels of contrast. So a passage played evenly delivers every note’s full spectrum however fast it goes, and a passage with a large dynamic contrast in it does not.

That is a useful narrowing. It says the third rung’s warning applies to texture — how many notes are sounding and how they are spaced — and not to speed, which is the thing the debt expected it to apply to.

How much of a voicing's roughness the voicing delivers. Five spacings of a chord at 70 decibels, with the share of each one's computed roughness that lies between partials the listener is actually given. The close middle voicing delivers 87 per cent; the widest delivers 40. That is the ordinary orchestration rule arriving from a direction it is never derived from — a spread voicing is not smoother because its partials are further apart, it is smoother because its own bass masks the roughness its top would otherwise have. The best-delivering spacing here is close, middle at 87 per cent.
Fig. 5 The earlier figure: how the count of audible partials moves with voicing, at one level and one instant. That effect is much larger than anything on this page.

Both of those are simultaneous masking, which is where this anchor’s effects are large. The two figures below are the temporal machinery that has to be composed with them, and the composition is what returns nothing.

The threshold before, during and after a burst. The level a brief probe needs in order to be heard, plotted against when it happens relative to a 70 dB burst that occupies the shaded band. To the right is forward masking, which decays over about 200 milliseconds. To the left is backward masking: the threshold is raised for a probe that has already finished before the masker begins.
Fig. 6 The decay everything on this page rests on, computed earlier. The cap at the left-hand edge — ten decibels below the masker at the instant it stops — is the number the proof is made of.
The partials a chord delivers and the partials it hides. Every partial of a 3-note voicing at 90 decibels, against the threshold the rest of the chord puts up at that frequency. A partial below the line is in the arithmetic and not in the ear. 12 of 24 stand above it, and the roughness computed over the ones that survive is 67 per cent of the roughness computed over all of them. That second number is the one that reaches furthest: every roughness figure here counts partials that are present in the score, and a partial the chord masks is not a partial the listener has.
Fig. 7 And the effect found earlier, at one instant: how much of a chord’s own spectrum the chord removes. Everything successive is a correction on this.

Which computation produced the numbers

The line is eight notes of a scale on middle C, at ten partials each with a one-over-n-squared envelope, at seventy decibels unless the figure says otherwise.

Each note’s own partials mask each other by the spreading function this collection has used since the first rung: 27 decibels per Bark downward and 24 + 0.23/fₖ − 0.2L upward, which is between six and sixteen over the levels here.

The predecessors are folded in with two functions composed rather than one. The spreading function gives what a partial of an earlier note would contribute at the probe’s frequency if it were still sounding; temporalMasking then reduces that by the published forward-masking decay, which is logarithmic in time and reaches zero at two hundred milliseconds.

The argument of that decay is milliseconds after the masker ends, which in a legato line is milliseconds into the new note. Getting that wrong — measuring from the masker’s onset instead — makes every number on this page smaller by a factor that grows with the tempo.

Each note is sampled at twelve instants through its own duration and the audible count is averaged, because the count is not constant through a note and the interesting quantity is what fraction of it is masked.

Where the model stops

Ten decibels is a cap and not a measurement. The whole proof rests on temporalMasking returning at most Lm − 10 at the instant a masker stops, which is this collection’s reading of the published time courses. A cap of Lm − 3 would let an equal predecessor take partials away and would change the answer entirely, so the number is load-bearing and is worth naming.

Notes do not stop when the next one starts. A piano’s note decays for seconds and a bowed note stops when the bow does, so “the predecessor ends when the successor begins” is true of a wind or a voice and false of most of a keyboard.

And there is no reverberation. A hall keeps every note alive for a second or two, which is a masker this model has not got and which would apply at full level rather than at a decayed one.

The line is one instrument. Two notes of different timbre have partials at different frequencies, so the predecessor’s components no longer coincide with the successor’s and the proof’s premise weakens — which is the case an orchestra is in and the one this rung has not computed.

There is no room. A hall keeps every note alive for a second or two at full level, and what a room does to a chord’s own spectrum is a masker this model has not got.

Nor is there an accompaniment. The case where forward masking obviously matters in music is a quiet line under a loud one, and that is a simultaneous masking problem with a temporal component, not the successive one computed here.

What the picture cannot show

It cannot show what a lost partial costs. A partial below threshold is not a note that disappears; it is a note whose timbre is thinner, and how much thinner is a question about what the spectrum is for rather than about how many of it there are.

Nor can it show a listener attending. Masking thresholds are measured on listeners told to detect a tone, and a listener following a melodic line is doing something else with the same signal — the ear builds objects out of the same seventy decibels, and an object that has been decided on is not one a threshold decides.

It cannot show the release. The staccato claim above uses the decay curve at a gap length and assumes the note is silent in the gap, which is true of a wind instrument and generous to a string.

It cannot show a repeated note. The worst case for the proof is a note repeated at the same pitch, where the predecessor’s components coincide exactly with the successor’s — and even there the ten-decibel cap holds, so a repeated note is the case where the calculation is tightest and still returns nothing.

And it cannot show two hundred years of dynamics. Fifteen decibels is a forte against a piano on the model this collection uses; the actual level difference a player produces between two markings is an ordinal scale laid over a continuous one and varies by instrument and by period.

Whose music, and when

The line is a generic scale and the levels are generic. Nothing here is a measurement of a repertoire.

The observation that fits one is the echo. Answering a loud phrase quietly with the same material is a device with a long history — it is a baroque commonplace, it is written into terraced-dynamic keyboard music by the instrument itself, and it survives into the nineteenth century as a marked effect. What the model adds is that the echo is not a copy at a lower level: it arrives into a threshold the first phrase left behind and is measurably thinner, by about a tenth of its partials at an ordinary tempo.

Whether that is why the device works is not a question this figure can answer, and it is a nicer mechanism than “it is quieter”.

Where this ladder goes next

Four rungs. One sound hides another, and it hides upward; a sound hides what came before it; a chord hides itself, by an amount that depends on how loud it is and how it is spaced; and now a line, which hides nothing at all unless there is a dynamic contrast in it.

What the ladder owes now is the accompaniment. Every rung above puts the masker and the probe in the same voice — the same chord, or the same melodic line — and the case that made this anchor worth opening is the other one: a quiet inner part under a loud outer one, sounding at the same time, where the spreading function and the forward decay both apply and the levels are set by the scoring rather than by a marking. This collection has an orchestration ladder that chooses those levels and a masking ladder that prices them, and the two have never been introduced. What would come out is how much of an inner voice a listener is actually given, which is a question about counterpoint that nobody has asked in decibels.

Part 4 of 8

One essay in the series on masking. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AudibilityCritical bandwidthDynamicsForward maskingMaskingPartialTempo