Form and structure

A soft chord has to fade in

Forward masking sits between the two integration times already in play — two hundred milliseconds against a thirty-seven millisecond roughness window and a two-second loudness release — and it was owed as the term that might eat the dissonance contrast. It does not. It widens it, by three per cent at a chorale's pace and fifty-nine at four chords a second, because it can only ever remove partials and it can only reach the chord a dynamic has already made quiet. What it does instead is stranger: one chord in the passage is entirely inaudible for its first twelve milliseconds and takes a quarter of a second to arrive whole.

Assumes: The dissonance arrives and the dynamic does not · A sound hides what came before it

The dissonance arrives and the dynamic does not put two integration times under one passage and found they differ by a factor of fifty. A roughness has to be integrated over several cycles of its own fluctuation, which is thirty-seven milliseconds; a running loudness decays with a measured release of two seconds. So at the pace music is played, a listener receives every chord’s dissonance whole and receives its dynamics averaged.

Its last line named the term that could undo that. There is a third temporal constant in this collection and it is neither of those two: forward masking, whose couple of hundred milliseconds sits squarely between them. A loud chord does not merely raise the running loudness. It raises the threshold, and part of the next chord is then not heard at all.

What the chord before takes out of the chord after. Five chords at 1.2 seconds each, with the roughness each one has on its own — its simultaneous masking and the threshold of hearing already applied — and the roughness it actually has once the chord in front of it has raised the threshold. Four of the five are untouched. The fifth, i, two parts, follows the only step in this passage that falls more than fifteen decibels, and it arrives into a hole: it is entirely below threshold for its first 13 milliseconds and takes 240 to get all of itself back. Masking can only remove partials, so it can only lower a roughness — and the chords it can reach are the ones a written dynamic has just made quiet, which are already the smooth ones. Across the passage the dissonance contrast goes from 3021 to 3113: the mask widens it by 3.0 per cent rather than eating it.
Fig. 1 Five chords, with the roughness each has on its own and the roughness it has once the chord in front of it has raised the threshold. Four of the five lie on top of their own baselines. The fifth arrives into a hole.

Which chord it can reach, and it is one of five

The passage is the one this anchor has been using: five chords with a written dynamic for each, at levels solved so that each chord’s own loudness matches its marking. Written out as decibels of the level knob, the steps between them are +0.9, +12.4, +2.3 and −16.4.

Only the last of those is a fall, and only the last is large. That matters because the size at which forward masking begins to take anything is not a free parameter here — another ladder measured it, and the number is fifteen.

The threshold before, during and after a burst. The level a brief probe needs in order to be heard, plotted against when it happens relative to a 50 dB burst that occupies the shaded band. To the right is forward masking, which decays over about 200 milliseconds. To the left is backward masking: the threshold is raised for a probe that has already finished before the masker begins.
Fig. 2 The time course this page borrows, drawn at the level of the passage’s loudest chord. The threshold a burst leaves behind starts about ten decibels below the burst’s own level and decays to nothing over roughly two hundred milliseconds, linearly in the logarithm of the delay.

The fifteen decibels, which is somebody else’s result

An equal note cannot be masked is a proof rather than a measurement, and it is the constraint everything on this page has to survive.

Forward masking leaves a threshold at most ten decibels below the masker’s own level, at the instant the masker stops, decaying from there. A chord also masks itself: every partial raises the threshold at every other partial, and a loud chord is a smaller chord counts what that alone removes. So a successor at the same level as its predecessor is already masking itself harder than the predecessor can, and nothing is taken. The contrast needed before a predecessor takes anything at all is about fifteen decibels.

The passage has exactly one step that large. So of the five chords, four are untouched and one is not, and which one it is was decided by the written dynamic rather than by anything acoustic.

It is worth being explicit that this is a piece of luck about the passage rather than a general result. A passage of five chords all at one dynamic has no masked chord in it at all; a passage of alternating forte and piano has one after every other chord. What the fifteen decibels buys is that the question which chords does the mask touch has a purely notational answer — read the dynamic markings, look for the falls larger than about one and a half steps of the eight-step scale a dynamic mark is measured on, and those are the chords.

The chord that is not there when it starts

The chord after the fall is a two-part I, marked at the quietest dynamic in the passage, and on its own it puts eleven partials above threshold. Arriving where it does, it puts none.

I, two parts arriving into the hole the chord before it left. How many of i, two parts's 11 audible partials are above threshold, against how long the chord has been sounding, on a logarithmic axis because that is the shape forward masking decays with. The chord is entirely inaudible for its first 13 milliseconds, has 3 partials back at 16 and all 11 at 240. Both windows are drawn on it, and their order is the point: the roughness window is 37 milliseconds and the mask is two hundred, so a listener resolves the climb rather than averaging over it. The chord is received as something that grows into itself — a fade-in nobody wrote, produced by the chord in front of it.
Fig. 3 The last chord coming back out of the hole, on a logarithmic time axis because that is the shape the decay has. Nothing is above threshold for the first twelve and a half milliseconds; three partials are back at twenty-eight; all eleven at two hundred and forty. Both windows are marked on it.

Twelve and a half milliseconds of complete inaudibility, then a climb: three partials at twenty-eight milliseconds, four at sixty, five at a hundred and forty, and all eleven at two hundred and forty. A chord written as an event with an onset is received as something that grows into itself over a quarter of a second.

And nobody wrote any of it. The dynamic marking is a level; the chord is a block; the score has an attack and a duration. The fade-in is produced entirely by the chord in front of it, and its shape is a property of the step rather than of either chord.

Which order the partials come back in is the other half of it and it is not random. The ones that return first are the ones furthest in frequency from where the previous chord had its energy, because masking spreads upward much further than downward and a partial sitting under the masker’s own region is the last to clear. So the chord does not merely grow louder as it arrives; it grows from a different spectrum, thin and oddly weighted at first, filling in toward its own shape over the quarter second. A soft chord after a loud one is a timbral event before it is a harmonic one.

Why a listener gets to see it rather than average it

Which brings the two windows back, and this is where the ordering does real work.

If the roughness window were long compared with the mask, the climb would be integrated away: a listener would receive one number for the chord, somewhere between what it has at the start and what it has at the end, and there would be nothing to notice. The window is thirty-seven milliseconds and the mask is two hundred, so the ratio runs the other way. A listener resolves the mask lifting.

One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics.
Fig. 4 The two windows that decide it, drawn just before: how much of each quantity’s contrast a listener still has by the end of a chord, against how long a chord lasts. The roughness window is short enough to sit inside the mask, which is what makes the climb an event rather than an average.

That is a slightly unusual thing for a limit of hearing to do. The three constants here are a filter’s ringing, a nerve’s recovery and a memory’s decay, and they have no reason to be ordered conveniently — but the shortest of them is five times shorter than the middle one, so the middle one’s shape is available rather than smeared. A listener has the resolution to hear a mask lift.

Whether they use it is a different question and this page cannot answer it. What can be said is that the information is delivered, which is the same standard the rung below held itself to.

The ordering also settles something the two windows on their own could not. A quantity with a thirty-seven millisecond window and a quantity with a two-second release differ by a factor of fifty, and it is tempting to read that as one is fast and one is slow. The mask shows they are not merely fast and slow, they are on opposite sides of something: a two-hundred-millisecond event is an event to the first and is invisible to the second. The loudness smoother sees a chord that fades in over a quarter of a second as a chord starting slightly late and slightly soft, which is a change of a fraction of a sone it will not have finished registering before the next chord arrives. Same physical event, two instruments, and only one of them has a reading.

What it does to the contrast, and it is the opposite of what was owed

The term was owed as a threat. The rung below found that a passage’s dissonance contrast survives to the listener whole — a factor of three thousand across five chords — and the mask was named as the thing that might eat it.

It does not eat it. It makes it larger.

The mask widens the dissonance contrast, and faster music more. The ratio between the passage's dissonance contrast with forward masking in it and without, against how long each chord lasts. Above one, the mask makes the difference between the roughest chord and the smoothest LARGER than it was played. At 3 seconds a chord it is 1.014 and at 0.1 it is 1.588 — a fixed two hundred milliseconds is a sixth of a slow chord and the whole of a fast one, so the effect grows as the music moves. The direction is the useful part: masking removes partials and cannot add them, and a chord is only masked if its predecessor was fifteen or more decibels louder, which has already divided its roughness by thirty-two because roughness is quadratic in pressure. So the mask reaches the chords the dynamic has already made smooth, and pushes them further down.
Fig. 5 The passage’s dissonance contrast with the mask in it, over the same contrast without, against how long a chord lasts. Above one at every pace, and further above it the faster the music moves: three per cent at a chorale and fifty-nine per cent at ten chords a second.

Three per cent at a second and a bit per chord. Twenty-three per cent at a fifth of a second. Fifty-nine per cent at a tenth, which is about ten chords a second — faster than harmonic rhythm ever goes, and included because it is where the curve is going.

The trend has a mechanical reason. The mask is a fixed two hundred milliseconds and a chord is not, so the same hole is a sixth of a slow chord and the whole of a fast one. Nothing about that is a fact about hearing; it is a fact about a fixed interval divided by a variable one.

Why the sign is not an accident

The direction is the part worth keeping, and it comes out of two things this collection already knew.

Masking removes partials and never adds one. Roughness is a sum of non-negative terms over pairs of partials, so a chord with fewer partials is never rougher than the same chord with more. Whatever the mask does to any individual chord, it does downward.

And it can only reach a chord whose predecessor was fifteen or more decibels louder. Roughness is quadratic in pressure — that is the second rung’s finding — so a chord fifteen decibels down has already had its roughness divided by thirty-two before the mask arrives. The mask’s entry fee is paid in the same currency as its effect.

Put the two together and the mask reaches the chords a written dynamic has already made smooth, and pushes them further down. The maximum of the passage is untouchable and the minimum is pressed. That is a widening, and it is a widening by construction rather than by luck.

The passage where it goes the other way

Which is a claim strong enough to test, and it is not a theorem, so the honest thing is to find where it fails.

For the mask to narrow a contrast it has to reach the passage’s roughest chord, which means that chord must be the quieter of a pair fifteen decibels apart and still the rougher — so its spacing must make it more than thirty-two times rougher than its predecessor’s at the same level. Over four-note chords inside four octaves the whole range of spacing roughness at one level is a factor of about a hundred and eight, from a semitone cluster at the bottom of the range to an open chord of octaves. So there is room, and it is narrow: taking every ordered pair of those chords, 0.019 per cent of them clear the factor of thirty-two.

The passage where the mask narrows the contrast instead. Two passages, each with and without forward masking, over four paces. The standing passage is widened at every pace; the second is narrowed at every pace, by up to 2.5 per cent. It is built to be the exception and the construction says what the exception needs: a chord whose SPACING makes it more than thirty-two times rougher than its predecessor's while being the quieter of the two, because thirty-two is what fifteen decibels of level does to a quantity quadratic in pressure. Over four-note chords inside four octaves the whole range of spacing roughness at one level is a factor of about a hundred, so the room for this is real and narrow — and the largest narrowing available is 2.5 per cent against a widening that reaches 31.
Fig. 6 The standing passage beside one built to be the exception: a loud open chord of octaves, then a soft low cluster, then the open chord again. The second is narrowed at every pace and the first is widened at every pace, which is the whole of the claim that the widening is not a theorem.

Built deliberately, the exception works and it is small. The constructed passage narrows by up to two and a half per cent, against a widening that runs to fifty-nine. So the mask’s power to widen a contrast is roughly twenty times its power to narrow one, and the arrangement that reverses it — a soft dense cluster immediately after a loud open chord — is a real orchestral gesture rather than an invented one.

At the pace music is actually played

One more reading, because the numbers above are at a chorale’s pace and most music is not a chorale.

Two quantities, two windows, one passage. Five chords, each held 0.3 seconds, with the roughness and the loudness the scoring produces — both raw and both as a listener integrating the past would have them. The loudness smoother's long release is two seconds and rounds every corner off. The roughness window is 37 milliseconds — several cycles of the fluctuation the roughness itself has, at a median rate of 108 hertz — and follows the chords exactly. Both axes are logarithmic over four decades, and the reason is the size of the two: the roughness varies by a factor of 4146 across this passage and the loudness by 2.8.
Fig. 7 The earlier figure at four chords a second rather than one. The roughness track still follows the chords exactly; the loudness track has lost most of its shape; and the mask’s contribution, invisible at a slow pace, is now fifteen per cent of the dissonance contrast.

At three hundred milliseconds a chord — four chords a second, a moderate harmonic rhythm in almost any repertoire — the dynamic contrast a listener receives is about two-thirds of what was written, the dissonance contrast is all of it, and the mask has added fifteen per cent on top.

So the picture the rung below drew gets more lopsided rather than less. The quantity a composer marks explicitly is flattened by the ear’s own release; the quantity nobody marks at all arrives whole, and then arrives slightly exaggerated by a mechanism whose usual description is that it takes things away.

Which computation produced the numbers

The passage, the levels and the two windows are the rung below’s, unchanged.

The masking machinery is the masking ladder’s, also unchanged. Each partial of the preceding chord is spread across frequency by the site’s own spreading function — twenty-seven decibels per Bark downward and a level-dependent slope upward — and decayed by the published forward-masking time course, which is linear in the logarithm of the delay and reaches nothing at about two hundred milliseconds. A partial of the current chord is dropped when its own level falls below the sum, in power, of what the predecessor leaves, what the chord masks in itself, and the threshold of hearing.

The survivors are then handed to the same roughness sum the anchor has used since its first rung: every partial of one part against every partial of another, through the Plomp–Levelt curve scaled to the critical bandwidth, quadratic in pressure.

The baseline each chord is compared against is not its raw partial list. It is the same chord with its own simultaneous masking and the threshold of hearing already applied, so that what is measured is the temporal effect and not a second helping of the simultaneous one.

The samples inside each chord are spaced logarithmically, because that is the shape the decay has and because a uniform grid at any affordable density steps straight over the first fifteen milliseconds, where the whole of the fade-in happens. The mean over a chord is trapezoidal in real time, so the dense end is not overweighted.

Where the model stops

Two hundred milliseconds is asserted and the ten decibels is asserted. Both come from the published time course rather than from anything here, and both carry the caveats the masking ladder records: the recovery is shorter at lower levels in the literature and is not in this model, whose level dependence is entirely in the depth.

Every part has the same timbre. The chords are scored with one partial list for every note, so the differences between them are differences of voicing and register — which means the one variable this anchor exists to study, who plays what, is absent from the passage that is being masked.

And the masking here is between chords rather than inside one. A chord masks itself as well, and what that leaves of each part is a separate and larger effect: the listener is given the top voice counts a four-part texture down to one partial on the bass. That is already applied to each chord as the baseline here, so what this page adds is only the part the predecessor takes — a smaller quantity, in one chord of five.

The chords are blocks. Real voice-leading holds some parts and moves others, so a real transition presents the mask with a partly unchanged spectrum, which is the case where forward masking is strongest and where this model has nothing to say.

And nothing decays. Each chord is a steady state for its whole length and then stops. A real chord’s own decay is the largest thing happening in its last two hundred milliseconds and would change what the next chord arrives into.

What the picture cannot show

It cannot show the room. Two seconds of reverberation puts each chord’s energy into the next one at a level the mask would have to be computed against, and that is a much bigger masker than the direct sound this page uses.

Nor can it show backward masking. A chord raises the threshold for a few milliseconds before it begins as well as after it ends, which is the odd half of the temporal masking result and which would shave the end of every chord rather than the start of the next.

It cannot show a listener attending. A masking threshold is measured on somebody told to detect a tone, and a listener following a harmony is doing something else with the same signal.

And it cannot show whether the fade-in is heard as a fade. Twelve milliseconds of silence and a quarter-second climb are inside the window where the ear stops timestamping events separately, so the honest description may be that the chord is heard as having a soft attack rather than as arriving late. Those are different claims and no figure here separates them.

Whose music, and when

The passage is generic and the dynamics are generic, so nothing here is about a repertoire.

The observation that fits one is about the subito piano. The gesture the rung below found unproducible in loudness — a full chord followed at once by a quiet one — is exactly the gesture that produces the only masked chord in this passage, and what the mask does to it is give it a soft, slow entry that no player is producing and no marking asks for. A device that is famously difficult to bring off is being helped by a mechanism nobody has ever credited, and hindered by a smoother nobody can defeat, at the same moment.

That is consistent and it is not evidence. It is stated because the temptation to make it an explanation is worth seeing.

Where this ladder goes next

Six rungs. Who plays what and how loud is one question; the ranking survives the dynamic and the chord does not; a dynamic mark changes what a note is; a subito piano is a rate rather than a level; the other quantity that succession has arrives whole; and now the mask, which arrives with it and pushes the same way.

What is owed after this is the fourth player. Every figure on all six rungs puts exactly one instrument on each note of the chord, which makes an arrangement a permutation and is the assumption the whole assignment machinery is built on — and it is the one thing orchestration does not do. The commonest single operation in the craft is a doubling, and a doubled note has two players on it, so the arrangements stop being permutations and become surjections. What that changes is not obvious in either direction: it might add a whole dimension to the search, or it might add nothing at all, because the spectrum ladder has already shown that two players on one note sound like one of them. Both machineries exist and neither has been pointed at the other. It needs arithmetic and nothing else.

Part 6 of 14

One essay in the series on orchestration. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DynamicsForward maskingIntegration windowLoudnessMaskingOrchestrationPartialRoughness