One sound hides another, and it hides upward
Assumes: The quietest thing audible, and why the volume knob is a tone control
A quiet tone at 1,600 Hz is perfectly audible in a silent room. Add a loud tone at 1,000 Hz and it is not merely harder to hear — it is gone, as completely as if it had been switched off, and switching it off makes no audible difference at all. Reverse the two, putting the loud tone above and the quiet one below, and the quiet one survives.
The effect is called masking, its asymmetry is the interesting half, and both are computable.
The left-hand slope is the same in all three curves. The right-hand slope is not: it gets shallower as the masker gets louder, so a loud tone reaches much further up the spectrum than a quiet one does and no further down it at all.
Why upward
The mechanism is in the shape of the travelling wave on the basilar membrane, and it has a direction built into it.
Sound entering the cochlea travels from the base — which responds to high frequencies — towards the apex, which responds to low ones. A wave at a given frequency builds gradually as it travels, peaks at the place tuned to it, and then dies away abruptly. The rise is slow and the fall is sudden.
Turn that shape around and read it as which places are disturbed by a tone of frequency f, and the answer is: everywhere between the base and f’s own place, which is to say everywhere above f in frequency. The regions tuned to higher frequencies are on the wave’s route. The regions tuned to lower ones are past its cliff.
Louder makes it worse because the response is compressive. The cochlear amplifier provides most of its gain at low levels and progressively less at high ones, which flattens the tuning of each place as the input grows — a filter that is sharp for a quiet sound and broad for a loud one. The standard model puts the upward slope, in decibels per Bark, at
with the masker frequency in kilohertz and its level in decibels. At 40 dB that is about 16 dB per Bark; at 90 dB it is about 6. The downward slope is taken as a constant 27 dB per Bark and does not care about the level at all.
The asymmetry runs opposite to intuition drawn from a graphic equaliser, where turning up a band affects that band. Here, turning a band up affects every band above it.
What this does to an arrangement
The prediction is specific and it is testable against three hundred years of orchestration practice, which is a useful thing for a psychoacoustic model to be testable against.
A low, loud instrument buries a high, quiet one, and not the reverse. This is why a piccolo can be heard over a tutti and a bassoon in the same relative position cannot. The reason usually given is the asymmetry itself, and the section on partial masking below shows that reason is wrong and the conclusion right.
Doubling at the octave costs less than doubling at the unison, and the excitation model prices it. Splitting one note’s power between two players at the unison gives a sonority exactly as loud as the one note was — nothing is bought at all. The same power split at a fifth gives 1.13 times, at an octave 1.21, at two octaves 1.41. So the practice of doubling melodies at the octave rather than in unison, near-universal in nineteenth-century orchestration, is worth about twenty per cent of loudness for nothing, and the unison doubling that a young arranger reaches for first is the one arrangement of two players that buys none.
The width of the ear’s analysis band is what decides all of this, and putting it in Bark is what makes the shape simple: on that axis a masking pattern is two straight lines, and the reason a low note hides so much more than a high one is that a fixed number of Bark is a much wider stretch of hertz down there.
Register decides who wins, not dynamics. A cello line in the bass at mezzo-forte can be perfectly audible under a fortissimo brass chord in the middle register, because the masking is running the wrong way to reach it. Composers who never heard of a Bark scale worked this out empirically and it is embedded in the textbooks as a rule about spacing: leave the bass open, crowd the middle if necessary, and never crowd the top.
Whose practice. All three claims above are claims about the Western orchestral repertoire from roughly Berlioz onward, written for instruments in the standard ranges, in halls of a particular size. A gamelan is organised on entirely different principles and a rock mix, where the balance is set with faders rather than with numbers of players, has different failure modes — the classic one being a bass guitar and a kick drum in the same octave, which is a masking collision two instruments deep and the reason the standard fix is to move one of them.
A note masks itself
The figures so far have used a pure tone as the masker, and there are almost none of those in music. A real note is a whole harmonic series, and every partial in it is a masker.
This is why a note is heard as one object rather than as a stack of tones — a result the streaming essays return to from the other direction, where what makes two partials one note turns out to depend on behaviour over time and not only on the spectrum. Each partial sits in the masking shadow of the ones below it, and only a partial that escapes the shadow can be heard out of the note as a separate pitch. A trained listener can just manage it for the lowest few partials of a sustained tone, and nobody manages it for the twelfth.
It is also the mechanism behind a fact every arranger knows and few can justify: a chord voiced with its notes close together in the bass sounds muddy, and the same chord voiced widely does not. In the bass the critical bands are wide in musical terms — a minor third at 110 Hz sits inside one band — so close bass notes both mask each other and beat, and the masking sheet each one throws upward covers everything the upper voices are trying to do.
The floor beneath the floor
There is a masked threshold beside every sound, and there is a threshold when there is no sound at all. They are the same quantity and the second is the limiting case of the first.
The two thresholds combining is also why masking is much more consequential at the edges of the range. At 60 Hz the threshold in quiet is already near 40 dB, so it takes very little masking to push a bass partial under it. At 3 kHz the threshold is near zero and a sound has to be pushed thirty or forty decibels down before it disappears. The ear’s sensitive band is also its hardest band to hide anything in, and both facts come out of the same curve.
The number that made it an industry
Masking’s largest consequence is not musical. It is that most of a digital audio signal is inaudible, and can be worked out in advance which parts.
A perceptual codec computes exactly this figure for every short frame of a signal: it finds the peaks, applies the spreading function, adds up the resulting curve into a global masked threshold, and then spends its bits only on what pokes above the curve. Everything below can be replaced by quantisation noise, provided the quantisation noise is also kept below the curve. The compression is not throwing away subtlety; it is throwing away signal that the ear demonstrably does not receive.
The ratio is roughly ten to one, which is what MP3 at 128 kbit/s achieves against a CD, and the number is a psychoacoustic result rather than an information-theoretic one. It says the ear discards about ninety per cent of what arrives at it before anything else happens.
Release from masking, which is the other half
Everything so far describes how a sound is hidden. The complementary question — what un-hides it — turns out to have answers that have nothing to do with making it louder, and they are the reason listening to an orchestra is not the ordeal the figures above imply.
Being in a different place. A masker and a probe presented to both ears identically mask maximally. Move the probe’s apparent position by giving it a different interaural delay and it becomes audible again, at a level ten to fifteen decibels below the masked threshold. The improvement is called binaural masking level difference, it is largest below about 500 Hz, and it is the same machinery two ears use to work out a direction being spent on detection instead. A soloist standing to one side of an orchestra is easier to hear for a reason that has nothing to do with the soloist.
Starting at a different time. A probe that begins before the masker, or after it, is much easier to detect than one that begins with it, even when the overlap is nearly total. This is the same fact the streaming essays use, and it is why an instrument with a distinct attack survives a texture that would swallow a sustained one.
Being modulated differently. A masker whose amplitude wobbles releases a probe that does not wobble with it — comodulation masking release. Vibrato is a frequency wobble rather than an amplitude one, but it has the same consequence: a vibrato tone in a steady texture is detected at a lower level than a steady tone would be. Whether that is why singers use vibrato is a question this essay is not equipped to answer, and it is a claim frequently made without evidence.
All three say the same thing. Masking is not a property of a spectrum; it is a property of a spectrum considered as one object, and any cue that makes the ear treat two things as two objects reduces it.
The measurement, and how it is made
The curve above is a model, and it is worth being clear what the underlying measurement is, because the experiment shapes the result.
The procedure is a threshold-tracking task. A masker is presented continuously and a brief probe tone is added at a chosen frequency; the probe’s level is moved up and down according to the listener’s answers until the level at which it is detected on some fixed fraction of trials is bracketed. Repeating over probe frequencies traces one curve.
Two known distortions come with it, and both are named rather than hidden.
Beats contaminate the region near the masker. When the probe is within a few hertz of the masker the two beat against each other, and the listener detects the beating rather than the tone — a much easier task, which pushes the measured threshold down and produces a notch in the curve right at the masker’s own frequency. Published masking patterns for tonal maskers usually show it. The model here does not reproduce it, and the figure would be wrong in that small region if it claimed to.
Combination tones do the same thing further out. A loud masker and a probe generate frequencies in the ear that were in neither, and the listener may detect one of those rather than the probe. This is a real effect and it has its own essay: the ear makes its own sound, and one of the things it makes is a detection cue for an experiment that was not asking about it.
What the picture cannot show
The figure draws a threshold, which is a binary. Masking is not binary. Below the curve a sound is undetectable; a few decibels above it the sound is detectable but quieter than it would have been alone, and it goes on being partially masked for another twenty or thirty decibels. Partial masking is the condition most music is actually in, and the curve has nothing to say about it.
The excitation model does, and it needs nothing this collection has not already built for the loudness of a chord: the partial loudness of a probe is what the whole sonority weighs with it, less what the sonority weighs without it. Running that for one note at a fixed level against a ten-note tutti, swept across the register:
| the probe | its loudness alone | what it keeps under the tutti |
|---|---|---|
| F2, 87 Hz | 22.9 sones | 4.3% |
| C4, 262 Hz | 38.0 | 4.1% |
| A5, 880 Hz | 41.8 | 15.0% |
| C7, 2,093 Hz | 34.9 | 37.2% |
| G7, 3,136 Hz | 29.0 | 46.5% |
A piccolo keeps nearly ten times the share a bassoon keeps, at the same level in the same texture, which is the orchestration claim made three sections above and now has a number on it. What it does not have is the reason given there.
The asymmetry predicts the opposite. Masking runs upward, so a probe above the tutti has the whole orchestra’s spectrum throwing a sheet at it and a probe below the tutti has nothing above it to be reached by. On the direction of masking alone the bassoon should win, and it loses by a factor of nine.
What is actually happening is a question of where the masker’s energy is rather than which way it spreads. The upward sheet decays at six to sixteen decibels per Bark, and three kilohertz is many Barks above anything in a normal tutti, so almost nothing survives the journey. A bassoon at 87 hertz is not being reached from below by anything — it is sitting in the same critical band as the bass line, the cellos’ fundamentals and the lower partials of everything else, and it is masked by its neighbours rather than by its inferiors. The register that wins is the empty one, and in an orchestra the empty one is the top.
That is a better rule than the direction rule and it makes a different prescription. Do not ask which way the masking runs; ask where the spectrum is thin, and put the line that has to be heard there.
The figure also draws one masker at a time. Real signals have many, their patterns overlap, and the combined threshold is not a simple sum: two maskers together mask more than either alone but less than their arithmetic total, and how much more is a question about whether the ear is adding intensities or excitations. Codecs use a fudge factor here, chosen by listening tests, and it is the least principled number in the whole apparatus.
Finally, and most importantly, the figure is drawn for tones. A noise masker of the same total level produces a different and generally deeper pattern, because it fills the band continuously and offers no beating cue. Tonal maskers and noise maskers are separated in every codec by an explicit tonality estimate for exactly this reason.
That asymmetry is the whole of why loudness changes which sounds hide which. A quiet accompaniment masks only its own neighbourhood; a loud one throws a shadow upward across the spectrum, and the note that disappears is not the nearest one but the one above.
That last case is worth stating plainly, because it connects masking to a claim the site makes elsewhere. A bell has no single pitch partly because its partials are not harmonically related — and partly because, not being harmonically related, they fail to hide each other, so several of them survive as separately audible objects. Fusion and masking are doing the same work.
Where the ladder goes next
Everything above is simultaneous: masker and probe at the same instant. The effect does not stop at the instant. A sound raises the threshold for about a fifth of a second after it ends, which is not surprising, and for a few milliseconds before it begins, which is — a sound can hide what came before it, and that is the next rung.
Sideways from here sits the other consequence of the same filter bank. If two tones inside one band are not resolved from each other, they are not merely masking each other; they are beating, and roughness is what that feels like. Masking and roughness are two readings of one piece of machinery, and the essays about consonance have been leaning on it since the site’s first phase.
Part 1 of 8
One essay in the series on masking. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 34.
- A chord is not as loud as its notes
- A clarinet keeps what a string loses
- A loud chord is a smaller chord
- An entrance is a change of colour
- Room is used up by whoever enters first
- The arch belongs to hearing, not to the series
- The dynamics are in the score already
- The listener is given the top voice, and the bass as a sine
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Bark scaleCritical bandwidthMaskingOrchestrationPerceptual codingSpreading function
- A low chord stops being rough by stopping being a chord critical bandwidth, masking, spreading function
- The chord that has room for an entrance critical bandwidth, masking, orchestration
- A chord is a register critical bandwidth, orchestration
- A page has two decibels critical bandwidth, orchestration
- An equal note cannot be masked critical bandwidth, masking
- The dissonance arrives and the dynamic does not critical bandwidth, orchestration