A sound hides what came before it
A tone hides another tone that is sounding at the same time, and the region it hides is asymmetric in frequency. Everything about that account assumed the two sounds were simultaneous. Nothing about hearing is.
The right-hand side of that figure is a mechanism nobody finds surprising. The left-hand side is a claim that a sound has been affected by something that had not happened yet, and it needs a better account than “the ear is slow”.
Forward masking, which is the easy half
A loud sound leaves the auditory system in a state that takes time to recover, and while it is recovering it is less sensitive. Three mechanisms are usually offered and they are not exclusive.
Ringing. The mechanical response of a place on the basilar membrane does not stop when the input does; a filter with a bandwidth of 100 Hz has an impulse response several milliseconds long, and while it is ringing at the masker’s frequency it is a poor detector of anything else. This accounts for the first few milliseconds and nothing beyond.
Adaptation. Auditory nerve fibres fire hard at the onset of a sound and then settle to a lower rate; after the sound stops, the rate is below its resting value for tens of milliseconds. A probe arriving during that dip has less to work with. This accounts for a great deal of the middle of the curve.
Persistence. The representation of the masker survives centrally for longer than either of the above, and a probe has to be discriminated from it rather than merely detected. This accounts for the long tail and is the least mechanical of the three.
The measured decay is close to linear in the logarithm of the delay, which is why the figure’s horizontal axis is what it is, and it reaches the unmasked threshold at somewhere between 100 and 200 ms depending on the masker’s level and duration. The one thing everybody agrees on is that the recovery is much longer than any mechanical ringing could explain.
Backward masking, which is the hard half
A probe presented five milliseconds before a masker is harder to detect than the same probe in silence. The effect is real, it has been replicated for sixty years, and it does not have a settled explanation.
What can be said about it is unusually specific:
It is short. Ten to twenty milliseconds, against forward masking’s two hundred. The two are not the same effect running in two directions.
It varies between listeners more than almost anything else in psychoacoustics. Published thresholds differ by twenty or thirty decibels between individuals with identical audiograms.
It gets much smaller with practice. A listener who has done the task a hundred times shows a fraction of the effect they showed at the start. Nothing that is happening in the cochlea does that.
That last point is the one that decides the interpretation. Forward masking is at least partly peripheral. Backward masking is not peripheral at all: it is a claim about how a brief event is processed, and the standard reading is that processing takes time, that a loud event arriving during the processing of a quiet one interrupts it, and that what is being masked is not the sound but the analysis of the sound.
That reading has a consequence worth stating plainly. It means the ear does not deliver events in the order they arrive; it delivers them after an interval, and events inside that interval compete rather than queue.
The window, and what else lives in it
The number that comes out of all of this is an integration window of a few tens of milliseconds, and it turns up in several places on this site that were not obviously related.
Twelve of the thirty partials are below it. A loud chord is not the sum of its notes’ spectra, and the shortfall is not small — which is the simultaneous version of the temporal effect this essay is about, and the one that is easier to compute because nothing has to be remembered.
A groove is measured in this window. The systematic timing departures that make a rhythm feel a particular way are twenty to fifty milliseconds — the same scale as backward masking and well inside the integration window. That is a striking coincidence and it should not be over-read: nothing here shows that a groove is a masking phenomenon, and the honest statement is that expressive timing operates at the scale where the ear stops treating events as separately timestamped, which is at least why the effect is felt rather than counted.
A room’s early reflections are in it too. The first fifty milliseconds of a room’s response are integrated with the direct sound rather than heard after it, and the boundary between “part of the note” and “an echo” is at the same scale. Three quite different literatures — masking, room acoustics and rhythm perception — have independently arrived at a few tens of milliseconds as the size of the auditory present.
How fast music can be before it stops being legible
Forward masking puts an upper bound on how quickly a listener can be given distinct events, and the bound is close enough to real musical tempi to be interesting.
A note recovers from the previous note’s masking over a hundred to two hundred milliseconds. Ten notes a second — a fast run, a trill, a tremolo — puts a hundred milliseconds between onsets, which is squarely inside the recovery. Each note in such a passage is heard against the residue of the one before it, and the effect is not that the notes become inaudible but that they stop being separately audible: the run becomes a gesture with a shape rather than a sequence with members.
That is exactly how fast passages are used. A scale run in a concerto is not written to have its notes counted; it is written to sweep. A trill is notated as an ornament rather than as thirty-two demisemiquavers because what it produces is a single object with a quality, and the notation is honest about that. The threshold at which counting becomes sweeping is a listener’s property and it sits where this figure says it does.
The rate, computed rather than gestured at
“Squarely inside the recovery” is a phrase, and the curve it refers to will give a number. Take a run of notes all at one level, ask how much masking each one still carries from the note before it, and read off the rate.
| notes a second | gap | 40 dB | 55 dB | 70 dB | 85 dB |
|---|---|---|---|---|---|
| 4 | 250 ms | 0.0 | 0.0 | 0.0 | 0.0 |
| 6 | 167 | 1.5 | 2.2 | 3.0 | 3.7 |
| 8 | 125 | 3.8 | 5.7 | 7.6 | 9.6 |
| 10 | 100 | 5.6 | 8.5 | 11.3 | 14.1 |
| 20 | 50 | 11.3 | 16.9 | 22.5 | 28.2 |
At ten notes a second — the rate a fast run reaches — a loud passage leaves each note eleven decibels of masking from its predecessor and a quiet one leaves under six. Below four a second nothing is masked at all, at any level, which is a plain statement about where the effect starts to matter: it is not a constraint on melody and it is a constraint on ornament.
Turning it round gives the number the section wants. The rate at which each note still carries six decibels from the one before it is 10.5 a second at 40 dB and 6.7 at 85 — so the tempo at which a run stops delivering separate notes moves by a factor of a little over one and a half across an orchestral dynamic range. A run that is legible at piano is a smear at fortissimo, at the same tempo, and the performer’s remedy — play the loud one slower — is exactly what the arithmetic asks for.
One caution about the second figure, because it is a caution about this whole model. Its caption says a quiet masker’s effect is not only smaller but shorter, and that is true of the drawing rather than of the arithmetic: the recovery time is an input to the function, and the two figures pass 200 milliseconds and 120. The model’s own level dependence is entirely in the depth of the masking, which starts at ten decibels below the masker whatever the masker is. The published literature does report shorter recovery at lower levels, so the figure is not wrong; it is asserting rather than deriving, and the table above is the part that comes out of the model instead of going into it.
Whose music. That is a claim about European keyboard and orchestral writing from roughly 1700 onward, where the run-as-gesture is a standing device and the notation for it is stable. It generalises unevenly. A gamelan’s fastest interlocking parts are played by two players alternating precisely so that neither is producing events at the rate the ear would merge, which is a different solution to the same limit; and the tabla repertoire treats very fast strokes as a syllable with a name, which is a third. Every tradition that plays fast has had to decide what happens at ten events a second, and they have not decided the same thing.
The shortest silence
The complementary measurement to all of the above is gap detection: the shortest silence that can be noticed in an otherwise continuous sound. For a broadband noise it is two to three milliseconds. That is a remarkably small number and it is the one place where the auditory system looks fast rather than slow.
The apparent contradiction — a system with a fifty-millisecond window that resolves a three-millisecond gap — is not a contradiction, and the resolution is the most useful thing in this essay. Integration and resolution are different operations on different quantities. The window over which energy is summed for the purpose of deciding how loud something is runs to tens of milliseconds. The machinery that detects a change is much faster, because a change produces an onset and onsets are what the system is built to find.
Music uses both. A staccato quaver is short enough that its loudness is reduced by the integration window, and its silence is long enough to be unmistakable. The note is quieter than it should be and its separation from the next note is perfectly clear, and both facts come from the same ear.
Which computation produced the numbers
The curves above are the standard fits rather than a derivation, and the difference matters more here than in most essays on this site.
Forward masking is drawn as a decay that is linear in log delay from the end of the masker to the unmasked threshold, with the amount of masking at short delays set by the masker’s level. That is the shape reported by Jesteadt, Bacon and Lehman in 1982 and by many others since, and the constants are theirs rather than anything derived here.
Backward masking is drawn as a shorter decay of the same kind, and it is drawn with much less confidence, because the published range for its size at a given delay spans tens of decibels. A single curve through that is a summary, not a measurement, and the essay’s claim rests on the sign of the effect and its duration, both of which are robust, rather than on its size, which is not.
The site’s habit is that every claim gets a test it could fail. The test here is negative and it is in the figure’s own construction: if the backward branch had been drawn at the same length as the forward one, the picture would assert a symmetry that no experiment reports, and the shape would have been invented to look tidy.
The artefact that made this audible to everybody
Temporal masking’s practical consequence arrived with digital audio compression, and it arrived as a failure.
A codec works in blocks. It analyses a window of the signal, computes a masked threshold from it, and spends bits accordingly — the frequency-domain half of that computation is the previous rung. The quantisation noise it introduces is spread evenly across the whole block.
Now put a sharp transient near the end of a block: a castanet, a triangle, a snare rim. The block’s average energy is high, so the codec allows a lot of quantisation noise. That noise is spread across the entire block, including the part before the transient, which was silent.
The result is a smear of noise arriving before the attack. It is called pre-echo, and it is audible because backward masking is only ten to twenty milliseconds long while a block is fifty. The noise that arrives in the last ten milliseconds before the strike is masked; the noise that arrives forty milliseconds before it is not, and it sounds like a little burst of hiss preceding the hit.
Every argument in this essay is invisible in a waveform, because a waveform draws amplitude against time and masking is a statement about frequency and level. The asymmetry is the mechanism: a loud sound reaches upward across the spectrum and a quiet one does not, so what a sound hides depends on how loud it was and not only on when.
The fix is block switching: detect a transient, drop to a much shorter window for that region, and the noise is confined to a window short enough to be masked. Every modern codec does it, and the detector that decides when to switch is the least psychoacoustic and most heuristic part of the whole design. When a codec’s transient detector misfires, pre-echo is what is heard, and it is the single most recognisable artefact of lossy audio.
What the picture cannot show
The figure is drawn for a probe and a masker at the same frequency. Move them apart and the whole thing weakens sharply — forward masking is far more frequency-specific than simultaneous masking is, which is more evidence that different mechanisms are involved.
It says nothing about what happens with many events. Real music is a continuous stream of onsets a few hundred milliseconds apart, each partially masking its neighbours in both directions, and the combined effect is not a sum of these curves. Nobody has a model of that which works on real signals.
The level dependence is thinner than the figures suggest. In the arithmetic, a masker’s level sets how deep the masking is at the start and nothing else; the depth at any delay is a fixed fraction of the masker’s level minus ten decibels. Everything the figures show about a quiet masker recovering sooner comes from a recovery time supplied to the drawing rather than computed by it, and the table above deliberately holds that time fixed so that what is left is the model’s own contribution.
Duration is missing entirely. A masker’s forward effect grows with how long it lasted, up to about 200 ms, and then stops growing. A short burst and a long note of the same level leave very different holes behind them, and the figure draws one burst.
And the vertical axis is a detection threshold, not an audibility. Everything at or below the curve is undetectable; everything above it is detectable and, for another twenty or thirty decibels, quieter than it would have been in silence. Music lives almost entirely in that partially-masked region, which neither this figure nor the frequency-domain one draws at all. Both figures show the boundary of a region and say nothing about the interior, and the interior is where the notes are.
Measured timing departures in recorded performances run to tens of milliseconds, which is the same order as the forward masking this essay is about. That coincidence is worth naming and is not evidence of anything: the two quantities are independent, and a performance’s departures are large enough that masking cannot be inferred from a score’s rhythm alone.
A consequence, and it is one of the few places where a psychoacoustic limit explains a notational one. A fixed physical delay is a different note value at every tempo, so a groove is unwritable in principle; and a delay small enough to be inside the integration window is not heard as a delay at all but as a quality of the event. Notation records order and proportion. It has no symbol for a quantity that is not perceived as a duration, and this is a whole class of them.
Where the ladder goes next
Masking is the ear removing things. The complementary question is how it assembles what is left: given a mixture of sounds from several sources, arriving continuously, which parts belong together. That is the ear building objects and it turns on the same few tens of milliseconds — a sequence fast enough splits into streams, and how fast is fast enough is the next figure.
Sideways from here is the room. The reason a reflection ten milliseconds behind a direct sound is not heard as a second sound is not masking exactly, but it lives in the same window and is measured in the same units: the first wavefront wins, and how late a reflection has to be before it stops winning is the number that decides whether a hall is clear or muddy.
Part 2 of 8
One essay in the series on masking. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 17.
- A soft chord has to fade in
- A dissonance has to last
- An equal note cannot be masked
- An exit is worth nothing until the tutti is given up
- Loud is relative, and it comes down slowly
- The dissonance arrives and the dynamic does not
- A bass chord low enough to balance has already hidden its tenor
- A clarinet keeps what a string loses
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Backward maskingForward maskingIntegration windowPre-echoTemporal maskingTransient