The first wavefront wins
Assumes: Two ears, and the whole of the difference is 655 microseconds
Working out where a sound is from a delay of a few hundred microseconds is impressive in a free field and looks impossible in a room. A hall sends the listener the direct sound and then, over the next tenth of a second, dozens of reflections from walls, floor, ceiling and every hard object in the space — each a full copy of the sound, each arriving from a different direction with its own interaural delay, most of them wrong.
If the ear averaged over the evidence it would place every source at the centre of the room. It does not. It answers as though the reflections were not there.
The picture has four regions and each is a different thing happening.
Under a millisecond: one sound, pulled
With a delay shorter than about a millisecond the two copies are inside a single period of most of the audible spectrum, and the ear does not treat them as two events at all. It computes one interaural delay from the combination, and the result is a single image sitting somewhere between the two sources, pulled towards the earlier one.
This is summing localisation, and it is the entire basis of stereo. Two loudspeakers playing the same signal at the same level produce an image midway between them — a phantom source where there is no source — and adjusting the relative level or the relative timing moves it. Every pan control ever built is exploiting this region.
It is worth noticing what a strange thing that is. A stereo image is not a reproduction of a spatial field; it is a systematic exploitation of a failure mode, in which the auditory system is handed evidence it was never designed to receive and produces an answer that happens to be useful.
One to thirty-five milliseconds: one sound, in the first place
Past about a millisecond the two copies are separately resolvable in principle — and the second one is denied. The image stays firmly at the position of the first arrival, and the second contributes loudness and timbre while contributing nothing at all to direction.
This is localisation dominance, and it is the region that makes rooms usable. The direct sound from an instrument always arrives first, because it takes the shortest path. Every reflection is by definition later. So a rule that says believe the first arrival is a rule that says believe the direct sound, and it is exactly right almost all of the time.
The strength of the suppression is the surprising part. Helmut Haas measured it for speech in 1949 and found that the delayed copy has to be about ten decibels louder than the direct sound before it starts to take over the perceived direction. The second sound is not merely ignored; it is discounted by a full order of magnitude in power.
Past the echo threshold: two sounds
Eventually the second copy escapes. It becomes a separately audible event with its own location, and the listener hears an echo.
Where that happens depends enormously on the sound, and this is the number most often quoted wrongly.
For clicks and other sharp transients: five to ten milliseconds. A click has a single unmistakable onset, and a second onset ten milliseconds later is a second event.
For speech: thirty to fifty milliseconds. A continuous signal with its own internal fluctuation hides a delayed copy far more effectively, and the copy has to be a good deal later before it stands out.
For sustained music: later still, and highly variable — a legato orchestral texture can absorb reflections at eighty milliseconds that would be obvious echoes in a spoken announcement.
The often-cited “35 milliseconds” comes from Haas’s speech experiments and is a fact about speech. Applying it to a snare drum is wrong by a factor of five in the dangerous direction.
What this makes possible: the delay line in the ceiling
The clearest practical use is sound reinforcement, and it is a piece of engineering built directly on the middle region of the figure.
A large hall with a speaker on a platform needs loudspeakers down the room, because the direct sound from the platform is far too quiet at the back. But a loudspeaker halfway down the hall is nearer the back rows than the platform is, so its sound arrives first — and by localisation dominance, the audience at the back hears the voice coming from a box above their heads. The effect is disconcerting and it is the reason badly designed reinforcement sounds wrong even when it is perfectly intelligible.
The fix is to delay the distributed loudspeakers so that their sound arrives after the direct sound from the platform, but by less than the echo threshold. The direct sound wins the localisation and the loudspeaker supplies the level. In practice a delay is chosen equal to the acoustic propagation time between the platform and the loudspeaker, plus a deliberate margin of ten to fifteen milliseconds to make certain the platform is first.
The margin has to fit in the window this essay is about. Too small and the loudspeaker occasionally wins; too large and the reinforcement becomes an audible echo. The whole design lives between one millisecond and thirty-five, and the numbers come from psychoacoustics rather than from acoustics.
The rule is on-axis, and an audience is not
That account is exact along the line through the platform and the loudspeaker, and it is exact for a reason worth seeing: on that line the two path lengths grow at the same rate, so the arrival difference at every seat behind the loudspeaker is the margin and nothing else. Which is why the rule can be a single number, and why it has never needed a picture.
An audience is two-dimensional. Off the axis the two paths grow at different rates, the triangle inequality gives something away, and the arrival difference becomes the margin plus a term belonging to the room. In a hall twenty metres by thirty-five with a pair of delay loudspeakers sixteen metres downstream, that term is 16.8 milliseconds across the seats those loudspeakers are for — half the whole precedence window, present before any margin is added, and invisible to the arithmetic that designed the system.
Classifying every seat against its own echo threshold, which depends on how much louder the loudspeaker is than the direct sound there:
| signal | margins at which every served seat stays fused |
|---|---|
| an orchestral chord | 1 to 63 ms |
| speech | 1 to 23 ms |
| a click | none, at any margin |
The middle row is the practice, vindicated and bounded. Ten to fifteen milliseconds sits comfortably inside the band, so the rule of thumb is a good one — but the band’s ceiling is 23 rather than the 35 this essay has been quoting, because the room term has already spent a third of the window. And the failure is not symmetric, which the sentence above gets wrong: too small bites hard, with 44 per cent of the served seats falling into summing localisation at a margin of zero, while too large does not bite at all until a margin no installer would use.
The bottom row is the one that matters, and this essay has already said why without applying it. A click has no workable margin. The best any value achieves is 94 per cent of the served seats at a margin of one millisecond, and by ten — the bottom of the recommended range — not one served seat is inside the window. Reinforcing a spoken voice through a delay ring is a solved problem; reinforcing a snare drum through the same ring is a system that delivers an audible echo to the entire audience, and the design that produced it was correct at every step, using the number for the wrong sound.
A hall, on two clocks
The precedence window and reverberation time are the same room measured on scales that differ by a factor of fifty, and putting them side by side is the quickest way to see what a hall is doing.
The division of labour is clean and it is why both quantities are worth having. The first thirty-five milliseconds decide where the sound is and what it sounds like: the direct arrival fixes the direction and the early reflections colour the timbre and add loudness without being separately heard. Everything after that decides how the space feels — its size, its liveness, whether a chord blooms or dies — and contributes nothing to direction at all.
A hall is engineered on both clocks at once, and the classic problem of hall design is that the two want different things. A strong early reflection from a nearby side wall is desirable: it arrives inside the window, adds level, and broadens the apparent source. A strong late arrival from a distant rear wall is a defect. The same reflecting surface can supply either depending only on where it is, which is why the shoebox shape of the nineteenth-century halls has survived every attempt to improve on it — its side walls are close and its rear wall is far, so its early reflections are strong and its long-path arrivals are weak.
Below a few hundred hertz a small room does not have reflections at all; it has modes, standing waves at frequencies fixed by its dimensions, and the notion of a “first arrival” stops being useful because the sound in the room is not travelling anywhere. The precedence effect is a story about rays, and rays are a high-frequency approximation — so everything in this essay is a claim about the part of the spectrum where a room can be described by paths, and about nothing below it.
Which computation produced the numbers
The three boundaries are published ranges from listening experiments, carried in a small table and drawn as bands, and nothing here derives them. Everything else is a length divided by a speed: sound travels 343 metres a second, so a reflection from a wall five metres further away than the direct path arrives 14.6 milliseconds later.
The reinforcement census is that division done at every seat rather than on the axis, over a grid at half-metre spacing, with the loudspeaker gain fixed by the requirement that it equal the direct sound at the rearmost centre seat and with fifteen decibels of off-axis attenuation for seats in front of the delay line. That last figure is not decoration: an omnidirectional loudspeaker in the same position delivers an echo to the front corners at every margin, and directivity is the reason a real installation does not.
That division is worth doing once, because it puts the whole essay onto a scale a person can picture.
One millisecond is 34 centimetres. The summing-localisation region is the size of a head and a bit — which is not a coincidence, since the largest delay a single source can produce between two ears is 655 microseconds.
Thirty-five milliseconds is twelve metres of extra path. In a hall twenty metres wide, a reflection from a side wall reaches a listener in the centre with an extra path of a few metres and arrives well inside the suppression window. In a hall sixty metres long, the reflection from the back wall reaches the front rows with forty metres of extra path — 117 milliseconds — and is a plain echo. Back-wall echo is the classic pathology of a long hall and it is a geometry problem with a psychoacoustic threshold.
Whose halls, and what was written for them
The window is a listener’s property and does not vary between centuries. What varies is what was built and what was written for it, and the pairing is close enough to be worth stating.
Music written for very reverberant spaces avoids fast harmonic change. A large stone church has a reverberation time of five seconds or more, so a chord is still sounding when the next four have arrived. Renaissance polyphony written for such buildings moves slowly in harmony and quickly in counterpoint, which is the combination that survives: the lines are distinguishable because they are separated in register and in onset, and the harmony is legible because it does not depend on hearing a progression cleanly.
Music written for a hall assumes early reflections. A classical or romantic orchestral score assumes an audible bass line, clear harmonic rhythm at the bar, and articulation that survives to the back of the room. Those require a reverberation time between one and a half and two seconds — which is what a shoebox hall of the period delivered.
And music written for a recording assumes none of it. A close-miked studio recording has almost no early reflection structure at all, and the reverberation is added afterwards and chosen. That is a genuinely new condition, less than a century old, and it is why studio production developed its own vocabulary — the artificial early-reflection pattern of a reverberation unit is a design parameter with no counterpart in any earlier music.
None of these is a claim that the acoustics caused the music. It is a claim that a repertoire and a building are constrained by the same listener, and that the constraints are visible on both sides.
What the picture cannot show
The listener’s two ears, which the figure has collapsed into one axis. Everything drawn here is a delay between two sounds; what the mechanism actually works on is the interaural delay each of them carries, and the suppression is a decision about which of several conflicting spatial estimates to keep. A figure with only time on it cannot show a mechanism whose input is direction.
One reflection, and rooms have hundreds. Every figure here is a direct sound and a single copy. A real room delivers a dense cloud in which no individual reflection is distinguishable, and the precedence effect in that condition is much harder to characterise. What is known is that suppression still operates and that its strength depends on the density of the arrivals as well as their delays.
The effect breaks down and then rebuilds. Present a click pair repeatedly and the suppression strengthens over the first few presentations — the auditory system learns the room. Change the geometry suddenly, by swapping which of the two sources leads, and the suppression collapses for a few presentations before rebuilding. This is the Clifton effect, and it shows that the mechanism is not a fixed reflex but something closer to a model of the room being maintained and revised.
Timbre is not on the axes at all. A reflection inside the suppression window is inaudible as an event and thoroughly audible as a change of quality: it adds comb filtering to the spectrum, which is heard as colouration. The whole of what a hall does to the sound of an instrument in its first fifty milliseconds is happening in a region this figure draws as “one sound, coloured”, and the word is carrying an entire subject.
The generalisation, and where the model stops
The rule the ear is following is a hypothesis about the world: sounds do not spontaneously duplicate themselves; a second copy of something arriving shortly after the first came off a surface, and surfaces do not make sound. It is an excellent hypothesis and it is why the effect exists.
It also states its own limits. The effect should fail when the world stops working that way, and it does.
Two genuinely independent sources close in time are misplaced. Two players attacking a chord fifteen milliseconds apart do not produce two locations; the second is pulled towards the first. In an ensemble that is a feature — it is part of what makes a section sound like an instrument rather than like twelve players — and it means the spatial resolution of an ensemble is much worse than the spatial resolution of a single source.
And it makes echo-cancelling hard for the wrong reason. Anything designed to reproduce a room — a recording, a simulation, a spatial audio system — has to reproduce reflections whose individual audibility is suppressed but whose collective effect is the entire character of the space. What must be got right is not what a listener can point at, but what a listener cannot.
Where the ladder goes next
This rung and the one before it are the localisation ladder’s foundation and its complication. What remains open is the thing the figure calls colouration: a reflection that is not heard as a sound but is heard as a change in the sound. That is where this ladder meets the room, and how long a room rings is the same physical space measured on a much longer clock — Sabine’s reverberation time starts where the precedence window ends.
Sideways, the same fixed window keeps appearing. Thirty-five milliseconds is the scale of expressive timing, it is the scale on which one sound masks another that has already finished, and it is the scale at which an onset difference will break a note into two objects. Four literatures, one number, and the number is the length of the auditory present.
Part 2 of 12
One essay in the series on localisation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 22.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Early reflectionsEcho thresholdHaas effectPrecedence effectReinforcementSumming localisation