A position and a width
Assumes: An echo is prevented by the crowd around it · A room with directions in it
An echo is prevented by the crowd around it laid the published echo threshold over the echogram and sorted a hall’s arrivals into two boxes: heard as a separate event, or not. Its last paragraph said what that sort was throwing away.
Fusion is not a yes or a no: a reflection that is not heard as a separate event still moves the apparent source, widens it, and colours it, and the published apparatus for that — image shift, apparent source width, colouration — takes the same three inputs this figure already has.
It does. The echogram the fifth rung built carries an arrival time, a level and a direction for every image source in the room, and those three quantities are the inputs to all three of these. Nothing new is measured here; a list that was being asked one question is asked three.
Three quantities out of one list
The position is where the sound seems to be. The first wavefront wins is the essay about why it is not simply the average of the arrivals: the earliest one dominates by a very large margin, and a reflection has to be about ten decibels louder than the direct sound to pull the image as hard. That trading ratio is the published quantity, and with it and a discount for how late each arrival is, the image is a weighted mean of directions in which the direct sound holds most of the weight.
The width is how large the source seems. It is not a property of the source at all — a violin on a platform is a metre wide and can be heard as five — and the quantity it tracks is the early lateral energy fraction, which the fifth rung already computes: the sideways-weighted energy arriving between five and eighty milliseconds, over the total up to eighty.
The colouration is what a strong early reflection does to the spectrum. A single reflection at delay Δt and amplitude ratio g is a comb filter: notches every 1/Δt hertz, of a depth 20 log₁₀((1+g)/(1−g)). Both numbers are in the echogram.
The three are independent readings of one list, and they do not vary together across a hall.
Down the hall, one of them moves
Taking seats along the centre line of a symmetric hall gives a clean result and a check.
The image shift is zero at every seat, exactly. A symmetric room’s side walls produce mirror-image arrivals whose directions cancel, so the weighted mean of directions is the direct sound’s own. That is not a measurement; it is a consequence of the geometry, and it is here as a check that the arithmetic has the signs right.
The width rises and then flattens. At eight metres from the platform the apparent source is twenty degrees wide; by twenty metres it is thirty-one; beyond that it hovers between twenty-seven and thirty-two. The mechanism is that the direct sound falls off as one over the distance while the reflected energy is roughly constant across a hall, so the lateral fraction — which has the direct sound in its denominator — grows until the reflections dominate and then stops.
The colouration deepens and its notches move apart. At the front the strongest early reflection is 17 decibels deep with notches every 400 hertz; at the back it is 26 decibels with notches every three kilohertz. Both changes are the floor bounce getting relatively later, in the sense that the direct path lengthens faster than the reflected one.
So a listener walking back through a hall gets a source that widens and then stops widening, a spectral comb that gets deeper and coarser, and an image that does not move at all.
Across the hall, the other one does
Moving sideways instead gives the complementary picture.
The image shift grows toward the walls, from nothing on the centre line to twelve degrees at a seat nine metres off it, and it is always toward the near wall. A listener at the side of a hall hears the orchestra pulled away from where it is, in the direction of the wall they are sitting against.
That is the direction that surprises people who expect the reflection to push the image away. It does not: a reflection adds a source in its own direction, and the near wall’s reflection arrives from the near side. Precedence keeps the image close to the direct sound and the pull is what is left over.
The width barely changes across the hall, from thirty to thirty-three degrees, because the lateral fraction at a fixed distance from the platform is not a strong function of which side of the room the seat is on.
And the colouration barely changes either, at about twenty-five decibels everywhere, because it is dominated by the floor bounce — a vertical reflection whose geometry depends on the distance and not on the sideways offset.
So the two axes of a hall carry different information. Depth decides how wide and how coloured; width decides how far the image is displaced. A seat is not one number.
The floor is the reflection nobody designed
The colouration reading has one arrival in it almost everywhere, and it is worth naming because it is the least discussed surface in a hall.
Every seat in every room has a hard floor about a metre below the listener’s ears and a source a few metres away, so there is always a reflection from the floor arriving less than a millisecond after the direct sound. Sub-millisecond delays give comb notches spaced kilohertz apart, which is a coarse spectral change spanning the whole audible range rather than a fine one.
And its amplitude ratio is high — the reflected path is barely longer than the direct one and the floor absorbs little — so the comb is deep. At the seats in this hall it runs from 17 to 26 decibels.
That is by far the largest single spectral event in the model, and it is produced by a surface that nobody treats as an acoustic element. Ceilings are shaped, side walls are angled, rear walls are diffused, and the floor is whatever the seats are bolted to.
There is a partial defence: the same comb is present when a listener hears anything at all, indoors or out, so the ear has every opportunity to learn to discount it. The ear builds objects is the ladder about what a listener does with a set of components that arrive together, and a reflection under a millisecond is well inside every fusion window there is.
What a hall designer would take from this
The three quantities correspond to three things halls are praised and blamed for, and the correspondence is not one to one.
Width is what “spaciousness” means, and it is why halls are built narrow. A narrow hall puts its side walls close to the seats, which gets lateral energy to the listener early and raises the lateral fraction; a wide fan-shaped hall sends its first reflections from the ceiling, which is not lateral, and sounds narrower. This collection has that argument on the room ladder and it is the standard one.
It also interacts with everything else a hall is judged on. The first eighty milliseconds are a different room shows that the early energy decides clarity, and the lateral fraction is a slice of the same energy sorted by direction — so a hall trying to be both clear and wide is spending the same eighty milliseconds twice, and how long a room rings is the third claim on it.
Colouration is what a hall is blamed for when a seat sounds hollow or nasal. A 25-decibel comb with notches every three kilohertz is a very audible spectral change, and it is produced by the floor — which is the one reflecting surface in a hall that nobody thinks of as an acoustic design element and which every seat has directly under it.
Image shift is what nobody complains about, and the figure suggests why: it is small in the middle of the hall where most people sit, and where it is large — at the sides — the listener already has a strong visual cue about where the platform is, and vision dominates auditory localisation comprehensively.
That last point is worth stating as a limitation rather than as a finding. The figure computes an auditory image, and a listener with their eyes open does not have one.
What was thrown away, and what it was worth
It is worth being explicit about how much information the two-box sort was discarding, because that is the case for this rung existing.
At a middle seat in this hall the census finds three hundred and six arrivals inside the window it examines and calls none of them an echo. The sixth rung’s answer to “what does this seat sound like” is therefore a single word — clean — and the same list, asked differently, gives a direction, a width and a spectrum.
That is the general shape of the thing. A threshold applied to a rich list produces a verdict, and a verdict is the smallest possible summary of the list. The sixth rung needed the verdict, because the question it was asking was whether a hall has an echo problem; every other question a listener might ask about the same list needs something else.
There is a caution attached, and it is the reason the sixth rung stopped where it did. Turning a list into a verdict uses one published criterion; turning it into three continuous readings uses three, and two of them are much weaker than the echo threshold. What has been bought is a richer description at the cost of leaning on softer evidence, and each of the three should be read with that in mind.
Which computation produced the numbers
The echogram is the fifth rung’s image-source model: every arrival up to order six in a shoebox room, with spherical spreading and a stated absorption per reflection, giving a time, an energy and a unit direction vector for each.
The position is a weighted mean of azimuths relative to the direct sound’s, with weight 1 for the direct arrival and, for each fused reflection, its amplitude ratio times 10^(−10/20) for the trading ratio times exp(−Δt/10 ms) for the delay discount. Only arrivals below the echo threshold are included, because an arrival heard as a separate event is not contributing to one image.
The width is a + b√LF with the lateral fraction from the fifth rung’s own definition. The coefficients are chosen to put a typical hall in the range published measurements report, and they are ordinal — the figure claims the shape of the curve across the hall, not the degrees.
The colouration is the comb of the strongest early reflection alone, which is the standard first-order treatment and which the section below qualifies.
Where the model stops
The trading ratio and the delay discount are asserted. Ten decibels and ten milliseconds are both published figures with wide ranges around them, and sweeping both across those ranges says exactly how much of this essay they are carrying.
At the side seat that gives twelve degrees, the shift over trading ratios from 5 to 20 decibels and discounts from 6 to 20 milliseconds runs from 3.1 degrees to 21.1 — a factor of seven, with the two constants pulling the same way. A weaker trading ratio and a longer discount both let the reflections count for more, and the extreme corner of the sweep is nearly two octaves of shift away from the other one.
So no number of degrees in this essay is worth more than its order of magnitude, and twelve should be read as “of the order of ten” rather than as a measurement.
The shape is untouched by any of it, and that is the half the argument rests on. The image shift on the centre line stays at zero — under a tenth of a degree at every one of the twenty settings tried, against a side seat’s three to twenty-one — because it is a cancellation between mirror-image arrivals and not a quantity. The direction never changes either: the pull is toward the near wall at every trading ratio and every discount. Geometry decides which way and whether; the two constants decide only how far.
The width coefficients are not a model. They are a monotone function of a quantity that is known to track apparent source width, fitted to nothing. What the figure supports is that the width follows the lateral fraction, which is the established result; the degrees are a scale.
One reflection colours. A real spectrum is combed by every early arrival at once, which is not the sum of the individual combs — the notches of several reflections interleave and partly fill each other in. The single strongest is an upper bound on the depth and an underestimate of the raggedness.
And the direct sound is not always first. Under a balcony or behind a column the direct sound can be attenuated below a reflection, and the whole precedence apparatus inverts. The model has no obstructions in it.
What the picture cannot show
It cannot show the eyes. Auditory localisation is comprehensively overridden by vision when the two disagree by less than about twenty degrees, which is most of the range this figure computes. A concert listener watching the platform does not have the image the figure describes.
Nor can it show the signal. The echo threshold depends enormously on what is being played — a click’s is a few milliseconds and an orchestral chord’s over a hundred — so which arrivals count as fused, and therefore all three of these readings, change with the music. The figure is drawn for a sustained chord.
It cannot show the distance. How far away the room takes over puts a boundary at the point where the reflected field exceeds the direct sound, and every seat here is well past it — which is precisely why the width reading rises and then flattens, and is a constraint the figure inherits without stating.
And it cannot show two ears properly. Everything here is a direction per arrival, and a real listener extracts direction from a time difference and a level difference between two ears in a head that shadows and diffracts. Where the two ears stop agreeing is the essay about how much that matters, and none of it is in this model.
Whose halls, and when
The room is a shoebox of 22 by 40 by 15 metres with a mean absorption of 0.18, which is a plain rectangular concert hall of the nineteenth-century type — the Musikverein, the Concertgebouw, Boston Symphony Hall. That shape is the one the lateral-fraction argument was developed to explain, and it is the one the figure flatters.
The measurements the width reading is built on date from the 1960s and 1970s, when a generation of acousticians set out to say why the old rectangular halls sounded better than the new fan-shaped ones, and arrived at lateral energy as the answer. That is a specific historical claim about a specific comparison and it has held up.
What has not been settled the same way is colouration. The floor reflection is present in every hall of every shape and nobody designs it, and the figure’s largest single number — a 25-decibel comb — is a consequence of the fact that listeners sit above a hard surface.
Where this ladder goes next
Seven rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity in neither ear’s signal; the frequency above which the two ears stop agreeing; a room with directions in it; the threshold that sorts them; and now the three things the ones below the threshold do instead of nothing.
What is owed after this is the head. Every direction in this ladder above the fourth rung is a vector from a seat to an image source, and a listener does not have vectors — they have two pressure signals, differing in time and in level by amounts a head decides. The second and fourth rungs have the head’s own arithmetic, and running the echogram through it would give the quantity that is actually available: the interaural difference at every instant of a hall’s response, which is what a binaural recording is and what every result in this ladder has been standing in for.
Part 7 of 12
One essay in the series on localisation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
ColourationEchogramHall designLateral fractionLocalisationPrecedence effectReflectionSpaciousness