A room with directions in it
Assumes: Where the two ears stop agreeing · The first wavefront wins
The fourth rung of this ladder found the frequency above which a room stops sending the two ears comparable signals, and computed the spaciousness that follows from it. Its last paragraph said what it had assumed:
Lateral energy is the term every published spaciousness measure weights and this model omits; computing it needs a room with directions in it, which means an image-source model rather than a statistical one.
That is not a small omission and it is not this ladder’s alone. Sabine’s model, which nine rungs of the room ladder are built on, has no directions in it whatever. It is a reservoir of energy with an absorbing boundary, and every quantity it produces — reverberation time, clarity, critical distance — is a scalar. A hall in it has a volume and a surface and nothing else, and two halls of the same volume and surface are the same hall.
They are not the same hall, and the thing that distinguishes them has been sitting outside every model this site has built.
The construction, which is one line
A reflection off a flat wall is geometrically identical to a direct sound from a source mirrored in that wall. That is all the image-source method is, and it has no approximation in it beyond flatness.
In one dimension, for a room of length L with a source at s, the mirror images sit at
and |n| counts how many times the sound has reflected. The zeroth is the source itself; the first is its mirror in the far wall; the minus-first is its mirror in the near wall. A shoebox is that rule applied three times, so there is one image per integer triple, and each of them is a point in free space at a computable distance and a computable direction from any seat.
That gives every reflection three things a statistical model cannot: when it arrives, from its distance; how strong it is, from spherical spreading and what it lost at each surface; and where it comes from, which is the direction of the vector from the seat to the image.
What lateral means, exactly
The measure this rung is for is the lateral energy fraction, and it is defined in a way that is worth stating precisely because two of its details do most of the work.
The numerator is measured with a figure-of-eight microphone pointed sideways, whose response is the cosine of the angle from the lateral axis, so its energy weight is that cosine squared. The denominator is measured with an omnidirectional one.
The numerator starts at five milliseconds and the denominator at zero. So the direct sound is excluded from the top and included in the bottom, deliberately, because the quantity being asked about is how much sideways sound arrives relative to everything, and the direct sound is the thing it is competing with.
That detail is why the front seats of a hall score badly, and the reason is not what it looks like.
Why a hall has a best width
The hero figure is the result this rung was opened for, and it is a design fact rather than a perceptual one.
Widen a hall and three things happen to its side-wall reflections at once. They arrive later, because the walls are further away. They arrive weaker, because energy spreads as the inverse square. And — this is the term a statistical model has no way to represent — they arrive from a smaller angle, because a wall that is further to the side is also, in the geometry of a rectangle, further ahead or behind, so the vector from the seat to its image swings toward the median plane where the cosine-squared weighting discounts it.
All three run the same way, and against them runs only one thing: in a very narrow hall, the side reflections arrive inside the first five milliseconds and drop out of the numerator entirely — they are heard as part of the direct sound rather than as a separate arrival, which is the precedence effect doing exactly what it does.
So there is a maximum, and the arithmetic puts it at about eighteen metres — for the seat the hero figure is computed at. Which of the hall’s other dimensions that number depends on is the question a design claim has to answer, and the answers separate sharply.
| held fixed, swept | best width |
|---|---|
| absorption, 0.08 to 0.35 | 18.0 m at every value |
| height, 9 to 24 m | 16 to 18 m |
| depth, 24 to 60 m | 10.0 to 20.5 m |
| seat, front fifth to back fifth | 10.0 to 32.0 m |
The optimum is completely insensitive to absorption and nearly insensitive to height, and it is not a property of the hall at all — it is a property of the seat. Widening the hall from a depth of 24 metres to 60 moves the best width from ten to twenty; moving the listener from the front fifth to the back fifth of one hall moves it from ten to thirty-two.
The two sweeps are the same sweep, because both of them are moving the listener away from the source. The best width tracks roughly the distance from the platform to the seat: about ten metres of width for a listener four metres back, eighteen for one sixteen metres back, thirty-two for one twenty-eight metres back. The geometry is not mysterious — the angle a side wall’s image subtends at a seat is set by the ratio of the half-width to the distance, and holding that ratio near its best value means growing the width with the distance.
So a hall of one width cannot be optimal for its own seats, and the eighteen metres is the answer for the middle of a forty-metre hall and for nobody else in it. The front third of that hall wants a ten-metre room and the back fifth wants a thirty-two-metre one, which is a factor of three across one audience.
That is a better statement than the one this section was reaching for, and it is what a fan-shaped hall is an attempt at: a room whose width grows with its depth is a room trying to hold the ratio constant down its length. It fails for a different reason — splayed side walls send their reflections toward the back rather than across — which is why the shape with the good measured numbers is the one that gets the ratio wrong everywhere except the middle and gets the directions right. The shoebox is not the shape that maximises this quantity. It is the shape that keeps it from collapsing anywhere, and the eighteen to twenty-five metres real halls are built at is the width at which the middle two thirds of the audience are all somewhere near their own optimum.
The other number the same construction gives
Once every reflection has a time and a level, a second quantity comes out of the identical list at no extra cost, and it is one this site already has by another route.
Clarity is the ratio of the energy arriving in the first eighty milliseconds to everything after it, in decibels, and it is what the first eighty milliseconds are a different room is about. Sabine’s model gives it from a decay rate and a distance; the image-source list gives it by adding up two subsets of the same arrivals. That the two agree to about a decibel across the seats above is worth stating, because it is the only external check this construction has: a model with directions in it must still get the direction-free quantity right, and it does.
Where they part company is in the scatter. The statistical model says clarity is a smooth function of distance from the source, so two seats equidistant from the platform have the same clarity. The image-source list says they do not — a seat under a side wall gets a strong early reflection that a seat in the middle of the same row does not, and the difference is a decibel or two. Every hall’s measured clarity map is patchy in exactly that way, and the statistical model has no way to be patchy.
The same geometry answers a question about the one reflection that is not part of the spatial impression at all, because it arrives too late to be.
Delay wins. A deeper hall quietens its rear wall and delays it by more than the quietening buys back, so past a certain depth the reflection stops contributing to spaciousness and starts being heard as a separate event. Every reflection in this essay has an angle and a distance; this is the one where the distance takes the angle out of the account entirely.
What this connects to that it should not have to
The fourth rung of this ladder found two limits set by the width of the head — the interaural delay’s phase ambiguity at 762 hertz and the diffuse field’s coherence zero at 980 — and concluded that everything two ears do well, they do below a kilohertz.
The lateral energy fraction is a quantity with no ears in it at all. It is measured with two microphones, one of them a figure-of-eight, and it has no frequency in its definition beyond the octave bands it is conventionally averaged over. Yet it is the measure that correlates best with what listeners call spaciousness, and the reason is presumably the one the previous rung computed: a sound arriving from the side produces a different signal at each ear, and difference between the ears is what a wide auditory image is made of.
So the two rungs are the same statement in two vocabularies, and neither can be derived from the other with what this site has. The fourth rung has the ears and no directions; this one has the directions and no ears. Joining them would mean computing, for each image source, the interaural difference it produces at a real head — which needs the head model the fourth rung built and the geometry this one does, and is the shortest path to a model that says why a wide image is pleasant rather than merely why it is wide. It is also the only way to get at what a listener who cannot separate two sources is actually being given.
Which computation produced the numbers
The hall is a rectangular box 22 metres wide, 40 deep and 15 high — the proportions of a nineteenth-century shoebox, and close to the Musikverein’s. The source is on the platform 4 metres in and 1.6 up; seats are at ear height, 1.2 metres.
Images are enumerated to a total reflection order of six, which is 377 of them within the cutoff. Six is not arbitrary: the shape of the width curve is present at two reflections and stops moving by about six, and the figure’s own handle sweeps it so that a reader can check that the answer is the geometry rather than the truncation.
The source is a point radiating equally in all directions, which no instrument is — an instrument points, and a violin aimed at a side wall is a different set of image strengths from the same violin aimed at the audience. Each image’s energy is one over the square of its distance times the surviving fraction after each reflection, with 18 per cent absorbed per bounce — a hall with wood and plaster and an audience in it. Its lateral weight is the square of the component of its direction along the hall’s short axis, which is a listener facing the platform.
Two conventions in the calculation are worth naming because they are choices.
Reflections are specular and surfaces are flat. A real hall’s walls have boxes, columns, statuary and a coffered ceiling, all of which scatter rather than reflect. Scattering redistributes energy toward later times and toward all directions, which raises the denominator and averages the numerator’s cosine toward a half — so a real hall’s lateral fraction is flatter across seats than this model says.
Absorption is one number for all surfaces and all frequencies. A real hall’s floor, ceiling, walls and audience differ by a factor of several, and the audience alone absorbs far more at high frequency than at low. The published lateral fraction is an average over four octave bands for exactly that reason, and running this model per band would need the band-by-band absorption the room ladder already has.
Where the model stops
There is no head in it. The cosine-squared weight is a microphone’s directivity, not an ear’s, and it was chosen historically because a figure-of-eight microphone is an instrument that exists. Whether a listener weights a reflection from thirty degrees the way that microphone does is a question about head-related transfer functions, and the answer is certainly not exactly.
There is no diffraction. Below a few hundred hertz a wavelength is comparable with the objects in a hall, and the image-source method’s assumption that sound travels in straight lines and bounces is worst exactly there — which is unfortunate, because the fourth rung’s whole finding is that the interesting binaural behaviour is below a kilohertz.
And a shoebox is not a hall. Real halls have balconies, raked floors, side galleries and a ceiling that is not flat, and every one of those is a departure from the one geometry this method handles exactly — including the raked floor, which is what makes an audience absorb what it does and is drawn nowhere here. The image-source method generalises to any polyhedron and the arithmetic gets rapidly worse; a fan-shaped hall — which the width curve above predicts should be bad, and which the literature agrees is bad — cannot be computed this way at all without a great deal more machinery.
Whose halls, and when
The shoebox is a specific object with a date. The Musikverein in Vienna opened in 1870 at about 20 metres wide; the Concertgebouw in 1888 at about 28; Boston Symphony Hall in 1900 at about 23. Those three are, by common consent, among the best halls in the world, and they were built before anybody could compute anything about them — the Boston hall is the first designed with any acoustic theory at all, and the theory was Sabine’s, which is the model with no directions in it.
So the width was not chosen for the reason this figure gives. It was chosen because it is roughly what a masonry roof can span, and the number that fell out of the structural constraint happens to sit on the maximum of a curve nobody drew until the 1970s.
Which makes the fan-shaped halls of the mid twentieth century the natural experiment. They were built wide because a fan seats more people with better sightlines, and they are consistently described as sounding distant and lacking envelopment. The hero figure says why: at forty metres across, the lateral fraction at the middle seat is less than half what it is at eighteen, and the walls that would have supplied it are too far away and at the wrong angle.
What the picture cannot show
Whether more lateral energy is better without limit. The measure correlates with spaciousness, and spaciousness is one of several things a listener wants; too much of it is a hall in which an orchestra sounds smeared and a soloist cannot be located. The curve here has a maximum in a quantity, not in a preference, and the two are not the same picture.
And it cannot show the listener’s own orientation. The cosine-squared weight assumes a listener facing the platform. A listener who turns their head has just changed every angle in the numerator, which means the measure is a property of a pose as much as of a seat — and the thing a real listener does with a spacious sound is precisely to stop turning their head, because there is nothing to locate.
Where this ladder goes next
Five rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity in neither ear’s signal; the frequency above which the two ears stop agreeing; and now a room with directions in it and the term every published measure weights.
The rung after it is the one the precedence effect has been owed since the second. The first wavefront wins established that a reflection inside a window of a few tens of milliseconds is fused with the direct sound rather than heard as an echo, and it could not say which reflection wins or when a particular hall crosses the threshold — because it had no way to know when any individual reflection arrived. It does now: the echogram above is a list of arrivals with times, levels and angles, and the echo threshold is a published function of exactly those three. Running one against the other would say, for a given seat in a given hall, which reflections are fused, which are heard separately, and which single surface is responsible.
Part 5 of 12
One essay in the series on localisation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
ClarityEarly reflectionsImage-sourceInteraural time differenceLocalisationReverberationRoom acousticsSpaciousness
- A rest needs a dry room reverberation, room acoustics
- How far away the room takes over localisation, reverberation
- Only the player hears a staccato end reverberation, room acoustics
- The beat that is not in the air interaural time difference, localisation
- The error that moves straight ahead interaural time difference, localisation
- Two ears, and the whole of the difference is 655 microseconds interaural time difference, localisation