Perception and the listener

A room with directions in it

Every room until now has been a reservoir of energy that drains at a rate. That model has no directions in it at all, so it cannot say the one thing every published measure of spaciousness is about: how much of what arrives comes from the side. Mirror the source in six walls and every reflection acquires an angle and a time — and the answer to why a concert hall is narrow falls out at eighteen metres.

Assumes: Where the two ears stop agreeing · The first wavefront wins

The fourth rung of this ladder found the frequency above which a room stops sending the two ears comparable signals, and computed the spaciousness that follows from it. Its last paragraph said what it had assumed:

Lateral energy is the term every published spaciousness measure weights and this model omits; computing it needs a room with directions in it, which means an image-source model rather than a statistical one.

That is not a small omission and it is not this ladder’s alone. Sabine’s model, which nine rungs of the room ladder are built on, has no directions in it whatever. It is a reservoir of energy with an absorbing boundary, and every quantity it produces — reverberation time, clarity, critical distance — is a scalar. A hall in it has a volume and a surface and nothing else, and two halls of the same volume and surface are the same hall.

They are not the same hall, and the thing that distinguishes them has been sitting outside every model this site has built.

Why a concert hall is narrowThe lateral energy fraction at the middle seat as the same hall is widened, everything else held. It peaks at 12 metres across at 0.235 and falls to 0.000 at 44. A wide hall's side walls are further away, so their reflections arrive later, weaker and — this is the part Sabine's model cannot say — from nearer the front, where the sideways weighting discounts them. The shoebox halls the orchestral repertoire was written for are all between about eighteen and twenty-five metres wide, and this is the arithmetic they are the answer to.12 m1520253035404500.050.10.150.20.25hall width, metreslateral energy fractionthe shoebox hallsare 18 to 25 macross, and thisis the arithmeticthey answer
Fig. 1 The lateral energy fraction at the middle seat as one hall is widened and everything else held: same depth, same height, same absorption, same source, same seat. It peaks at eighteen metres and falls away in both directions, and by forty-four metres it has lost more than half of what it had. The shoebox halls the orchestral repertoire was written for are all between about eighteen and twenty-five metres wide.

The construction, which is one line

A reflection off a flat wall is geometrically identical to a direct sound from a source mirrored in that wall. That is all the image-source method is, and it has no approximation in it beyond flatness.

In one dimension, for a room of length L with a source at s, the mirror images sit at

xn = nL + (n even ? s : L − s)

and |n| counts how many times the sound has reflected. The zeroth is the source itself; the first is its mirror in the far wall; the minus-first is its mirror in the near wall. A shoebox is that rule applied three times, so there is one image per integer triple, and each of them is a point in free space at a computable distance and a computable direction from any seat.

That gives every reflection three things a statistical model cannot: when it arrives, from its distance; how strong it is, from spherical spreading and what it lost at each surface; and where it comes from, which is the direction of the vector from the seat to the image.

The room, and the sources it is equivalent to. A 22 by 40 metre hall from above, with the 25 image sources of order 3 or less that lie in its own plane. A reflection off a wall is the same thing as a direct sound from a source mirrored in that wall, so a room full of reflections is a lattice of sources in free space and every one of them has an angle and a distance. The seat sits 16.0 metres from the real source; the nearest image is 16.2 metres away and arrives 0.7 milliseconds later.
Fig. 2 A twenty-two by forty metre hall from above, with the images of order three and below that lie in the plane of the floor. The lattice extends in every direction and is cut off here for legibility; the full computation to six reflections has 377 of them. A room full of reflections is a room-shaped grid of sources in an infinite empty space, and every one of them is a direction as well as a delay.

What lateral means, exactly

The measure this rung is for is the lateral energy fraction, and it is defined in a way that is worth stating precisely because two of its details do most of the work.

LF = ∫5ms80ms E·cos²θ  /  ∫080ms E

The numerator is measured with a figure-of-eight microphone pointed sideways, whose response is the cosine of the angle from the lateral axis, so its energy weight is that cosine squared. The denominator is measured with an omnidirectional one.

The numerator starts at five milliseconds and the denominator at zero. So the direct sound is excluded from the top and included in the bottom, deliberately, because the quantity being asked about is how much sideways sound arrives relative to everything, and the direct sound is the thing it is competing with.

That detail is why the front seats of a hall score badly, and the reason is not what it looks like.

Lateral energy at five seats in a 22 by 40 metre hall. The lateral energy fraction — sideways-weighted energy arriving between 5 and 80 milliseconds, over everything arriving in the first 80 — computed from the image sources of a shoebox with 18 per cent absorption. It is highest at mid side at 0.23 and lowest at front centre at 0.12, and the reason the front seat is worst is not that its reflections are weak: it is that the direct sound is 6 metres away rather than 18, and the direct sound is in the denominator.
Fig. 3 Five seats in one hall. Front centre has the lowest lateral fraction at 0.116 and rear centre is not much better at 0.154; the best is mid-side at 0.232. A listener in the front row is not short of side reflections — they are getting the same walls as everybody — but the direct sound reaching them from six metres away is four times the energy that reaches the middle from sixteen, and the direct sound is in the denominator.

Why a hall has a best width

The hero figure is the result this rung was opened for, and it is a design fact rather than a perceptual one.

Widen a hall and three things happen to its side-wall reflections at once. They arrive later, because the walls are further away. They arrive weaker, because energy spreads as the inverse square. And — this is the term a statistical model has no way to represent — they arrive from a smaller angle, because a wall that is further to the side is also, in the geometry of a rectangle, further ahead or behind, so the vector from the seat to its image swings toward the median plane where the cosine-squared weighting discounts it.

All three run the same way, and against them runs only one thing: in a very narrow hall, the side reflections arrive inside the first five milliseconds and drop out of the numerator entirely — they are heard as part of the direct sound rather than as a separate arrival, which is the precedence effect doing exactly what it does.

So there is a maximum, and the arithmetic puts it at about eighteen metres — for the seat the hero figure is computed at. Which of the hall’s other dimensions that number depends on is the question a design claim has to answer, and the answers separate sharply.

held fixed, swept best width
absorption, 0.08 to 0.35 18.0 m at every value
height, 9 to 24 m 16 to 18 m
depth, 24 to 60 m 10.0 to 20.5 m
seat, front fifth to back fifth 10.0 to 32.0 m

The optimum is completely insensitive to absorption and nearly insensitive to height, and it is not a property of the hall at all — it is a property of the seat. Widening the hall from a depth of 24 metres to 60 moves the best width from ten to twenty; moving the listener from the front fifth to the back fifth of one hall moves it from ten to thirty-two.

The two sweeps are the same sweep, because both of them are moving the listener away from the source. The best width tracks roughly the distance from the platform to the seat: about ten metres of width for a listener four metres back, eighteen for one sixteen metres back, thirty-two for one twenty-eight metres back. The geometry is not mysterious — the angle a side wall’s image subtends at a seat is set by the ratio of the half-width to the distance, and holding that ratio near its best value means growing the width with the distance.

So a hall of one width cannot be optimal for its own seats, and the eighteen metres is the answer for the middle of a forty-metre hall and for nobody else in it. The front third of that hall wants a ten-metre room and the back fifth wants a thirty-two-metre one, which is a factor of three across one audience.

That is a better statement than the one this section was reaching for, and it is what a fan-shaped hall is an attempt at: a room whose width grows with its depth is a room trying to hold the ratio constant down its length. It fails for a different reason — splayed side walls send their reflections toward the back rather than across — which is why the shape with the good measured numbers is the one that gets the ratio wrong everywhere except the middle and gets the directions right. The shoebox is not the shape that maximises this quantity. It is the shape that keeps it from collapsing anywhere, and the eighteen to twenty-five metres real halls are built at is the width at which the middle two thirds of the audience are all somewhere near their own optimum.

Every arrival at one seat, against the delay at which it would be an echo. The echogram: each reflection at its delay after the direct sound and its level relative to it, in a 22 by 40 by 15 metre hall with 18 per cent absorption. The line is the published echo threshold for speech — 40 milliseconds at equal level and about 3.0 more for each decibel of attenuation — so anything to the RIGHT of it is late enough and loud enough to be heard separately. Nothing here stands clear enough of its neighbours to be heard as an echo, and 29 arrivals are fused with the direct sound instead.
Fig. 4 What arrives at one seat in the first 128 milliseconds, reflection by reflection, with the heavy lines those coming from the side. The first eighty milliseconds is the window the measure integrates over, and it is the same window the clarity measure uses for a completely different purpose. Beyond it the arrivals become too dense to count, which is where a room stops being a room and where the statistical model takes over. Sabine’s model can produce the envelope of this picture and cannot produce a single one of its lines.

The other number the same construction gives

Once every reflection has a time and a level, a second quantity comes out of the identical list at no extra cost, and it is one this site already has by another route.

Clarity is the ratio of the energy arriving in the first eighty milliseconds to everything after it, in decibels, and it is what the first eighty milliseconds are a different room is about. Sabine’s model gives it from a decay rate and a distance; the image-source list gives it by adding up two subsets of the same arrivals. That the two agree to about a decibel across the seats above is worth stating, because it is the only external check this construction has: a model with directions in it must still get the direction-free quantity right, and it does.

Where they part company is in the scatter. The statistical model says clarity is a smooth function of distance from the source, so two seats equidistant from the platform have the same clarity. The image-source list says they do not — a seat under a side wall gets a strong early reflection that a seat in the middle of the same row does not, and the difference is a decibel or two. Every hall’s measured clarity map is patchy in exactly that way, and the statistical model has no way to be patchy.

One decay, two verdicts, and the line is the listener's. The early-to-late energy ratio against reverberation time, for 50 millisecond and 80 millisecond windows. Nothing about the room differs between the curves; only where the line is drawn across its decay. Zero comes at 1.00 seconds for the 50 ms window and 1.59 seconds for the 80 ms window — which are, to two figures, the published rules of thumb for a room for speech and a room for music. The design targets were not put in; they came out.
Fig. 5 Clarity from the statistical model, for a hall of this volume. The image-source computation reproduces its numbers to about a decibel at every seat drawn above, which is the check on the construction. What it adds is the departure from them: a smooth prediction is the average of a lumpy reality, and the lumps are where the walls are.

The same geometry answers a question about the one reflection that is not part of the spatial impression at all, because it arrives too late to be.

How deep a hall has to be before its rear wall is late enough to be an echo. The once-reflected rear wall at a seat 16 metres from the source, against hall depth. Up the page is how far past its own echo threshold the reflection arrives — the threshold rises with attenuation, so a deeper hall raises it too, and the question is which rises faster. Delay wins: the reflection crosses zero at about 31 metres for speech and 41 for an orchestral chord, so every concert hall anybody builds has a rear wall arriving late enough and loud enough to be heard as an echo. It is not heard as one, and the reason is not in this figure.
Fig. 6 The once-reflected rear wall at a seat sixteen metres from the source, against how deep the hall is. Up the page is how far past its own echo threshold the reflection arrives — attenuation raises that threshold as the hall deepens, so the question is which rises faster.

Delay wins. A deeper hall quietens its rear wall and delays it by more than the quietening buys back, so past a certain depth the reflection stops contributing to spaciousness and starts being heard as a separate event. Every reflection in this essay has an angle and a distance; this is the one where the distance takes the angle out of the account entirely.

What this connects to that it should not have to

The fourth rung of this ladder found two limits set by the width of the head — the interaural delay’s phase ambiguity at 762 hertz and the diffuse field’s coherence zero at 980 — and concluded that everything two ears do well, they do below a kilohertz.

The lateral energy fraction is a quantity with no ears in it at all. It is measured with two microphones, one of them a figure-of-eight, and it has no frequency in its definition beyond the octave bands it is conventionally averaged over. Yet it is the measure that correlates best with what listeners call spaciousness, and the reason is presumably the one the previous rung computed: a sound arriving from the side produces a different signal at each ear, and difference between the ears is what a wide auditory image is made of.

So the two rungs are the same statement in two vocabularies, and neither can be derived from the other with what this site has. The fourth rung has the ears and no directions; this one has the directions and no ears. Joining them would mean computing, for each image source, the interaural difference it produces at a real head — which needs the head model the fourth rung built and the geometry this one does, and is the shortest path to a model that says why a wide image is pleasant rather than merely why it is wide. It is also the only way to get at what a listener who cannot separate two sources is actually being given.

Where a room stops sending the two ears the same soundThe correlation between the two ears' signals against frequency, for a seat 15 metres from the source in a 15,000 cubic metre room with a 2 second reverberation time. The pale curve is the diffuse field alone — sin(kd)/(kd) for an ear separation of 17.5 centimetres, which first crosses zero at 980 hertz. The heavy curve adds the direct sound, which is coherent and lifts the whole thing by an amount the direct-to-reverberant ratio sets. At 125 hertz the coherence is 0.98 and at 1000 it is 0.16.980 Hz62.51252505001k2k4k8k-0.20.20.40.60.81interaural coherencefrequency, hertzthe diffuse field alonesin(kd) / kdwith the direct soundat 15 mspaciousness 0.31at 500 Hz
Fig. 7 The other half of the pair, from an earlier essay: how alike the signals at the two ears are in a diffuse field, against frequency, for a head 17.5 centimetres across. This has ears and no directions; the figures above have directions and no ears. The essay that joins them is the one still owed.

Which computation produced the numbers

The hall is a rectangular box 22 metres wide, 40 deep and 15 high — the proportions of a nineteenth-century shoebox, and close to the Musikverein’s. The source is on the platform 4 metres in and 1.6 up; seats are at ear height, 1.2 metres.

Images are enumerated to a total reflection order of six, which is 377 of them within the cutoff. Six is not arbitrary: the shape of the width curve is present at two reflections and stops moving by about six, and the figure’s own handle sweeps it so that a reader can check that the answer is the geometry rather than the truncation.

The source is a point radiating equally in all directions, which no instrument is — an instrument points, and a violin aimed at a side wall is a different set of image strengths from the same violin aimed at the audience. Each image’s energy is one over the square of its distance times the surviving fraction after each reflection, with 18 per cent absorbed per bounce — a hall with wood and plaster and an audience in it. Its lateral weight is the square of the component of its direction along the hall’s short axis, which is a listener facing the platform.

Two conventions in the calculation are worth naming because they are choices.

Reflections are specular and surfaces are flat. A real hall’s walls have boxes, columns, statuary and a coffered ceiling, all of which scatter rather than reflect. Scattering redistributes energy toward later times and toward all directions, which raises the denominator and averages the numerator’s cosine toward a half — so a real hall’s lateral fraction is flatter across seats than this model says.

Absorption is one number for all surfaces and all frequencies. A real hall’s floor, ceiling, walls and audience differ by a factor of several, and the audience alone absorbs far more at high frequency than at low. The published lateral fraction is an average over four octave bands for exactly that reason, and running this model per band would need the band-by-band absorption the room ladder already has.

Where the model stops

There is no head in it. The cosine-squared weight is a microphone’s directivity, not an ear’s, and it was chosen historically because a figure-of-eight microphone is an instrument that exists. Whether a listener weights a reflection from thirty degrees the way that microphone does is a question about head-related transfer functions, and the answer is certainly not exactly.

There is no diffraction. Below a few hundred hertz a wavelength is comparable with the objects in a hall, and the image-source method’s assumption that sound travels in straight lines and bounces is worst exactly there — which is unfortunate, because the fourth rung’s whole finding is that the interesting binaural behaviour is below a kilohertz.

And a shoebox is not a hall. Real halls have balconies, raked floors, side galleries and a ceiling that is not flat, and every one of those is a departure from the one geometry this method handles exactly — including the raked floor, which is what makes an audience absorb what it does and is drawn nowhere here. The image-source method generalises to any polyhedron and the arithmetic gets rapidly worse; a fan-shaped hall — which the width curve above predicts should be bad, and which the literature agrees is bad — cannot be computed this way at all without a great deal more machinery.

Whose halls, and when

The shoebox is a specific object with a date. The Musikverein in Vienna opened in 1870 at about 20 metres wide; the Concertgebouw in 1888 at about 28; Boston Symphony Hall in 1900 at about 23. Those three are, by common consent, among the best halls in the world, and they were built before anybody could compute anything about them — the Boston hall is the first designed with any acoustic theory at all, and the theory was Sabine’s, which is the model with no directions in it.

So the width was not chosen for the reason this figure gives. It was chosen because it is roughly what a masonry roof can span, and the number that fell out of the structural constraint happens to sit on the maximum of a curve nobody drew until the 1970s.

Which makes the fan-shaped halls of the mid twentieth century the natural experiment. They were built wide because a fan seats more people with better sightlines, and they are consistently described as sounding distant and lacking envelopment. The hero figure says why: at forty metres across, the lateral fraction at the middle seat is less than half what it is at eighteen, and the walls that would have supplied it are too far away and at the wrong angle.

What the picture cannot show

Whether more lateral energy is better without limit. The measure correlates with spaciousness, and spaciousness is one of several things a listener wants; too much of it is a hall in which an orchestra sounds smeared and a soloist cannot be located. The curve here has a maximum in a quantity, not in a preference, and the two are not the same picture.

And it cannot show the listener’s own orientation. The cosine-squared weight assumes a listener facing the platform. A listener who turns their head has just changed every angle in the numerator, which means the measure is a property of a pose as much as of a seat — and the thing a real listener does with a spacious sound is precisely to stop turning their head, because there is nothing to locate.

Where this ladder goes next

Five rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity in neither ear’s signal; the frequency above which the two ears stop agreeing; and now a room with directions in it and the term every published measure weights.

The rung after it is the one the precedence effect has been owed since the second. The first wavefront wins established that a reflection inside a window of a few tens of milliseconds is fused with the direct sound rather than heard as an echo, and it could not say which reflection wins or when a particular hall crosses the threshold — because it had no way to know when any individual reflection arrived. It does now: the echogram above is a list of arrivals with times, levels and angles, and the echo threshold is a published function of exactly those three. Running one against the other would say, for a given seat in a given hall, which reflections are fused, which are heard separately, and which single surface is responsible.

Part 5 of 12

One essay in the series on localisation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

ClarityEarly reflectionsImage-sourceInteraural time differenceLocalisationReverberationRoom acousticsSpaciousness