The hall through a head
Assumes: A position and a width · Two ears, and the whole of the difference is 655 microseconds
A position and a width gave this ladder its seventh rung and ended by naming what all seven of them assume:
Every direction computed so far is a vector from a seat to an image source, and a listener does not have vectors — they have two pressure signals, differing in time and in level by amounts a head decides.
The head’s arithmetic has been here since the second rung. Two ears and one difference computes the interaural delay a spherical head produces at every angle, and the fourth adds the level difference that takes over where the delay stops being decidable. Neither has ever been shown a room.
Position and width, in the right units
The seventh rung’s phrase was that a room does not give a listener two boxes — a direct sound and a diffuse remainder — but a position and a width. In interaural units both of those are numbers.
At the centre of the stalls the direct sound arrives at zero microseconds and the energy-weighted mean of the whole response is also zero, which it has to be by symmetry. The root-mean-square spread of the reflections about that mean is 315 microseconds, against a geometric maximum of 656.
So the picture a listener has of a hall from the best seat in it is a point at dead centre with a cloud around it half as wide as the head allows. The position is exact and the width is enormous, and both come out of one line applied to each of fifty-four reflections.
The seat moves the whole image and the standard metric barely notices
Move eight metres to the side and the mean interaural delay becomes 158 microseconds.
A listener detects a change of about fifteen microseconds in an interaural delay. So a side seat shifts the entire acoustic image of the hall by ten times a threshold, and the direct sound with it.
Over exactly the same seats the lateral energy fraction — the measure concert-hall acoustics actually reports, which this ladder’s fifth and sixth rungs computed — moves from 26 per cent to 30. That is a few points on a scale whose whole useful range is about twenty, and it is a measure of width rather than of position: it is designed to be insensitive to which way the energy comes from, only how far off the median plane.
So a hall’s standard number is nearly blind to a change that dominates what a listener hears. That is not a criticism of the measure, which was built to describe spaciousness rather than image position. It is a statement about what it leaves to the seat, and this collection now has both numbers on one axis.
The width, meanwhile, barely moves: 315 microseconds at the centre and 322 at the side. A side seat changes where the hall is and not how big it is, which is the opposite of what “sitting off to the side” is usually described as doing.
And then the head throws away the back of the room
Woodworth’s interaural delay is r(θ + sin θ)/c, and it rises from zero straight ahead to 656 microseconds at ninety degrees. Past ninety it comes back down, reaching zero again at 180. The level difference does the same.
A sound at 120 degrees produces exactly the delay and exactly the level difference of one at 60. That is the cone of confusion, it is a property of any head shaped like a sphere, and it means the two cues this ladder has been using cannot tell front from back at all.
In a hall that matters, because a hall is behind a listener as well as in front. At a centre-stalls seat two of fifty-four reflections arrive from behind and carry two per cent of the energy. Move to the back of the room and it is forty-eight of a hundred and fifteen, carrying thirty per cent.
Every one of those is drawn on the first figure exactly where a front reflection from its mirror angle would be, and nothing in the signals separates them.
So the reverberant field a listener receives is not the reverberant field the image-source model computes. The model knows the difference between a first reflection from the side wall ahead and one from the side wall behind; a two-cue listener does not, and this collection’s account of localisation is a two-cue account.
It is worth being clear that this is not a small correction to the geometry. Thirty per cent is not a rounding error on the reverberant field; it is the part of a hall that arrives from the rear wall, the rear side walls and every second-order path that goes back before it comes forward — which on the standard picture of what makes a hall enveloping is much of what is doing the work.
The cue that is left, and it is a different one
The one thing that does survive the fold is how much energy arrives late relative to the direct sound, which has no direction in it at all and is the quantity a hall’s reverberation time describes. So a two-cue listener at the back of a hall has an accurate sense of how reverberant the room is and a folded sense of where the reverberation is.
That distinction has a name in the literature — source width against listener envelopment — and it is usually drawn between early and late energy. This figure draws it somewhere else: between the energy whose direction survives the head and the energy whose direction does not, which cuts the same response along a different line.
What that means for everything above
The sixth and seventh rungs sorted reflections by a prominence threshold and asked what the ones below it do. The answer they gave was that they contribute to a width rather than to a position, which survives here and gets a number.
What does not survive is the direction half of the sixth rung’s picture. A room with directions in it treated the echogram’s azimuths as available, and half of them are not — not because they are weak, but because the geometry of a head maps two of them to one signal.
Real listeners do better than this, and how they do it is not on this page. Pinnae impose a direction-dependent notch on the spectrum above about six kilohertz, which breaks the front-back symmetry, and head movement breaks it by a completely different route: turning the head changes an interaural delay in opposite directions for a front source and a back one. Neither is in a spherical head and both are why the confusion is a laboratory phenomenon rather than an everyday one.
That is worth saying plainly. This rung’s finding is a limitation of the model, and it is a limitation the model shares with every claim this ladder has made about direction.
The direct sound is the one thing the fold does not touch
There is a reassuring half to this and it is worth stating before the frequency argument.
The direct sound arrives from straight ahead by construction, where both curves are at zero and both are steepest. The interaural delay changes by 8.9 microseconds per degree near the median plane and by 4.5 near ninety, so the head’s own arithmetic is twice as sensitive where a listener is pointed as it is at the sides.
That is why the first wavefront wins is such a strong effect and why it has to be. The one arrival whose direction matters is the one the head resolves best, and every arrival that is folded, compressed or ambiguous arrives later and is suppressed by the precedence mechanism before its direction is asked for.
So the ladder’s own order of business turns out to be the head’s. The prominence threshold of the sixth rung and the fold of this one are not two independent facts: the reflections the threshold discards are largely the ones whose direction the head could not have told anybody anyway.
What the fold does damage is the width, because a width is built from exactly those discarded arrivals. Thirty per cent of the energy at a rear seat is interaurally indistinguishable from energy in front of the listener, so the cloud on the first figure is symmetric in a way the room is not.
The two cues do not describe the same hall
There is one more thing the conversion exposes, and it is about frequency.
An interaural delay has no frequency in it: the geometry is the geometry, and 656 microseconds is 656 microseconds at every pitch. An interaural level difference is entirely a function of frequency, because a head only shadows what is small compared with itself. At the same seat and the same reflections, the mean absolute level difference is 0.3 decibels at 250 hertz, 2.2 at a kilohertz, and 8.8 at eight.
And the two do not simply add. Above about 762 hertz — half a wavelength shorter than the longest delay the head can make — an interaural phase is ambiguous, so the delay cue stops working on the fine structure exactly where the level cue starts working.
So a hall’s low frequencies are located by timing and its high frequencies by level, and this figure is two different pictures of one room stacked on top of each other. The seventh rung’s width is a timing width. A level width would be a different number, and the two need not agree about where anything is.
That gap between the two cues is also where the frequency at which the two ears stop agreeing came from, four rungs ago, and it was computed there for a single source. A hall is fifty sources at once, so the disagreement is not between two accounts of one direction — it is between two accounts of fifty, and there is no reason for the mean of one to equal the mean of the other. Measured at this seat they do agree about the position, because both are odd functions of azimuth and the room is symmetric; at an asymmetric seat they would not have to.
Which computation produced the numbers
The room is the ladder’s own: a 22 by 40 by 15 metre hall with a source on the stage, image sources to the sixth order, an absorption coefficient of 0.18 at every surface, spherical spreading and nothing else.
Each image source gives a direction. The azimuth is measured from the direct sound’s own horizontal direction, because a listener faces the stage — measuring it from the room’s y axis puts the soloist behind the listener and reports 98 per cent of the energy as arriving from the rear, which is a statement about the coordinate system rather than about the hall.
The interaural delay is Woodworth’s r(θ + sin θ)/c with a head radius of 8.75 centimetres, folded at ninety degrees. The level difference is the low-frequency shadow approximation this ladder’s fourth rung uses, folded the same way.
The mean and the width are energy-weighted over every reflection within 120 milliseconds of the direct sound. The lateral energy fraction is the ladder’s own, unchanged: the sideways-weighted energy over the total, with a cosine-squared weight.
The fifteen-microsecond detection threshold is the ladder’s own from its second rung.
Where the model stops
A sphere is not a head. Everything about the front-back fold follows from the head being symmetric front to back, and a real one is not — a pinna is a direction-dependent filter and that is precisely the cue the model is missing.
The listener does not move. Head movement resolves the fold completely and costs nothing, so the confusion computed here is what a listener would suffer if they sat perfectly still, which nobody does.
There is no head-related transfer function. Both cues are analytic approximations. They are right about the size and the shape of each cue and they are not measurements, which the fourth rung says at length.
And the reflections are points. An image source is a specular mirror image, and a real hall’s surfaces scatter — so a real reverberant field is much more diffuse than fifty-four discrete arrivals and its interaural picture is correspondingly smoother.
What the picture cannot show
It cannot show the running correlation. What a listener has is not a list of arrivals with delays attached; it is two continuous signals whose cross-correlation has a shape. The interaural cross-correlation coefficient is the measure that captures that, and computing it needs the signals rather than the echogram.
Nor can it show the precedence effect. The threshold that sorts the reflections operates before any of this: the first wavefront wins, and the later arrivals contribute to a width without ever being heard as separate. This figure draws all of them equally.
It cannot show a moving source. A soloist who walks two metres across a platform changes every image source’s direction and the interaural picture with it, which is a large effect nobody has computed and which opera staging depends on.
It cannot show the periodicity cue. A periodicity in neither ear’s signal is this ladder’s third rung and is a third thing two ears do with a room, and it is not a delay or a level.
It cannot show elevation. Every azimuth here is horizontal, and the image sources from the ceiling and the floor are folded into the horizontal plane by taking their azimuth alone. A spherical head has no elevation cue at all, which is the same missing pinna.
It cannot show the crowd. An echo is prevented by the crowd around it is the ladder’s own account of what an audience does to a rear wall’s return, and an audience is exactly where the folded energy comes from.
And it cannot show two ears listening to music. Everything here is a response to an impulse, and what arrives at a listener is that response convolved with a signal that is itself changing.
Whose halls, and when
The room is a generic large European concert hall of the shoebox kind, and the head is a generic adult one. Nothing here is a measurement of an instrument or a building.
The one historical observation available is about the metric rather than the music. The lateral energy fraction and the interaural cross-correlation coefficient both entered concert-hall acoustics in the 1970s and 1980s as measures of spaciousness, and both were designed to be insensitive to image position — a hall is judged by whether it sounds enveloping, not by whether the orchestra sounds like it is where it is. That is a reasonable thing to have optimised and it is not the only thing a seat decides, and the figure above is the size of what it leaves out.
Where this ladder goes next
Eight rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity in neither ear’s signal; the frequency above which the two ears stop agreeing; a room with directions in it; the threshold that sorts them; what the ones below it do instead of nothing; and now the whole response in the units a listener has, which loses the back of the room.
What is owed after this is the movement. The fold is complete for a stationary head and it is broken instantly by a turn: rotating by ten degrees moves a front source’s interaural delay one way and a rear source’s the other, so a listener who moves has a cue with a sign in it where a listener who does not has nothing. This collection has the head’s curve and the room’s directions and could compute the delay change per degree of rotation for every reflection in a hall — which would say how much of a turn it takes to resolve the ambiguity, and whether the answer is a movement anybody would notice making.
Part 8 of 12
One essay in the series on localisation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Image-sourceInteraural time differenceLateral energy fractionLocalisationPrecedence effectReverberationSpaciousness
- How far away the room takes over localisation, precedence effect, reverberation
- The first eighty milliseconds are a different room precedence effect, reverberation