The turn is half the angle
Assumes: The hall through a head · Two ears, and the whole of the difference is 655 microseconds
The hall through a head ran this ladder’s echogram through the head it had been assuming since its second rung, and lost a third of the room doing it. Both of the cues a spherical head produces fold at ninety degrees: a reflection arriving from 120 degrees makes exactly the interaural delay and exactly the interaural level difference of one arriving from 60, so a listener with those two cues and nothing else cannot tell them apart. At a rear seat that is thirty per cent of the energy.
It ended by naming the way out, and the way out is that a listener moves.
The two hypotheses move opposite ways
Turn the head to the right by an angle and every source moves to the left by that angle in the head’s own coordinates. That is all the geometry there is, and everything below follows from applying it twice.
A source in front at some azimuth has that azimuth as its folded value, and after the turn its folded value is the azimuth minus the turn. Its mirror behind sits at a hundred and eighty degrees less the azimuth; after the turn it sits at a hundred and eighty less the azimuth less the turn — and folding that back gives the azimuth plus the turn.
So one goes down and the other goes up, by the same amount, from the same starting point. That is not an approximation and it does not depend on the size of the turn. It is a property of the fold: the two candidate directions are on the two branches of a curve that has a maximum at ninety degrees, and moving along the axis moves points on opposite branches in opposite directions.
A listener who does not move has a number. A listener who moves has a number and a sign, and the sign is the answer.
That is worth dwelling on, because it explains why the cue is as strong as it is. To decide front from back a listener does not need to know how far the head turned. They do not need a calibrated estimate of the rotation, or a model of the head, or a memory of what the delay was before. They need to know only which way they turned and which way the delay went, and those are two binary facts. Every quantitative claim below is about how big a turn it takes for the change to be detectable; none of it is about how accurately the turn itself has to be known.
And the turn costs exactly half a minimum audible angle
The interaural delay a spherical head makes is Woodworth’s, r(θ + sin θ)/c, and its rate of change with direction is (r/c)(1 + cos θ) per radian. That is 8.9 microseconds per degree straight ahead and 4.5 at ninety — the slope halves across the quadrant rather than going to nothing, which is easy to guess wrong from the shape of the curve alone.
That slope decides two different quantities and they are related in a way that is not obvious until it is written down.
The first is the minimum audible angle: how far a source must move before the same listener notices it has moved. At a threshold of fifteen microseconds it is 1.68 degrees straight ahead and 2.68 at seventy-five.
The second is the turn that resolves the fold. A turn moves the front hypothesis one way and the rear one the other, so the two are separated by twice what either has moved. One threshold of separation therefore arrives after half the turn one threshold of displacement would need.
0.84 degrees straight ahead and 1.34 at seventy-five, at a fifteen-microsecond threshold. At the ten microseconds the ladder’s second rung reports for a trained listener on the median plane, it is 0.56 and 0.89.
This is a small number and it is worth saying how small. It is less than a degree, and this collection’s habit is to price a threshold against what a listener can otherwise do: it is smaller than the precision of an intention: a listener who decides to hold still is already turning by more than that. The confusion the eighth rung computed at length is not resolved by a deliberate movement — it is dissolved by the ordinary business of having a neck. That is why the cone of confusion is a laboratory phenomenon and why it takes a head clamp to demonstrate.
Except where the two directions are the same direction
There is one place the argument fails, and it fails in a way that turns out to be reassuring.
Near ninety degrees the two candidate directions have nearly converged: a source at 89 degrees in front and its mirror at 91 are two degrees apart. Turning cannot spread them further than the difference between them allows, so the separation stops growing with the turn and saturates. At 88 degrees the resolution still costs 1.6 degrees of turn; at 88.5 it costs 7; at 89 it costs 43; and past about 89.2 no turn resolves it at all.
So the ambiguity that survives movement is the one that costs almost nothing to get wrong. A listener who cannot tell 89 degrees from 91 has an error of two degrees in azimuth and a categorical error about front and back, and only the second sounds serious. Everywhere the confusion would place a sound somewhere genuinely different, a turn of under a degree and a half settles it.
What the same turn does to a room
Everything above is about one source. A hall is fifty of them, and they do not behave alike.
At a rear-centre seat in this ladder’s own hall, fifty-three of a hundred and twenty-five arrivals come from behind the listener and carry thirty-three per cent of the energy. Every one of those moves the opposite way from every arrival in front — same geometry, opposite branch — so under rotation a hall is not one object swinging. It is two populations swinging apart.
The consequence is arithmetic. The front population’s energy-weighted rate is 8.2 microseconds a degree and the rear population’s is 8.4 in the other direction, and the response’s own mean moves at 2.7. Almost the whole of it has cancelled.
The direct sound has not. It arrives from straight ahead by construction, it is never behind, and it moves at the full 8.9.
So a head turn separates the source from the reverberation. After twenty degrees the direct sound sits at 176 microseconds and the mean of the whole response at 54, and the gap between them opened at 6.2 microseconds a degree — a fifth of a detection threshold every degree, so about two and a half degrees of turn to make it noticeable.
That is a second cue, and it is one this ladder has not had. A position and a width argued that a room gives a listener a position and a width rather than a direct sound and a diffuse remainder, and the eighth rung converted both into microseconds. Neither of them could say how a listener would separate the two. Rotation separates them, because the two are made of arrivals with different distributions of azimuth and the fold treats those distributions differently. It is also the separation the direct-to-reverberant ratio makes in energy rather than in direction, arriving from a completely different quantity.
And it is strongest exactly where the room is worst
The size of that separation is not the same everywhere in the hall, and its dependence is the most useful thing on this page.
At a mid-stalls seat only two of fifty-four arrivals come from behind and they carry under two per cent of the energy, so nothing cancels: the response’s mean moves at 7.7 microseconds a degree against the direct sound’s 8.9, and the two separate at 1.2. At the rear seat the same difference is 6.2.
The cue that pulls the source out of the room is five times stronger at the back of the hall than in the middle of the stalls, and it is stronger for exactly the reason the back of the hall is a worse seat: because more of what arrives there comes from behind — and the back of a hall is where the reverberant field has taken over altogether.
That is a genuinely odd result and it is worth stating what makes it true rather than leaving it as a coincidence. The rear energy fraction does two things at once. It is what makes the eighth rung’s fold destructive, since a third of the room is being misplaced. And it is what makes the response’s centroid move slowly under rotation, since that third is pulling the other way. One quantity, two consequences, and they point in opposite directions: the seats that lose the most to the fold are the seats where movement recovers the most.
There is a third consequence in the same figure and it is smaller but real. The width of the image at the rear seat goes from 260 microseconds to 297 over a twenty-degree turn, because the two populations are separating. At the mid-stalls seat it goes from 315 to 310 — very slightly narrower, since with almost nothing behind, a turn simply compresses the front arrivals toward the flatter part of the curve. So the sign of the width’s change under rotation is itself a report of how much of the room is behind the listener, which is a quantity nothing in a signal otherwise carries and which no measure of how long a room rings can supply.
What the fold does to the picture the eighth rung drew
Read the arrivals on that figure against the rates on the one three sections above and the relationship is exact. Every dark point moves right as the head turns right; every pale point moves left; the direct sound is the pale point at the top, moving fastest of all. Nothing about the picture at rest says which dots are which, and one degree of turn says it about all of them at once.
It is worth being precise about what “says it” means, because a listener is not reading a scatter plot. What arrives at two ears is two continuous signals, and their running cross-correlation has a shape with a peak in it. A turn moves that peak, and it moves the parts of the distribution that are behind the listener the other way — so what a real listener has is a change in the shape of the correlation and not a list of arrivals with signs attached. The claim this ladder can support is that the information is present and how much of it there is; the claim it cannot support is that a listener uses it in this form.
The frequency the whole argument lives at
Two caveats about frequency, and the first is the ladder’s own.
An interaural delay stops naming a direction above the frequency where half a wavelength is shorter than the largest delay the head can make, which for this head is 762 hertz. Above that the two ears stop agreeing about the fine structure, and the level difference takes over.
The level difference folds too, and it has a sign under rotation for the same reason. But its rate of change with azimuth is much smaller near the median plane than the delay’s — the shadow depends on which side of the head the source is, and near straight ahead it is on neither — so the turn that resolves the fold by level is considerably larger than the turn that resolves it by delay. This ladder has not computed how much larger, because the level cue’s own threshold is a different quantity from the delay’s and this collection has only the delay’s.
The second caveat is about the room. A reflection’s energy is spread over frequency and the arrival times drawn here are broadband. A real reverberant field is diffuse rather than fifty discrete arrivals, and what a turn does to a diffuse field is to change its interaural coherence rather than to move any arrival — a quantity this collection computes elsewhere and has not put under rotation.
Which computation produced the numbers
The head is Woodworth’s sphere at 8.75 centimetres and 343 metres a second, unchanged from this ladder’s second rung. The room is its own: a 22 by 40 by 15 metre shoebox, a source on the platform, image sources to the sixth order, eighteen per cent absorption at every surface and spherical spreading.
The turn threshold is found by scanning rather than solved, because past about eighty-eight degrees the separation has a maximum in it and a bisection would report the first crossing of a curve that later comes back down. The minimum audible angle is the analytic slope, so the factor of two is a comparison of two independently computed quantities rather than an algebraic identity restated.
The detection threshold is fifteen microseconds, which is what the eighth rung used for a comparison between two hall conditions. Ten is the figure for a trained listener discriminating on the median plane, which is the smallest difference this collection credits anywhere; every number here scales linearly with whichever is chosen, and the factor of two between the turn and the angle does not depend on it at all.
Where the model stops
Nothing here knows how far the head turned. The cue as computed is a sign, and a sign needs no calibration — but a listener who wants to know where a source is, rather than merely whether it is in front, has to combine the change in delay with an estimate of the rotation, and the precision of that estimate is a proprioceptive quantity this collection does not have.
The turn is instantaneous. A real turn takes time, during which the source may move, stop or change, and the binaural system’s own integration window is tens of milliseconds. Nothing on this page has a duration in it.
A sphere is still not a head. The pinna breaks the front-back symmetry by a completely different route, imposing a direction-dependent notch above about six kilohertz, and that cue is available to a listener sitting perfectly still. The two are not alternatives; a real listener has both.
And the sources are points. Every reflection here is a specular mirror image with a single direction. A real surface scatters, so a real hall’s rear arrivals are spread rather than discrete, and their rates under rotation are spread with them.
What the picture cannot show
It cannot show a moving source. A soloist walking across a platform changes every image source’s direction, and a listener turning while the source moves is receiving the sum of two rotations they have no separate access to.
Nor can it show the precedence effect. The threshold that sorts the arrivals operates before any of this, and the first wavefront wins whatever the later ones do. Every arrival on these figures is drawn as though its direction were equally available, and it is not.
It cannot show the effort. Turning to resolve an ambiguity is something a listener does, and whether they do it depends on whether the ambiguity matters to them. Nothing in a geometry says when a listener bothers.
And it cannot show a hall’s own asymmetry. The room here is a symmetric shoebox and the seats on its centre line are symmetric in it, so several of the numbers above are zero by construction rather than by measurement. A hall with a balcony over the listener is a different arithmetic and the same method.
Whose halls, and when
The room is a generic large European shoebox and the head is a generic adult one; nothing here is a measurement of a building or a person.
The one observation that is not the model’s own is about practice rather than about the model. Concert-hall acoustics measures seats with an omnidirectional microphone and a figure-of-eight beside it, both bolted to a stand. Every quantity a hall is judged on is therefore a quantity measured by an instrument that cannot turn, and the cue on this page is unavailable to it in principle. That is not a criticism — a repeatable measurement has to hold still — but it does mean that the difference between a seat where movement recovers a great deal and one where it recovers very little is not in any hall’s published numbers.
Where this ladder goes next
Nine rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity that is in neither ear’s signal; the frequency above which the two ears stop agreeing; a room with directions in it; the threshold that sorts them; what the ones below it do instead of nothing; the whole response in the units a listener has, which loses the back of the room; and now the movement that gets it back, at a cost of under a degree and a half.
What is owed after this is the duration. Every number on this page is a change in a delay, and a change is not a quantity until something has had time to measure it twice — a turn takes tens or hundreds of milliseconds, the binaural system integrates over a window of its own, and a sound that stops before the turn is finished cannot be disambiguated at all however large the sign. This collection has the integration window from its masking ladder and the settling times from its instrument ladders, and putting them against a plausible head-turn rate would say how long a sound has to last for movement to help, and therefore which musical events are disambiguated by it and which are not. A held chord and a pizzicato would come out on opposite sides of that line, and the arithmetic is a division.
Part 9 of 12
One essay in the series on localisation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Head shadowImage-sourceInteraural time differenceJust-noticeable differenceLocalisationPrecedence effectReverberation