Two ears, and the whole of the difference is 655 microseconds
A listener asked where a sound is will point, confidently, and be right to within a few degrees. The information available for that judgement is startlingly thin: two pressure signals, from two points about seventeen centimetres apart, and nothing else.
The delay is computed from geometry alone. For a sphere of radius , the extra path to the far ear is — a straight-line segment plus an arc around the surface — so the interaural time difference is
with the speed of sound and the azimuth. Nothing in that is fitted. For a head of radius 8.75 cm the maximum, at ninety degrees, is 655 microseconds.
Six hundred microseconds, and ten of them matter
The whole usable range of the first cue is 655 microseconds. A listener resolves changes of about ten microseconds in it, which is a resolution of roughly one part in sixty-five of the total.
That number deserves to be stared at. Ten microseconds is a hundred-thousandth of a second. The period of a 440 Hz tone is 2,273 microseconds, so the ear is resolving a phase difference of about a six-hundredth of a cycle of the note it is listening to. No neuron fires that fast, no membrane responds that fast, and the resolution is achieved by comparing the timing of firings from the two sides in a structure — the medial superior olive — whose job appears to be nothing else.
That the front is the accurate region is not a coincidence either. It is where a listener puts things they are attending to, by turning towards them, and turning towards a source is a way of moving it into the steep part of a curve.
Why one cue is not enough
The delay cue has a hard ceiling, and the ceiling is arithmetic.
Interaural delay is detected from the phase difference between the two ears. A phase difference is unambiguous only while it is less than half a cycle: beyond that, a delay of Δ and a delay of Δ minus one period produce the same phase reading, and the two answers are on opposite sides of the head.
Half a period equals the maximum path difference when
so above roughly 760 Hz the phase cue starts to be ambiguous for sources at the extremes, and by about 1,500 Hz it is useless for everything. The number comes from the speed of sound and the width of a head and from nothing else. A larger head has a lower crossover; an elephant’s phase cue fails an octave earlier than a human’s, and a mouse’s has almost no low-frequency range to work in at all.
This is Lord Rayleigh’s duplex theory, from 1907, and it is one of the older results in the subject that has survived unchanged. Low frequencies are localised by time, high frequencies by level, and the region between 750 and 1,500 Hz is the worst part of the spectrum for locating anything — which is a testable prediction and it holds: localisation accuracy for pure tones has a measurable dip there.
What sits in that dip
The frequency at which localisation is worst is, unhelpfully, the frequency at which most music sits. The middle of the piano is 262 to 1,046 Hz; the human speaking voice has its fundamental well below the dip and its formants well above.
That is less of a problem than it sounds, and the reason is worth stating because it is a case of a limitation being circumvented rather than tolerated. A real musical note is not a pure tone — it is a harmonic series spanning several octaves — so a note whose fundamental is at 400 Hz has partials at 800, 1,200, 1,600 and beyond. The fundamental is localised by time, the upper partials by level, and the answers agree because they came from the same place. Complex tones are localised far more accurately than pure tones anywhere in the range, and particularly in the dip.
There is a second mechanism doing the same job. The envelope of a high-frequency sound carries timing information even when its carrier does not: a burst of 4 kHz tone arriving at one ear before the other has an unambiguous onset difference, even though its individual cycles are hopelessly ambiguous. The ear uses envelope timing at high frequencies, which is one reason a click is easier to localise than any steady tone, and why percussion is the most precisely placed thing in an ensemble.
That per-band picture also explains an everyday failure. A subwoofer can be put anywhere in a room and is conventionally described as being non-directional. It is not that low frequencies carry no direction information — they carry the best kind, an unambiguous phase difference — it is that a single low band in a small room is dominated by the room’s own modes, which arrive from the walls rather than from the loudspeaker. The cue is intact and the evidence is corrupted.
Both of those numbers are about a source in free space. In a room the same geometry produces a third one.
So the head’s width sets where a room stops sending the two ears the same sound, by the same arithmetic that sets the phase ambiguity — one separation, one wavelength, three different consequences. The two ears do not stop agreeing at one frequency; they stop agreeing at whichever of these the situation makes relevant.
What the two numbers cannot say
Interaural time and level differences fix one thing: the angle between the source and the plane of symmetry through the head. That is a cone, not a direction.
Every point on a cone opening out from one ear produces the same time difference and the same level difference — directly in front, directly behind, above, and below all lie on the same cone when the source is straight ahead. This is the cone of confusion, it is a hard geometric consequence of having two receivers, and it means front and back are indistinguishable and elevation is unavailable from these cues at all.
Two things resolve it, and both are outside the model above.
Moving the head. Turning slightly changes the time difference in one direction for a source in front and in the opposite direction for a source behind. A listener who has moved their head has resolved the ambiguity, and listeners do it constantly and without noticing. Anyone who has tried to locate a chirping smoke alarm has experienced the failure of every cue in this essay and then solved it by walking around.
The shape of the ear. The pinna is a small, complicated, asymmetric reflector, and its folds impose direction-dependent notches on the spectrum above about 5 kHz — a notch near 8 kHz for a source in front, moving in frequency as the source moves up. This is a monaural spectral cue: one ear alone can use it, it is the only source of elevation information, and it is idiosyncratic to the individual ear, learned during development.
That last fact is why recorded binaural audio is unconvincing for most listeners. A recording made with a dummy head carries somebody else’s pinna cues, and a listener who has spent their life learning the notches of their own ears is being given a stranger’s. The time and level differences transfer perfectly; the spectral cues do not, and the result is a sound image that is convincingly to one side and stubbornly inside the head.
Where an orchestra is, and whether it matters
The claims above are about a listener and a source. Applying them to a concert hall produces a result that is slightly deflating and worth having.
At thirty metres from the back of a hall, an orchestra eighteen metres wide subtends 33.4 degrees, so the first violins sit at +16.7 and the double basses at −16.7. Running those through Woodworth gives a delay difference between the two ends of 295 microseconds — 45 per cent of the 656 the geometry allows at all, a large signal and easily resolved. The seating plan is audible.
Anything finer is harder, and the reason is not the one it looks like. Two players two metres apart at that distance differ by 3.8 degrees and by 34 microseconds of delay, which is not at the edge of the resolution limit — it is more than three times it. Solving the geometry the other way round: in a free field, a listener at the back of that hall has enough delay resolution to separate two sources sixty centimetres apart on the platform.
So the geometry is not what defeats the listener. What defeats them is everything the free-field figure leaves out: eighty sources sounding at once, so that each ear receives a mixture whose delay is the delay of nothing in particular, and a reverberant field arriving from every direction with delays that are simply wrong.
That is a more useful conclusion than a resolution limit would have been, because it says where the loss happens. The ears are good enough to resolve the desks and the situation is not, and the familiar experience of hearing exactly where the oboe is comes mostly from having seen it. The accuracy of a listener’s spatial map of an ensemble is much lower with the eyes closed, and it is lower for reasons that would not improve with better ears.
Whose practice. Orchestral seating conventions are frequently justified in spatial terms — antiphonal violins, cellos to the outside, and so on — and this essay is not evidence for or against any of them. What it does say is that the differences those arrangements produce at a listener’s ears are at the scale of tens of degrees rather than of individual desks, and that the audible consequence of a seating change is more likely to be about which instruments are masking which than about where they appear to be.
A beat that is in neither ear
The clearest evidence that the two ears are compared rather than merely added is a sound that exists in neither of them.
Present 440 Hz to one ear and 444 Hz to the other, in isolation — headphones, no crosstalk — and a slow four-per-second throb is heard. Nothing in either ear is throbbing. The air at the left eardrum is a pure 440 Hz tone and the air at the right is a pure 444 Hz tone, and neither signal contains any amplitude variation at all.
What varies at four hertz is the interaural phase: the two signals drift in and out of step once every quarter of a second, so the apparent position of the sound rotates slowly inside the head. The binaural beat is heard as movement rather than as loudness variation, which is the giveaway.
It is also much weaker than an ordinary beat, and it disappears above about 1,000 Hz — the same ceiling as the phase cue itself, and for the same reason. That coincidence is the argument: binaural beats and interaural delay fail at the same frequency because they are the same machinery.
A caution, since this is a topic with a literature attached. Binaural beats are the subject of a large body of claims about brainwave entrainment, altered states and therapeutic effect. This essay makes none of them. What is being asserted here is exactly one thing — that a variation exists in the percept which exists in neither signal — and that is a claim about where two signals are compared, not about what the comparison does to anybody.
Which computation produced the numbers, and which did not
Two of the figures here are derivations and one is an approximation, and the difference should be visible.
The delay is exact, in the sense that it follows from a stated geometry with no fitted constants. The assumption is a rigid sphere with the ears at opposite poles, which is Woodworth’s 1938 idealisation and is good to within a few per cent for a real head.
The crossover frequency is exact for the same reason. It is one over twice the maximum delay, and it moves when the head size does.
The level difference is an approximation and is drawn as one. A real head is not a sphere, the ears are not at the poles, and the shadow depends on frequency in a way that a simple diffraction expression only gets roughly right. What the figure is entitled to claim is where the cue appears and its order of magnitude — nothing to speak of below 500 Hz, and up to about twenty decibels at the top. A measured head-related transfer function is what a real system uses, and it is a table with thousands of entries per ear, individual to the person it was measured on.
What the picture cannot show
A room. Every figure here assumes one source, in free space, with no reflections. Real listening is in a room, and a room’s reflections arrive from every direction with their own time and level differences, most of which are wrong. That the system works at all in a reverberant space is a much harder result than anything in this essay, and it needs a separate mechanism that discards the reflections rather than averaging over them.
Loudness. Nothing here says how the two ears’ signals are combined into one impression of level, and they are: a sound presented to both ears is louder than the same sound presented to one, by about six decibels, which is a fact about summation rather than about direction.
Motion. A moving source produces a changing delay, and the change is a strong cue in its own right — strong enough that a sound which is moving is localised better than the same sound standing still. None of the figures has a time axis at all.
Distance. Nothing above says how far away a source is, and the two cues are nearly silent on it: direction is a good deal easier than distance, which relies on loudness, on the ratio of direct to reverberant sound, and on the loss of high frequencies over long paths in air. Distance judgements are correspondingly poor and are systematically compressed — far things are judged nearer than they are.
Two sources at once. The whole apparatus is built to answer where a sound is. With two sources sounding together, both ears receive a mixture, and the delay of a mixture is not the delay of either component. Untangling that requires having already decided which components belong to which source, which is the streaming problem — and localisation, which looks like it ought to come first, in fact depends on it.
Where the ladder goes next
The single largest problem this essay has left open is the room, and the next rung is the mechanism that solves it. Between about one and thirty-five milliseconds, a second copy of a sound is not heard as a second sound, does not move the image, and does not get a vote on the direction: the first wavefront wins, and the numbers that bound the effect are what decide whether a hall sounds clear or confused.
Sideways in a different direction, the machinery here is what makes an ensemble’s parts separable at all: a source’s position is one of the cues that decides which components belong together, and two instruments in the same register segregate more readily when they are in different places. Localisation looks like a question about geometry and is also a question about grouping.
There is also a rung below this one that this site has not written and should note the absence of. Everything here assumes two working ears with matched sensitivity. Asymmetric hearing loss removes the level cue on one side and leaves the timing cue intact but miscalibrated, and the result is not a halving of spatial hearing but a systematic bias — sources are pulled towards the better ear. That is a large fact about a large number of listeners and it is a rung’s worth of argument on its own.
Sideways from here sits a use of the same machinery that has nothing to do with direction. Giving a masked tone a different interaural delay from its masker makes it audible again, ten to fifteen decibels below the masked threshold — a release from masking bought entirely with spatial information. The two ears are not only a direction finder; they are also, and perhaps more importantly, a way of hearing one thing in the presence of another.
Part 1 of 12
One essay in the series on localisation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 16.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Cone of confusionDuplex theoryHead shadowInteraural level differenceInteraural time differenceLocalisationPinna
- A room with directions in it interaural time difference, localisation