Perception and the listener

Where a wrong head gives itself away

Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.

Assumes: A smaller head in the same hall · The turn is half the angle

Eleven rungs of this ladder rest on one equation. Woodworth’s delay for a sphere is (r/c)(θ+sinθ)(r/c)(\theta + \sin\theta), every figure on the anchor evaluates it at r=8.75r = 8.75 centimetres, and the rung that finally varied that radius showed what happens when the head is a different size: the whole delay range scales with it, the detection threshold does not, and a smaller listener has fewer distinguishable positions in front of them.

That essay ended by naming what it had not touched. A listener does not know rr. Nothing tells an infant their head radius; the map from delay to direction is acquired, and it is acquired while the head is growing by seventy per cent underneath it. So the interesting question is not whether the map can be learned but what happens when it is wrong — and the closing paragraph made a prediction about that:

the error is largest exactly where the curve is steepest, which is the median plane, where the direct sound is

The arithmetic says the opposite, and says it in one line.

A head model a few millimetres out reads every azimuth but the front. The azimuth a listener reports against the azimuth a source is at, for internal head radii from 8.22 to 9.28 centimetres against a true radius of 8.75. The delay a source produces is (r/c)(θ + sin θ) and the listener inverts it with the radius they believe they have, so their answer solves θ̂ + sin θ̂ = (r/r̂)(θ + sin θ). Every curve passes exactly through the origin: on the median plane there is no delay and therefore no error, whatever the head model is. The error grows with azimuth and is largest at the side. An internal head 5.3 millimetres too small runs out of azimuth at 82 degrees: beyond that the world is delivering a delay larger than any its owner's model can produce, and every source out there collapses onto the side.
Fig. 1 Where a listener puts a source against where it is, for four internal head radii either side of a true 8.75 centimetres. Every curve passes through the origin exactly. The two low ones run out of azimuth before ninety degrees.

One substitution, and the one place it does nothing

Give the listener a head of radius rr and an internal model r^\hat r. The world hands them a delay D=(r/c)(θ+sinθ)D = (r/c)(\theta + \sin\theta) and they invert it with the radius they believe they have, so the direction they report solves

θ^+sinθ^  =  rr^(θ+sinθ).\hat\theta + \sin\hat\theta \;=\; \frac{r}{\hat r}\,(\theta + \sin\theta).

That is the whole model. It has no free parameter in it and nothing was fitted; r/r^r/\hat r is a single number and the rest is the same function the first rung drew.

Read it at θ=0\theta = 0. The right-hand side is zero whatever r/r^r/\hat r is, so θ^\hat\theta is zero too. A source straight ahead is straight ahead for every listener, however wrong their head model is, and not approximately — exactly, because a source on the median plane produces no interaural delay at all and there is nothing there to misread. That is the same zero the rung about a periodicity in neither ear’s signal works from, read for once as an absence of evidence rather than as a starting point.

The error grows from there. A head model two per cent too small — 1.75 millimetres on an adult — reports 30 degrees as 30.6 and 60 as 61.5. At the far end it does something stranger. The largest delay a head of radius r^\hat r can account for is (r^/c)(π/2+1)(\hat r/c)(\pi/2 + 1), and past some azimuth the world is delivering more than that: the listener is handed a delay their own model says is impossible. The 1.75-millimetre error runs out at 87 degrees and the 5.25-millimetre one the figure above draws at 82. Every source beyond that collapses onto the side, because there is nowhere else in the map to put it.

The steepness is in both quantities, and cancels between them

The prediction the last rung made is a reasonable one and it is worth being precise about why it fails.

The delay curve’s slope is (r/c)(1+cosθ)(r/c)(1 + \cos\theta) per radian, which is twice as steep straight ahead as it is at the side. That has two consequences and they point opposite ways. A steep curve turns a given error in the delay into a small angular error, so the mis-calibration displaces a frontal source less. And a steep curve makes a given angular displacement produce a large change in the delay, so the smallest detectable displacement — the quantity the turn rung computed as jndc/(r(1+cosθ))\text{jnd}\cdot c / (r(1+\cos\theta)) — is smaller at the front too.

Both carry the same factor. Write ε=r/r^1\varepsilon = r/\hat r - 1; then the angular error is ε(θ+sinθ)/(1+cosθ)\varepsilon (\theta + \sin\theta)/(1 + \cos\theta) and the smallest detectable angle is jndc/(r(1+cosθ))\text{jnd}\cdot c/(r(1+\cos\theta)), and their ratio is

εITD(θ)jnd\frac{\varepsilon \cdot \mathrm{ITD}(\theta)}{\text{jnd}}

with no trace of the slope left in it.

The steepness is in both quantities and cancels between them. Two angles against azimuth, for an internal head 1.75 millimetres too small. The rising curve is how far the mis-calibration displaces the source; the flatter one is the smallest displacement a listener could detect at that azimuth, which is 10 microseconds divided by the delay's own slope. Both carry a factor of one plus the cosine of the azimuth — the first because a steep curve turns a delay error into a small angle, the second because a steep curve makes a small angle detectable — so the ratio between them is the delay error over the threshold and nothing else. They cross at 60 degrees: the mis-calibration is invisible in front and visible at the side, which is the opposite of what the steepness of the delay curve suggests.
Fig. 2 The two angles for a head model 1.75 millimetres too small: how far the source is displaced, and the smallest displacement anybody could see. Both contain one plus the cosine of the azimuth and the ratio between them does not.

So the question “where is a mis-calibration visible” is not a question about slopes at all. It is the question “where is the delay large”, and the delay is largest at the side. The median plane is where a listener localises best and it is the one direction that carries no information about whether their map is right.

That is worth stating as a general shape, because it is not special to this model. A multiplicative error in a monotone map is invisible wherever the map is near zero, and every acuity measure that is a threshold on the mapped quantity inherits the same cancellation. The interaural delay happens to be zero on the axis a listener faces, which is a coincidence of anatomy, and it puts the blind spot of the calibration problem exactly where the attention is.

One and three tenths of a millimetre

Setting the ratio to one gives the smallest detectable mis-calibration at each azimuth, and it is a number rather than a shape.

A head model is wrong by 1.3 millimetres before anything gives it away. How far out an internal head radius has to be before the source it misplaces is misplaced detectably, against where the source is. At a threshold of 10 microseconds it is 16.1 millimetres for a source 5 degrees off centre and 1.31 at the side; At a threshold of 15 microseconds it is 22.1 millimetres for a source 5 degrees off centre and 1.96 at the side. The curve is jnd divided by the delay, so it falls the whole way and its minimum is at ninety degrees for every threshold. A newborn's head grows 30 millimetres on its way to an adult's, which is 23 of these steps: the map has to be relearned that many times over.
Fig. 3 How far out an internal head radius has to be before the source it misplaces is misplaced detectably, against where that source is. At the two thresholds used throughout — ten microseconds for the trained case and fifteen for a room — the best available anywhere is 1.31 and 1.96 millimetres, and both are at ninety degrees.

At the ten-microsecond threshold the first rung quotes, a head model 1.31 millimetres out is caught by a source at the side and nothing smaller is caught anywhere. At the fifteen-microsecond figure the rungs working in a room use, it is 1.96. Ten degrees off centre the same computation asks for 8.85 millimetres, and five degrees off centre for very much more.

Those two numbers are the substance of this rung, and the way to read them is as a tolerance. A listener’s internal head can be a millimetre and a third wrong and no source in the world will say so. On an adult head that is one and a half per cent; on a six-year-old’s it is nearly two. Whatever process acquires this map, it is not being asked for better precision than that, because better precision would have nothing to be checked against.

An impossible delay is its own alarm

There is one exception buried in the previous section and it is not really an exception.

The saturation above is the detection criterion at ninety degrees, arrived at from the other direction. When the internal head is too small by Δr\Delta r, the largest delay it can account for falls short of the largest delay the world produces by (Δr/c)(π/2+1)(\Delta r/c)(\pi/2 + 1), and at Δr=1.31\Delta r = 1.31 millimetres that shortfall is 9.85 microseconds — which is the threshold, to the width of the line. A mis-calibration large enough to be caught by a source at the side is exactly a mis-calibration large enough to make a source at the side impossible.

That matters because it removes the need for a teacher. Everything in the section above was scored against a listener who somehow knows where the source really is. A delay outside the listener’s own model’s range needs no such thing: the map itself refuses the input.

The same geometry at four sizes of head. Woodworth's interaural delay against direction, for 4 head radii from 5.8 to 9.8 centimetres. The whole range runs from 431 microseconds for a newborn to 735 for a large adult, and it scales exactly with the radius because the delay is (r/c)(θ + sin θ) and r is a multiplier. The detection threshold does not scale with the listener, so the number of distinguishable delays across the whole range falls from 147 to 86: a smaller head has the same directions in front of it and a shorter ruler to measure them with.
Fig. 4 The four head sizes used throughout, and the delay curve each produces. The vertical extent of each curve is the range its owner’s map has to cover, and it runs from 431 microseconds to 735.

It is a narrow instrument, though. It fires only when the internal head is too small, since a model too large simply has range to spare; and it fires only for sources within a few degrees of the interaural axis, where a listener’s angular resolution is at its worst and where — as the seat-by-seat rung found — a concert hall puts very little of anything. A hall’s reflections do arrive from the sides, and the rung that gave a room its directions counted them; but a reflection is not a source of known position, and the one that measured what the response looks like as an image found it spread over tens of microseconds rather than standing at one.

The check that needs nobody else

There is a second self-supervised signal, it is stronger, and it is available in the dark.

A listener turns their head. That is the manoeuvre the ninth rung priced for a different purpose — resolving the front-back fold — and it carries information this rung needs as a by-product. Proprioception knows the angle turned to a fraction of a degree; the delay before the turn and the delay after it are both measured; and the map has to be consistent with all three. Place a source at θ^1\hat\theta_1, turn by φ\varphi, and the delay afterwards should be the delay of a source at θ^1φ\hat\theta_1 - \varphi read through the same internal head. If the head is wrong the prediction misses, and the size of the miss is in microseconds and comparable directly to the detection threshold.

The listener can catch it without being told anything. What a head turn costs a wrong head model, for an internal radius 1.75 millimetres too small. The listener places a source, turns by a known angle — proprioception knows the angle exactly — and predicts the delay the turn should produce. The curves are how far that prediction misses, in microseconds, against how far the head turned. Nothing outside the listener is consulted: no seen source, no second opinion, only the requirement that the map agree with itself. The shaded band is one detection threshold of 10 microseconds. The signal is largest for a source at the side and a turn that brings it round to the front, and reaches 23 microseconds at 85 degrees with a turn of 80. Solving for the error that just clears the threshold gives 0.76 millimetres — better than a seen source manages, and available in the dark.
Fig. 5 What a head turn costs a wrong head model — how far the predicted delay after the turn misses the delay that arrives, for an internal radius 1.75 millimetres too small. The shaded band is one detection threshold. No source position is known and nothing outside the listener is consulted.

Solving for the error that just clears ten microseconds under a turn of at most eighty degrees gives 0.74 millimetres — nearly twice as sensitive as the best a seen source can manage, and it needs no source of known position, no vision and no second opinion. The map is only required to agree with itself.

Two things about it are worth carrying. The signal is largest for a source near the interaural axis brought round toward the front, which is the same geography as before; and it wants a large turn. Restricted to forty-five degrees the same computation gives 1.22 millimetres, which is no better than the seen source. The self-check is strong precisely because a big turn traverses the part of the curve where the two head models disagree most, and a listener glancing round is not running it.

The other cue knows, and is four times too quiet to say so

The duplex rung established that a head offers two cues and that each fails where the other works. Both read the same radius, and they read it differently — the delay through θ+sinθ\theta + \sin\theta, the level difference through the sphere’s shadow, which depends on the head in wavelengths. So a wrong radius should put the same source in two places at once, and a conflict between two cues is the classic shape of a recalibration signal.

It is there. It is not usable.

The second cue knows too, and it is 6 times too quiet to say so. The conflict between the two binaural cues, in decibels, for an internal head 1.75 millimetres too small. The listener places the source by its delay, reads that azimuth back through the head they believe they have to predict the level difference it should produce, and compares the prediction against what arrives. The shaded band is the published interaural level difference limen, about one decibel. The largest conflict anywhere drawn is 0.174 decibels at 90 degrees and 6 kilohertz — 6 times below the threshold. Reaching one decibel would take an internal head 10.0 millimetres out at 3 kilohertz, which is seven times the error the delay cue catches on its own. The conflict is real, it is present at every azimuth, and it is not a signal a listener could use.
Fig. 6 The conflict between the two cues in decibels, for an internal head 1.75 millimetres too small: the level difference the listener’s own model predicts for the direction the delay gave them, against the level difference that arrives. The shaded band is the published limen for an interaural level difference.

Scoring that conflict needs care about which domain it is scored in, and the obvious domain is the wrong one. Asking each cue for an azimuth and differencing the two answers puts the level cue through an arcsine, and an arcsine near ninety degrees turns a hundredth of a decibel into whole degrees — so a comparison drawn that way reports a conditioning problem as a finding. The level cue’s own threshold is in decibels, and that is where the conflict belongs: the listener places the source by its delay, reads that azimuth back through the head they believe they have, predicts the level difference it should produce, and compares.

The largest conflict anywhere on the figure is 0.13 decibels, against a limen of about one — six to eight times too small over the whole range. Reaching one decibel would take an internal head ten millimetres out, which is seven times the error the delay cue catches on its own.

The reason is that the shadow depends on the head only through log(1+k2a2)\log(1 + k^2a^2), which for a head several wavelengths across is nearly flat: a two per cent change in the radius moves the shadow at three kilohertz from 13.82 decibels to 13.66. The delay is proportional to the radius outright. The two cues are unequally sensitive to the quantity that is wrong, so the level cue’s reading is very nearly right even when the delay’s is not, and the difference between them is dominated by the delay’s error rather than being an independent check on it.

This is a negative result about a mechanism that would have been an attractive answer, and it is priced rather than merely denied: the level cue would have to depend on head size about eight times more steeply than a sphere makes it for the conflict to be usable.

A map whose target will not hold still

Put the tolerance beside the growth and the shape of the learning problem falls out.

A growing head goes out of calibration fifteen times over. The four head sizes used throughout, placed by how many detectable steps of mis-calibration each is from a newborn's. A step is 1.31 millimetres at a threshold of 10 microseconds, A step is 1.96 millimetres at a threshold of 15 microseconds. A newborn's 57.5-millimetre radius reaches an adult's 87.5 by growing 30.0 millimetres, 23 steps at 10 µs and 15 steps at 15 µs. The growth is front-loaded and this is not a growth chart: taken evenly over the six years to a six-year-old's head it is one step every 11 months, and the true early rate is faster than that.
Fig. 7 The four head sizes again, placed by how many detectable steps of mis-calibration each is from a newborn’s. A step is 1.31 millimetres at the ten-microsecond threshold and 1.96 at fifteen.

A newborn’s interaural radius of 57.5 millimetres reaches an adult’s 87.5 by growing thirty millimetres, which is twenty-three detectable steps at the trained threshold and fifteen at the room threshold. The map is not learned once. It is learned, invalidated, and relearned, on the order of twenty times, and every one of those relearnings has to happen from evidence available only at the sides of the listener.

The rate matters as much as the count and it is where this stops being arithmetic. Head growth is heavily front-loaded — the published growth charts have most of the infant year’s increase in its first months — so a rate taken evenly across the six years to a six-year-old’s 70-millimetre radius is a floor rather than an estimate. Evenly, twelve and a half millimetres over six years is one detectable step every eleven months at the room threshold and every seven and a half at the trained one. In the months where growth is fastest it is a step in weeks.

This is not a claim about musical practice and should not be dressed as one. It is a claim about a population — human infants, whose head circumference is one of the most heavily tabulated quantities in medicine — and the tabulation is what makes the count checkable by somebody who has the charts. What it does bear on musically is small and real: a child in a hall is not a small adult in a hall, and the rung before this one already showed the hall arriving on a shorter axis. This one adds that the axis is also, for part of childhood, out of date.

Which computation produced the numbers

The delay is Woodworth’s spherical-head formula throughout, at 343 metres a second, with the true radius at 8.75 centimetres and the internal radius stated per figure. The inversion is by bisection on θ+sinθ\theta + \sin\theta, which is monotone on the quarter turn, and it returns nothing rather than a clamped value when the argument exceeds π/2+1\pi/2 + 1 — the saturation is a real absence of a solution and not a drawing convention.

The smallest detectable angle is jndc/(r(1+cosθ))\text{jnd}\cdot c/(r(1+\cos\theta)), the analytic form the ninth rung derived rather than a search. Two thresholds are carried throughout because this collection uses two and they are for different tasks: ten microseconds is the trained figure for a source straight ahead and fifteen is what the rungs working in a room use, where the comparison is between two seats.

The turn figure predicts the delay after a turn as (r^/c)f(θ^1φ)(\hat r/c)f(\hat\theta_1 - \varphi) and compares it against the delay that arrives, (r/c)f(θφ)(r/c)f(\theta - \varphi), with the fold at ninety degrees handled by sign rather than by clamping. The detectable error under that criterion is found by bisection on the internal radius, with the maximum taken over a grid of source azimuths and turn angles.

The level cue is the same low-frequency shadow approximation the duplex rung uses, 10log10(1+k2a2)10\log_{10}(1 + k^2a^2) times the sine of the azimuth, and the conflict between the two cues is computed in decibels rather than in degrees for the reason that section gives. The threshold it is scored against is the published interaural level difference limen, taken at one decibel; the literature puts it between about a half and one and a half depending on frequency and on the listener, and the conclusion is unchanged anywhere in that range because the conflict is under a seventh of the smallest of them. Each figure asserts that its own quantity vanishes when the internal head is right, which is the check that matters: every number here is a difference between two readings of one geometry, and a difference that failed to vanish at r^=r\hat r = r would be an artefact of the arithmetic rather than a property of the ear.

Where the model stops, and what the picture cannot show

The internal head is one number here and a listener’s map is not. Nothing requires an acquired delay-to-direction map to be parameterised by a radius at all; it could be a lookup with no shape assumption in it, in which case a “mis-calibration” is not a scalar and the arithmetic above describes the wrong kind of error. What the one-parameter version buys is that it is the smallest family containing the truth, so the tolerance it produces is the tolerance for the most favourable case — a listener who has to learn only one number and has got it slightly wrong. A listener learning a whole surface has more to get wrong, not less.

A sphere is not a head, and the elevation is missing entirely. Woodworth’s formula has no pinna in it, and everything above is confined to the horizontal plane, where the delay is the whole story. The median plane the last rung named is not merely the place where this analysis finds no error — it is the place where the delay carries no information at all and directions are told apart by spectral cues the model does not contain. The claim here is narrow and should be read narrowly: the delay-based map has nothing to say about its own calibration in front.

The threshold is treated as fixed and it is not measured on children. The whole result turns on the detection threshold not scaling with the head, which is the previous rung’s argument and rests on the threshold being a property of the binaural comparison rather than of the skull. Interaural thresholds in infants are known to be considerably worse than adult ones — the same direction the duration effect runs in when a note is too short to specify its own frequency — and a worse threshold makes the tolerance wider, not narrower — so the twenty-three steps is an upper bound on how often the map is invalidated, and the true count over a childhood is smaller.

And there is no learning rule anywhere in this. The figures say what evidence exists and how large it is. They say nothing about what a listener does with it, how many exposures a recalibration takes, or whether it happens at all — a listener could simply live with a map a millimetre out, and the tolerance computed here says they would never find out.

Where this ladder goes next

Eleven rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity in neither ear’s signal; the frequency above which the two ears stop agreeing; a room with directions in it; the threshold that sorts them; what the ones below it do instead of nothing; the whole response in the units a listener has; the turn that recovers the third of it the head throws away; the listener who is not one size; and now the listener who does not know what size they are, whose map can be 1.3 millimetres wrong before anything says so and who can only find out by looking sideways.

What is owed after this is the asymmetry, and it is the parameter this anchor still sets to one value without saying so. Every head in every figure on this ladder has its two ears at ±r\pm r from a centre, so the delay curve is odd and a source at +θ+\theta produces exactly the negative of a source at θ-\theta. Real heads are not symmetric and real ears are not equally sensitive: a published interaural asymmetry of a few per cent in effective radius, or a fixed offset from a difference in the two ears’ thresholds, adds a constant to the delay rather than a factor, and a constant does something the factor above cannot. It displaces the median plane. A listener with a two per cent asymmetry has a straight ahead that is not straight ahead, by an angle this collection can compute from the same three functions — and unlike a gain error, a constant offset is largest exactly where the curve is steepest, which is where the previous rung expected to find the first error and where this rung found none. It needs no corpus and no listener: the arithmetic is the same substitution with an added term, and what would come out is whether the one direction a mis-calibrated map gets right is a property of two ears or a property of the model having only one radius in it.

Part 11 of 12

One essay in the series on localisation. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CalibrationHead shadowInteraural time differenceJust-noticeable differenceLocalisationSpeed of sound