A smaller head in the same hall
Assumes: Two ears, and the whole of the difference is 655 microseconds · The turn is half the angle
There is a way of finding the next thing to say about a subject that has worked on this collection more than once: look through every figure an anchor has drawn and find the parameter all of them set to the same value.
For this one the answer is not the room. The hall has been 22 metres wide and 45; the absorption has been 0.10, 0.18 and 0.70; the seat has been front, middle, rear and side; the model has run to six reflections deep and to two and to ten; the frequency has been 250 hertz and 8,000. Every one of those has been swept by some figure, and several were swept precisely because a rung existed to sweep them.
The head has been 8.75 centimetres in every drawing this anchor has ever made. It is the only quantity in the model that belongs to the listener rather than to the building, and no placement has ever passed another value.
One shape, four sizes
Woodworth’s result is r(θ + sin θ)/c, and r stands outside everything that varies. So the four curves above are not four different functions: they are one function at four heights, and every delay on the small one is 0.80 times the corresponding delay on the adult’s — exactly, to every digit, because 7.00 over 8.75 is 0.80.
That is a strong enough statement to be worth checking, and it is checked: the figure asserts the ratio of each head’s whole range to the adult’s against the ratio of the radii, at machine precision, and would refuse to draw if they came apart.
A newborn’s head is about 5.75 centimetres in radius, a six-year-old’s about 7.00, an adult’s 8.75, and a large adult’s 9.80. The whole interaural range those produce is 431, 525, 656 and 735 microseconds — a factor of 1.7 from the smallest listener in the room to the largest.
If that were the end of it there would be no rung here, because a scaled picture carries the same information as the picture it is scaled from. A photograph enlarged is not a photograph with more in it.
Except that the threshold does not scale
The reason there is a rung is that one quantity in this ladder does not have a head radius in it.
Two ears and one difference reports that a listener detects a change of about ten microseconds in an interaural delay, and the rungs that work in a room use fifteen. Neither number is a property of the skull. It is a property of the machinery behind the ears that compares two spike trains — the same machinery that makes a beat out of two tones that never meet in the air — and there is no reason for it to shrink when the head does.
So the axis shrinks and the ruler laid against it does not. The number of just-distinguishable delays across the whole range — the number of positions a listener has, in the crudest possible sense — goes from 98 for a large adult to 57 for a newborn.
That is the finding, and it is worth being exact about what is and is not happening in it. Put four listeners of different sizes in the same seat of the same hall. The same fifty-four reflections arrive, at the same instants, from the same directions, carrying the same energies. Nothing about the room is different for any of them. What differs is the axis those arrivals are mapped onto, and it differs by exactly the ratio of the radii.
The mid-stalls image is 315 microseconds wide for an adult and 252 for a six-year-old, which in thresholds is 21.0 against 16.8. Four fewer distinguishable positions across the width of a concert hall, from being six years old.
The same effect, worse where the hall is worse
At the back of the hall the effect compounds with the one the eighth rung of this ladder found. There, a third of the energy arrives from behind the listener and folds onto the front, which narrows the image before any head size is applied: the adult’s width falls from 315 microseconds to 260. Scale that by the head and a newborn is reading a hall that is 171 microseconds wide, which is 11.4 thresholds.
Eleven positions is not many for a room forty metres long. It is roughly the resolution of a bad photograph, and it is what the geometry gives before any of the things that would make it worse — noise, inattention, an unfamiliar room — are added.
Two honest qualifications go with that number and neither rescues it. The first is that a smaller listener is usually closer to the source, since children sit at the front, and proximity raises the direct-to-reverberant ratio, which strengthens the direct sound against everything that folds. The second is that a threshold of fifteen microseconds is an adult’s, measured on adults; nobody in this collection has a developmental figure, and if the threshold is worse in a child the two effects multiply rather than cancel.
The two cue boundaries move together and stay one interval apart
The delay is not this ladder’s only cue, and the second one has a head radius in it too — in a completely different place, which is what makes the next figure worth drawing.
A head casts an acoustic shadow only where it is large compared with the wavelength, so the level difference appears above roughly c/2πr. And the delay stops naming a direction above 1/(2·ITDmax), which is where half a wavelength becomes shorter than the longest path difference the head can make and two answers fit the same interaural phase. Where the two ears stop agreeing computed the second for an adult and got 762 hertz.
Both boundaries are proportional to 1/r, so a small head has both of them higher: a six-year-old’s delay cue survives to 953 hertz where an adult’s gives out at 762, and a newborn’s to 1,160.
That is a large musical distance. An adult’s changeover sits around G5 and a six-year-old’s around B♭5 — nearly a fifth higher — so the register in which a listener localises by timing rather than by level is not a fixed band of the spectrum. It descends through the top of the treble staff over a childhood.
And here is the part that does not move at all. Write the two boundaries out and divide:
The radius cancels and so does the speed of sound. The ratio is 1.2220 for every head that has ever existed, which is 347 cents — a neutral third, between the equal-tempered minor and major. A head decides where the changeover between the two cues sits and cannot decide how wide it is.
That is the kind of result this collection looks for and does not often find: a quantity that survives every parameter in the model. It is not a deep fact — it follows from both boundaries being one length over one speed, with the geometry supplying two different dimensionless constants — but it does mean that the awkward band where neither cue is at its best is a fixed interval rather than a fixed number of hertz, and that no listener has more or less of it than any other.
Both cues are weaker, and they are weaker at once
It would be a tidier story if the small head lost on one cue and gained on the other, and it does not.
The level difference a head makes rises as the square of its size against the wavelength, so at any frequency below the shadow’s onset a smaller head casts less of one. At a kilohertz and a source at the side, the four listeners get 3.2, 4.2, 5.5 and 6.3 decibels. A newborn has both a shorter delay axis and a shallower level difference at the same pitch, and the compensation — that their delay cue survives higher up the spectrum — only helps above the frequency an adult’s has already given out at.
So the honest summary is that the small listener is worse off across the band where music mostly is, and better off in a band where an adult has nothing. Whether that is a fair exchange depends on where the energy is, and the reverberant field of a hall is broadband: a room sends both cues at once and a listener uses whichever works at each frequency. Counting the exchange properly would need a spectrum for the reverberation, which this ladder has never carried — its arrivals are broadband impulses with a level and a direction.
The one thing the exchange does not touch is the interval between the boundaries, which is the same 347 cents for all four. Whatever a listener gains at the top they lose at the bottom, in exactly the proportion that keeps that band’s width fixed.
And then the parameter that changes nothing
The head radius is not the only quantity held fixed everywhere. So is the speed of sound: 343 metres a second in every figure, which is dry air at twenty degrees. A hall in January and a hall full of people in July differ by more than that.
The obvious expectation is that it behaves like the head, since it appears in the same formula. It does not, and the reason it cannot is the most useful sentence on this page.
Every interaural delay is a length over c, and every arrival time in the room is a length over the same c. Warming the air divides both by the identical number. So the whole binaural picture of a hall — the arrival times down one axis and the interaural delays across the other — is similar to itself at every temperature, and nothing about its shape can move.
The only thing temperature can do is slide the picture against a threshold that is fixed in microseconds, and over the whole range a hall is ever at, that is small. From freezing to forty degrees the largest delay runs from 679 microseconds to 634 — 45 microseconds, three thresholds at the extreme azimuth and none at all near the front, where the curve’s own value is zero however fast the air is.
So a cold hall and a warm hall are the same hall to two ears, and every acoustic consequence of temperature is somewhere else: in the pitch of the wind instruments, which this collection computes at length, and in the absorption of the air, which changes how much high-frequency energy survives a long path and therefore how long a room appears to ring at the top of the spectrum.
That is a null and it is worth recording as one. Two parameters sit in the same expression; one of them changes what a listener can distinguish and the other cannot, and the difference is not a matter of size. It is that c is shared between the head and the room and r belongs to the head alone. A parameter shared by both halves of a model can only rescale it. A parameter belonging to one half can reshape it.
What it does to the picture the hall is drawn on
Reading that figure with the four ratios in hand is the whole argument in one look. The vertical axis is time in the room and does not move. The horizontal axis is the head, and it contracts. The reflections keep their arrival times and lose their spread.
There is a consequence for the width the seventh rung named that goes slightly against expectation. A smaller listener does not hear a narrower hall in the sense of hearing the sound come from a smaller angular sector — the directions are the same directions, and a source forty degrees to the left is forty degrees to the left for everybody. What they have is a coarser representation of those directions. The hall is the same width and is measured in fewer marks.
That distinction matters because the two would be confused by any experiment that asked a listener to point. Pointing is a motor act calibrated on the listener’s own body, and a listener whose delays are all 20 per cent smaller can still learn that a given delay means a given direction. What they cannot learn is to resolve two directions the delay does not separate.
Which computation produced the numbers
The head is Woodworth’s sphere throughout, at four radii, with the level difference from the low-frequency shadow approximation this ladder’s fourth rung uses. Neither is a measured head-related transfer function and the ladder says so wherever it matters.
The room is unchanged from the fifth rung: a 22 by 40 by 15 metre shoebox, a source on the platform, image sources to the sixth order, eighteen per cent absorption everywhere and spherical spreading. The width of the image is the energy-weighted root-mean-square of the interaural delays over everything arriving within 120 milliseconds of the direct sound.
The head radii are round figures for a newborn, a six-year-old, an adult and a large adult, and none of them is a measurement of a person. The direction of every result is secure at any plausible set of values, since all of them are proportional to r; the third digit of any width is not.
The temperature sweep uses 331.3√(1 + T/273.15) metres a second, which is the standard expression for dry air.
Where the model stops
The threshold is an adult’s. Every claim about how many positions a small listener has multiplies a scaled delay by a threshold measured on grown-ups, and a developmental figure would change all of them. It is the single number this rung would most like to have.
A sphere is not a head, and a child’s head is not a small adult’s. Its proportions differ, its pinnae are relatively larger, and the flesh of it is a different acoustic obstacle. The scaling here is geometric and nothing else.
And nothing in it grows. A listener’s head changes size over about fifteen years while the mapping between a delay and a direction is being learned, so what a real listener has is not a fixed axis at all — it is one that lengthens under a calibration that has to keep up. Whether that is easy or hard is a question about learning and there is nothing about learning here.
What the picture cannot show
It cannot show the pinna. The one cue that breaks the front-back symmetry without movement is a direction-dependent spectral notch above about six kilohertz, and it scales with the pinna rather than with the head — so it does not follow the same law and this collection cannot say how it scales at all.
Nor can it show the shoulders. A real head sits on a torso that reflects, and the torso’s contribution to a direction cue arrives at a delay of its own.
It cannot show the precedence effect. The first wavefront wins whatever the later ones do, and nothing about that mechanism has a head size in it that this ladder can find.
It cannot show two listeners at once. Every figure here is one seat, and the reason head size matters musically is that an audience contains all of these listeners simultaneously, hearing the same performance at different resolutions. Nothing about a hall’s design can be optimised for all of them.
And it cannot show a moving listener. The turn that resolves the front-back fold costs half a minimum audible angle, and the minimum audible angle scales as 1/r — so a smaller head has to turn further, by exactly the same factor, and this rung’s arithmetic and that one’s compose without either of them saying so.
Whose halls, and whose heads
The room is a generic European shoebox, and the heads are round numbers rather than people.
There is one place where the practice has been ahead of the arithmetic, and it is worth naming. Every measurement of a hall that involves a head at all is made with a dummy — a manikin built to an adult standard, with an adult’s ear separation, and the two international standards for such things differ from each other by about a centimetre of radius. That difference alone is worth eight per cent of every interaural delay a hall’s published numbers contain, which is five thresholds at the extreme azimuth. Nothing on this page suggests that measuring halls with children’s heads would be useful. It does suggest that the one head the discipline measures with is a choice, and that the choice has a size.
Where this ladder goes next
Ten rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity in neither ear’s signal; the frequency above which the two ears stop agreeing; a room with directions in it; the threshold that sorts them; what the ones below it do instead of nothing; the whole response in the units a listener has; the turn that recovers the third of it the head throws away; and now the listener, who is not one size and whose threshold does not change when they are.
What is owed after this is the learning. Every claim in this ladder maps a delay to a direction through one fixed geometry, and a real listener acquires that map while the geometry is changing under them by seventy per cent. The interesting question is not whether the map can be learned but what happens to it when it is wrong: a listener whose internal head is a few millimetres out reads every azimuth with a systematic gain error, and the error is largest exactly where the curve is steepest, which is the median plane, where the direct sound is. This collection has the curve and the threshold and could compute how large a mis-calibration is detectable and where a listener would first notice it — an arithmetic question about a psychological process, which is the kind this ladder has answered before and which would need no corpus and no experiment to state precisely.
Part 10 of 12
One essay in the series on localisation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Head shadowImage-sourceInteraural time differenceJust-noticeable differenceLocalisationSpaciousnessSpeed of sound
- An echo is prevented by the crowd around it image-source, localisation