Perception and the listener

Where the two ears stop agreeing

A room sends both ears versions of the same sound, alike at low frequencies and increasingly unlike at high ones. Where they stop resembling each other is 980 hertz, and it is set by the 17.5 centimetres between the ears rather than by anything about the room — which is within a quarter of a frequency found earlier for a completely different reason. One minus that correlation is spaciousness, and it is computable from a room's own reverberation.

Assumes: The beat that is not in the air · How far away the room takes over

The beat that is not in the air ended by naming a question the two halves of this collection could answer together and never had. A room sends both ears versions of one sound that are alike low down and unlike high up; the frequency at which they stop resembling each other is computable from the room’s own reverberation; and the room ladder has had every quantity needed to compute it since its sixth rung.

Where a room stops sending the two ears the same soundThe correlation between the two ears' signals against frequency, for a seat 15 metres from the source in a 15,000 cubic metre room with a 2 second reverberation time. The pale curve is the diffuse field alone — sin(kd)/(kd) for an ear separation of 17.5 centimetres, which first crosses zero at 980 hertz. The heavy curve adds the direct sound, which is coherent and lifts the whole thing by an amount the direct-to-reverberant ratio sets. At 125 hertz the coherence is 0.98 and at 1000 it is 0.16.980 Hz62.51252505001k2k4k8k-0.20.20.40.60.81interaural coherencefrequency, hertzthe diffuse field alonesin(kd) / kdwith the direct soundat 15 mspaciousness 0.31at 500 Hz
Fig. 1 The correlation between the two ears’ signals against frequency, for a seat fifteen metres from the source in a shoebox concert hall. The pale curve is the reverberant field alone and the heavy one adds the direct sound, which is the same at both ears and lifts the whole thing. At 125 hertz the two ears receive nearly identical signals; at 1000 they receive almost unrelated ones.

The crossing is at 980 hertz and nothing about the room is in it. It is the speed of sound over twice the distance between the ears, and the only quantity in the formula is 17.5 centimetres.

The two limits the head sets

That number is worth putting beside the one this ladder started with.

Two ears, and the whole of the difference is 655 microseconds established the interaural time difference: the largest delay a head can produce is the time sound takes to travel round it, and above the frequency at which half a period equals that delay a phase difference stops naming a direction. That frequency is 762 hertz, and it is the ceiling on where the timing mechanism can work.

The coherence zero is 980 hertz, and it is where the two ears stop receiving comparable signals at all in a reverberant field.

Two limits, both between 750 and 1000 hertz, both from the head’s own size. They look like different calculations — one is a half-period against a travel time and the other is the first zero of a sinc — and the section on the arithmetic below shows they reduce to one expression, c over twice a distance, evaluated at the two distances a head has: the straight line between the ears and the path around them. The gap between them is the ratio of those two lengths and nothing else.

The whole of the delay is 655 microseconds. Interaural time difference against the direction of the source, from Woodworth's formula for a sphere of radius 8.8 cm. The entire usable range is 656 microseconds, from hard left to hard right; a listener resolves about ten of them, so the ear is doing arithmetic on a scale about a thousandth of the period of the note it is listening to.
Fig. 2 The earliest figure, drawn again: the interaural delay against the direction it implies, with the frequency above which a phase difference is ambiguous. The 762 hertz here and the 980 hertz above are the same object measured two ways, and between them they say that everything a head does with two ears is a low-frequency business.

Everything the binaural system does well, it does below a kilohertz. Above it, direction is carried by level rather than by timing, and the two ears’ waveforms are not comparable at all.

It is worth being explicit about what “stop agreeing” means, because it is easy to hear it as a claim about damage. The two ears are receiving the same sound — the same instrument, the same note, the same music — and what stops matching is the moment-to-moment waveform. Below the crossing a peak at one ear is a peak at the other; above it the two are as unrelated as two microphones in different corners. Nothing is lost, and something is gained: the difference between two ears’ signals is the entire physical basis of a sound having width, and a source with no difference between the ears is a source inside the listener’s head.

That is why the quantity is worth computing rather than merely bounding. Coherence is not a defect being measured; it is a resource, and the room is what supplies it.

The direct sound is what lifts the curve

The diffuse field is the room’s limit and no listener is only in a diffuse field. There is also the direct sound, which arrives from one place, and is therefore identical at the two ears up to a delay.

Where a room stops sending the two ears the same soundThe correlation between the two ears' signals against frequency, for a seat 4 metres from the source in a 15,000 cubic metre room with a 2 second reverberation time. The pale curve is the diffuse field alone — sin(kd)/(kd) for an ear separation of 17.5 centimetres, which first crosses zero at 980 hertz. The heavy curve adds the direct sound, which is coherent and lifts the whole thing by an amount the direct-to-reverberant ratio sets. At 125 hertz the coherence is 0.99 and at 1000 it is 0.75.980 Hz62.51252505001k2k4k8k-0.20.20.40.60.81interaural coherencefrequency, hertzthe diffuse field alonesin(kd) / kdwith the direct soundat 4 mspaciousness 0.09at 500 Hz
Fig. 3 The same hall from a seat four metres away rather than fifteen. Inside the critical distance the direct sound dominates, the coherence stays high across the whole spectrum, and the sound is precise and narrow. The curve’s shape is the same and its height is set entirely by the direct-to-reverberant ratio, which is the quantity the room model computes from the volume, the reverberation time and the source’s directivity.

So the coherence at a seat is an energy-weighted average: a coherent direct part and an incoherent reverberant part, in a proportion the room decides. That gives a number, and the number has a name in the hall literature.

Spaciousness is one minus the coherence, averaged over the low bands, and it is one of the two or three measures concert halls are actually judged on.

Spaciousness against where the listener is sitting. One minus the interaural coherence, averaged over the 125, 250, 500, 1000 hertz bands, against distance from the source in five rooms of the same volume and different reverberation times. Every curve rises steeply out to the critical distance and then flattens, because past it the sound is nearly all reverberant and there is nothing left to make coherent. A seat three metres from the source in a dry room scores 0.01; a seat forty metres away in a stone church scores 0.38.
Fig. 4 Spaciousness against distance in five rooms of the same volume and different reverberation times. Every curve rises steeply out to the critical distance and then flattens, because past it the sound is nearly all reverberant and there is no coherent part left to remove. The interesting consequence is the flattening: most of a hall’s seats have nearly the same spaciousness, and the ones that do not are the ones near the front.

There is one shape in the curve worth pointing at, because it looks like an artefact and is not. The diffuse coherence does not fall to zero and stay there — it crosses zero at 980 hertz, goes slightly negative, and then oscillates about zero with a decaying amplitude. The negative lobe means the two ears’ signals are anti-correlated over a band around 1400 hertz: a peak at one ear tends to be a trough at the other.

That is what a sinc does, and it is a consequence of the field being isotropic and the two points being a fixed distance apart. The lobe is centred at 1,402 hertz and reaches −0.217 at its deepest, and the zeros above the first are at 1,960 and 2,940 — evenly spaced, because a sinc’s zeros are, which means the anti-correlated bands and the uncorrelated ones alternate at a fixed spacing of 980 hertz all the way up.

Whether a listener’s two ears ever see any of that is another matter — a real head’s shadowing damps the lobes heavily — but it is in the model, it is drawn rather than hidden, and it is the reason the vertical axis in the figures goes below zero. What matters for the spaciousness measure is that everything above the first zero oscillates about zero with an amplitude under a quarter, so the whole region above a kilohertz is near-incoherent whatever the fine structure is, which is why averaging it into a single number loses nothing and why the hall literature does not.

What this joins up

The room ladder and the localisation ladder have been computing the same quantities without sharing them, and this rung is the join rather than a new measurement. Three things fall out of it.

A seat is a position on one curve. The spaciousness figure has one variable a listener controls — where to sit — and it says the control is nearly all spent in the first ten metres. Moving from the third row to the tenth changes the number a great deal; moving from the twentieth row to the fortieth changes it hardly at all. That is a fact about buying a ticket that falls out of an ear separation and a reverberation time.

A hall’s spaciousness is a property of its reverberation, not of its shape. In this model — and the model is doing a lot of work here — the only room quantities are the volume, the reverberation time and the source’s directivity, all of which the room ladder computes. Two halls with the same reverberation time have the same spaciousness at the same distance whatever they look like.

That is the model’s weakest claim and it is worth flagging as such rather than as a finding, because it is exactly what the hall literature disagrees with. Halls are judged on lateral energy, and a shoebox and a fan-shaped room of identical volume and reverberation time deliver very different amounts of it — the shoebox’s side walls return early reflections from the sides and the fan’s return them from behind. The model here has no directions, so it cannot see that difference and reports the two as identical. What it does supply is the part that is not about shape: the frequency at which coherence collapses, and the distance at which the direct sound stops defending it. A shape argument has to be built on top of those rather than instead of them.

Spaciousness saturates. Past the critical distance the direct sound is a small share of the energy and removing more of it changes little. How far away the room takes over computed the critical distance for exactly this reason and did not know it was computing this.

And the measure only discriminates at low frequency. Above 980 hertz the diffuse coherence is near zero in every room, so a spaciousness number computed over the whole spectrum would be nearly the same everywhere. That is why the hall literature scores the 125-to-1000 hertz bands, and it is a fact about the listener rather than a convention.

Direct and reverberant sound in a shoebox concert hall. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 7.8 metres in a room of 18700 cubic metres with a 2-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument.
Fig. 5 The distance at which the direct and reverberant energies are equal, which the room model computes and which the curves above bend at. It is proportional to the square root of the volume over the reverberation time, so a large dry room has a critical distance of tens of metres and a small live one of two or three — and inside it a listener is hearing an instrument, outside it a room.
500 Hz in one ear, 504 in the other. Two tones 4 hertz apart, one to each ear. They never meet in the air, so neither eardrum sees any modulation at all and there is no acoustic beat to hear. What changes is the phase between the ears, which advances a whole cycle every 250 milliseconds — and the direction that phase implies sweeps with it, drawn here as azimuth against time. The sweep is clipped at the edges, because the implied delay leaves the range a head can produce. A head 17.5 cm across gives at most 656 microseconds, so the phase stops naming a direction above 762 Hz.
Fig. 6 An earlier figure, which is the extreme case of an incoherent pair: a different tone to each ear, which cannot beat in the air because nothing sums, and which produces a fluctuation anyway. That essay found the fluctuation ends at a frequency the head’s size sets. This one finds a room producing the same kind of dissimilarity, continuously, at every frequency above a limit the same size sets — so the binaural beat and the spaciousness of a hall are the same mechanism given two very different inputs.

Which computation produced the numbers

Three formulae and one join.

The diffuse-field coherence between two points a distance d apart is sin(kd)/(kd), where k is 2πf/c. That is a standard result for an isotropic field and it is a sinc function, so its first zero is at kd = π and its frequency is c/2d. With d = 0.175 metres and c = 343 metres per second, that is 980 hertz.

The direct-to-reverberant ratio at a seat is the room ladder’s own directToReverb, unchanged: the direct energy falls as one over the square of the distance and the reverberant energy is constant across the room, so the ratio is fixed by the distance, the source’s directivity and the room constant, which comes from the volume and the reverberation time.

The seat’s coherence is then the energy-weighted average of one — the direct sound, coherent — and the diffuse value, weighted by that ratio. The spaciousness is one minus it, averaged over four octave bands.

The one number that is neither computed nor published is the ear separation. 17.5 centimetres is twice the head radius this site has used since the perception phase, and it is the straight-line distance rather than the path around the head. Using the path instead — which is what the interaural delay uses — puts the zero not at “about 700 hertz” but at 762.4, which is the interaural ceiling to a tenth of a hertz.

That is not a coincidence and it is worth writing out, because it says what the two limits actually are. The path around a sphere between the two ear points is a(π/2 + 1) = 22.49 centimetres, which is 1.285 times the straight-line 17.5. The interaural ceiling is one over twice the maximum delay, and the maximum delay is that path over the speed of sound — so the ceiling is c over twice the path. The coherence zero is c over twice the separation. Both limits are the same formula, c/2d, with two different d’s, and the only reason they are not identical is that the two calculations are entitled to different distances: a diffuse-field coherence between two points uses the straight line between them, and a diffraction delay uses the path around the obstacle.

So the “two independent limits landing within a quarter of each other” is a smaller coincidence than it looked and a cleaner fact than it looked. They are one expression of the head’s size, evaluated at the two lengths a head has, and the 29 per cent between them is the ratio π/2 + 1 to 2 — a number with no acoustics in it at all.

Where the model stops

The head is not two points in free air. A real head shadows, diffracts and has pinnae, and every one of those changes what arrives at each ear as a function of both frequency and direction. The sinc is the coherence between two microphones in a diffuse field, and the coherence between two ears on a sphere is measurably different — closer to one at low frequency and less deeply negative around the first zero.

A reverberant field is not isotropic. The formula assumes sound arriving equally from every direction, and a real hall’s tail arrives preferentially from the sides in a good hall and from above in a bad one. The published spaciousness measures are weighted to lateral energy for exactly that reason, and this model has no directions in it at all.

So the numbers here are too small. Measured halls report a low-band spaciousness of 0.4 to 0.7; this model gives 0.3 at a distant seat in a reverberant hall. The shape is right, the ordering is right, and the level is low because a real hall’s tail is more lateral than an isotropic one and a real head decorrelates more than two points do.

And nothing here is about the early reflections. The first eighty milliseconds are a different room, and the hall literature’s spaciousness measure is computed on exactly that window rather than on the whole decay, because early lateral energy is what widens a source and late energy is what envelops a listener. Those are two different perceptual quantities and this model computes something between them.

Whose music, and what a hall was built for

The arithmetic is about a listener and does not know what is being played. The consequence is about buildings, and the buildings were built for repertoires.

The shoebox halls of the nineteenth century — narrow, high, with side walls close enough to send strong early reflections from the sides — are the ones that score highest on lateral energy, and they were built before anybody could measure it. The fan-shaped halls of the twentieth century, designed for sightlines and for getting more seats near the stage, score badly on it and are the halls whose acoustics are complained about.

That is a well-known story and this rung adds a small piece to it. The reason the shape matters and the reverberation time does not settle the question is precisely the isotropy assumption above. A model with only volume and reverberation time — which is this one — says a fan-shaped hall and a shoebox with the same decay are the same. They are not, and the difference is entirely in where the energy comes from, which is the term this model drops.

For the earlier repertoire the reading is different again. The stone rooms the earliest notated polyphony was written in have reverberation times of six to eight seconds, so a listener anywhere but the front is far past the critical distance and the coherence is at its floor across the whole spectrum. Everything in such a room is maximally spacious and nothing in it is precisely located, which is a fair description of what that music sounds like in the buildings it was written for.

What the picture cannot show

Whether the model’s saturation is real. The flattening past the critical distance is a property of a statistical room in which the reverberant energy is the same everywhere, and a real hall’s late energy is not uniform — it falls with distance too, more slowly than the direct sound but not by nothing.

Whether coherence is what a listener is judging. Spaciousness is a report, the interaural cross-correlation is a signal statistic, and the connection between them is a correlation across halls rather than a mechanism.

Nor whether the low bands are the right bands. The measure is averaged over 125 to 1000 hertz because that is where it discriminates, and it is worth noticing that this is the same range in which a room’s own modes stop being countable and in which the bass ratio is measured. Three different hall measures agree on which end of the spectrum they live at, for three different reasons.

And the source is a point. An orchestra is thirty metres wide and its parts are at different distances, so the direct sound at a seat is not one coherent arrival but many, from different directions, with their own delays — which decorrelates the “coherent” half of the average and is not in the model.

Where this ladder goes next

Four rungs. Two ears and 655 microseconds; a room full of copies and the first wavefront winning; a periodicity in neither ear’s signal; and now, the frequency above which the two ears stop receiving comparable signals at all, and the spaciousness that follows from it.

The rung after it is the one the isotropy assumption names, and it is the one that would make this a model of a hall rather than of a decay. Lateral energy is the term every published spaciousness measure weights and this model omits; computing it needs a room with directions in it, which means an image-source model rather than a statistical one — a source, a shoebox, its mirror images in six surfaces, and each reflection arriving from a computable angle at a computable time. That is a construction this collection has never built and it is the same object the precedence rung would need to say which reflection wins.

Part 4 of 12

One essay in the series on localisation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Critical distanceDiffuse fieldHead shadowInteraural time differenceLocalisationPrecedence effectReverberationRoom acoustics