Perception and the listener

How far away the room takes over

Direct sound falls six decibels every time the distance doubles and the reverberant field does not fall at all, so the two cross once. In a concert hall the crossing is at about five and a half metres, which is nearer than nearly every seat — so almost everybody in almost every hall is hearing the building more than the players, and the number that says so is built from two quantities already computed — a room's reverberation time and an instrument's directivity.

Assumes: The room is part of the instrument

Two rungs of this ladder describe quantities that seem to belong to different discussions. Reverberation time is a property of a building — volume over absorption, one number, the same everywhere inside it. Directivity is a property of an instrument — how much of its output goes forward rather than sideways, a strong function of frequency and of the size of the radiating surface.

Combine them and a third quantity appears that is a property of neither: a distance. It is the radius at which the sound that came straight from the instrument and the sound that has been round the set of resonances the room is made of arrive at the same level, and everything about how a listener experiences a room turns on which side of it they are sitting.

Direct and reverberant sound in a shoebox concert hall. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 5.5 metres in a room of 18700 cubic metres with a 2-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument.
Fig. 1 Direct sound and reverberant sound in a concert hall, against distance. The direct falls six decibels for every doubling — it is spreading over a sphere — and the reverberant field is the same everywhere, because it is the energy that has stopped having a direction. They cross at five and a half metres, and the crossing is the whole of this essay.

The two lines, and why one of them is flat

A source radiating into free space spreads its power over a sphere whose area grows as the square of the radius, so intensity falls as one over the radius squared: six decibels per doubling, the inverse-square law, and nothing about a room in it.

The reverberant field is what is left after the sound has bounced enough times to lose track of where it came from. Its level is set by a balance — power going in against power being absorbed per bounce — and it is the same at every point in the room, because energy arriving from every direction has no reason to prefer a position. It does not fall with distance because it is not travelling anywhere.

One line sloping, one line flat: they cross exactly once. The critical distance is 0.057 times the square root of Q times the volume over the decay time — the volume in cubic metres, the decay time in seconds, and Q the directivity factor of the source in the listener’s direction.

Everything in that formula has already been computed on this site. The volume and decay time are Sabine’s; Q is the directivity index expressed as a ratio rather than in decibels. What is new is only that they combine.

How close a listener has to be to hear the instrument. The critical distance in each of six rooms, at three directivity factors. Inside it the direct sound is louder than the reverberant field and a listener hears the player; outside it they hear the room. The number is a few metres everywhere, and the seats in a concert hall begin at about five.
Fig. 2 The distance itself, for six rooms and three directivities. The whole range is between half a metre and ten, which is a much narrower spread than the rooms’ volumes — a factor of seven hundred in volume produces a factor of twenty in distance, because the volume is under a square root and the decay time partly cancels it.

The number for a concert hall, and what it implies

A shoebox hall of 18,700 cubic metres with a two-second decay gives 5.5 metres for an omnidirectional source.

The front row of such a hall is perhaps six metres from the players. Every other seat is further. So for essentially the whole audience, the reverberant field is louder than the direct sound, and what they are listening to is more the room than the instruments.

That is not a complaint. It is the design. A hall is built to add a reverberant field of a particular length and colour to whatever is played in it, and if the direct sound dominated everywhere the hall would be doing nothing. What the number gives is a way of saying how much: at twenty metres, a typical mid-stalls distance, the direct sound is eleven decibels below the reverberant field.

And it explains an experience. A listener in a good hall can still tell exactly where each instrument is, despite the direct sound being far quieter than the reverberation. That is not a contradiction, because localisation does not use the loudest arrival — it uses the first one, and for the first few tens of milliseconds everything that follows is denied a vote. Direction comes from the direct sound and loudness comes from the room, and the two are separately computed — the first from the difference between the two ears in the first arrival, the second from an energy sum over everything that follows.

How late a reflection has to be before it is an echo. What a single reflection does to the sound it follows, against its delay, on a logarithmic axis. Under a millisecond the two combine into one image that is pulled towards the earlier source. From there out to a few tens of milliseconds the reflection is not heard as a separate event at all and does not move the image — it only changes the timbre. Past the echo threshold it becomes a second sound, and the threshold is five times later for speech than for a click.
Fig. 3 The mechanism that lets a listener beyond the critical distance still hear where the players are. Everything arriving in the first few tens of milliseconds after a sound is fused with it and contributes loudness rather than direction, however loud it is. So the reverberant field can dominate the level entirely without disturbing the image, which is exactly what a concert hall relies on.

What it says about smaller rooms

The formula’s most useful readings are at the other end of the scale.

A carpeted bedroom at 34 cubic metres and a 0.4-second decay gives 0.53 metres. Half a metre. A microphone further from a singer than an outstretched hand is recording mostly the room, and a listener standing anywhere in such a room is well beyond the crossing.

A rehearsal room at 250 cubic metres gives 0.95 metres, so a player two metres from another player hears the room more than the colleague.

That last one has a consequence for ensembles that is not obvious. Beyond the critical distance the level of a source stops falling with distance, so moving further from a player in a small room does not make them quieter — it only reduces the direct component while the reverberant component stays put. Balance within an ensemble in a small room is therefore much less controllable by position than it is on a stage, which is a familiar frustration with a computable cause.

Why a bigger room does not always take over later

The rooms figure has a result in it that reads as a mistake and is not: a gothic cathedral of 25,000 cubic metres has a smaller critical distance than a concert hall of 18,700.

The reason is in the formula. Volume pushes the crossing out and decay time pulls it in, and the cathedral’s eight-second decay against the hall’s two seconds is a factor of four where the volume is a factor of 1.3. Under the square root that is a factor of two against a factor of 1.16, and the decay wins.

Stated physically: a long decay means the reverberant field is loud, because sound put into the room stays there. A room that hoards energy overwhelms its own direct sound sooner, however big it is. That is why a listener at the back of a cathedral hears almost nothing but the building and a listener at the back of a good hall still hears an orchestra, despite the cathedral being larger.

Direct and reverberant sound in a large stone church. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 2.4 metres in a room of 7000 cubic metres with a 4-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument.
Fig. 4 The same two curves in a stone church: seven thousand cubic metres and a four-second decay. The crossing is at 2.4 metres — against 5.5 in a hall nearly three times the volume.

Absorption beats size, and by a long way. The church is the smaller room and its critical distance is less than half the hall’s, because stone absorbs almost nothing and the reverberant field it sustains is correspondingly loud. A listener in a cathedral is past the crossing from the second row.

The frequency dependence nobody puts in the formula

Q is written as a constant and it is not one. An instrument radiates evenly below the frequency at which its radiating surface is about a wavelength across and beams above it, so its directivity factor rises with frequency — from 1 in the bottom octaves to 10 or more at the top for a directional source aimed at the listener.

Since the critical distance goes as the square root of Q, a source with Q rising from 1 to 10 has a critical distance three times further at the top of its range than at the bottom. The crossing is not a circle around the player; it is a shape, wide in front and at high frequencies and narrow behind and low down.

Putting the site’s own directivity model under the formula rather than a constant says the shape is much more extreme than that. For a violin-sized radiating surface of nine centimetres:

Q critical distance in the hall
100 Hz 1.0 5.6 m
500 Hz 1.3 6.4 m
1 kHz 2.4 8.5 m
2 kHz 6.4 14.0 m
4 kHz 22.7 26.3 m
8 kHz 88 51.7 m

The one-to-ten range for Q is a description of the middle of the spectrum, roughly a kilohertz to two and a half; above that the piston model runs away, and the critical distance with it. Across a violin’s actual range the crossing moves not by three times but by about nine.

Which turns the essay’s headline into half a statement. For essentially the whole audience the reverberant field is louder than the direct sound is true at 100 hertz, where the crossing is at 5.5 metres, and false at 4 kilohertz, where it is at 26 and most of the stalls are inside it. An audience sits outside the critical distance for the bass of an instrument and inside it for its treble, in the same note, and the seat at which the two swap is somewhere in the middle of the register.

That is a mechanism for something a listener notices and usually attributes to the room’s absorption alone: the reverberant field is darker than the direct sound, because the direct treble carries out into the room and the direct bass does not. A room does not decay evenly is the other half of it, and the two push the same way — one takes the top off what is in the room and the other keeps the top out of it.

The eighty-eight at the top of that column should be read as a direction rather than a number. Q = 1 + (ka)²/2 is a rigid piston in an infinite baffle, and a violin at eight kilohertz is not one: it is a small, irregular, multiply-resonant radiator whose real directivity is high, frequency-dependent and lobed rather than a smooth beam. What the model gets right is that Q rises as the square of frequency once the radiator is a wavelength across, and everything above follows from that alone.

Direct and reverberant sound in a shoebox concert hall. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 5.5 metres in a room of 18700 cubic metres with a 2-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument.
Fig. 5 The two fields drawn against distance rather than as a single crossing point: the direct sound falling six decibels per doubling, the reverberant field not falling at all, and the total that a listener actually receives.

The total is the part worth reading. Beyond the crossing it is almost flat — moving twice as far away costs a listener very little loudness, because what they are hearing is the room rather than the source. That is why the back of a reverberant hall is not quiet, and why it is also not clear.

Where each room stops being a set of resonances. The Schroeder frequency of 3 rooms — a carpeted bedroom at 217 Hz, a shoebox concert hall at 21 Hz, a gothic cathedral at 36 Hz — marked on a logarithmic frequency axis with the ranges of 2 instruments underneath. Below the mark a room is a handful of separable modes and a note's loudness depends on where the listener is standing; above it the modes overlap and the room is described by one decay time.
Fig. 6 The other crossover, for the same three rooms. The two are computed from the same two numbers and go in opposite directions: a long decay pushes the frequency boundary up and pulls the distance boundary in. So a cathedral is statistical almost everywhere in frequency and reverberant almost everywhere in space, which is one description rather than two.

Whose music, and when

The formula is standard architectural acoustics, deriving from Sabine’s work of the 1890s and the diffuse-field theory built on it in the middle of the twentieth century. It is used routinely in hall design and in sound-system design, and the constant 0.057 is the SI form of an expression that appears in every textbook in the field.

Its relevance to music is a claim about how a repertoire and a room fit together, and that claim has a period. Orchestral music of the nineteenth century was written for halls of exactly the size the formula puts almost every seat outside the critical distance in — the same halls whose decay times set a ceiling on how fast the harmony in them can change — and the balance conventions of orchestration — which instruments can cover which, how many strings against how many winds — were established under those conditions. Those conventions are not portable: the same score in a studio, where the critical distance is a metre or two and every microphone is inside it, balances quite differently, which is why an orchestral recording is a made object rather than a captured one.

And the same number explains why amplified music sounds the way it does in the wrong room. A loudspeaker is a highly directional source with a Q of five to twenty over most of its range, which pushes its critical distance out by a factor of two or four — one of the two things a directional loudspeaker is for. The other is that the energy it does not send at the audience is energy the room never gets.

What a microphone is choosing

The distance is the whole of one recording decision, and it is worth stating as one because it is usually taught as taste.

A microphone placed inside the critical distance records mostly the instrument and a microphone outside it records mostly the room, and the boundary is at a specific number of metres that the room supplies. In a hall a spot microphone at a metre and a pair at fifteen are on opposite sides of it by a factor of three, which is why they sound like different recordings of different things — because they are.

In a small room there is no outside-and-inside to choose between: at half a metre the critical distance has already been passed by anything except a very close microphone, so a small-room recording has one usable position and a hall has a continuum of them. That is the difference between the two kinds of room stated as a count of available choices rather than as a judgement.

And the same arithmetic decides whether a room can be treated. Moving the critical distance out means raising Q or lowering the decay time, and only the second is available to somebody who owns a room rather than an instrument. Halving the decay time moves the distance out by a factor of 1.4, and Sabine says what halving the decay time costs: the absorption has to double, which means either twice the treated area or twice the absorption coefficient over the area already treated. Doubling the absorption in a room to move the critical distance out by forty per cent is the whole trade, and it is why small-room recording relies on microphone position rather than on treatment.

The frequency table above sharpens that too, because absorption is not flat either. Soft treatment absorbs the top of the spectrum far better than the bottom, so treating a small room moves the critical distance out most at exactly the frequencies where it was already furthest — and hardly at all in the bass, where it is half a metre and the problem is.

Direct and reverberant sound in a rehearsal room. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 1.0 metres in a room of 250 cubic metres with a 0.9-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument.
Fig. 7 The same two lines in a room of 250 cubic metres. The crossing is at 95 centimetres, so a listener anywhere in the room and a microphone anywhere but very close are both in the reverberant field. The direct sound at four metres is twelve decibels down on the room — a smaller room does not mean less room, it means more of it, sooner.

Where the model stops

The reverberant field is not really uniform. It is uniform in the idealisation, and in a real hall it falls slowly with distance from the source, so the two lines are not quite a slope and a flat and the crossing is a little further out than the formula says. In very absorptive or very elongated rooms — corridors, under balconies — the assumption fails badly.

And the source is a point. Every line here treats the instrument as radiating from one place, which is what lets the direct sound fall as one over the radius squared. A cello is over a metre long and an orchestra is twenty metres wide, and inside a distance comparable with the source’s own size the inverse-square law does not hold at all — so the crossing computed for a bedroom at half a metre is being computed at a distance smaller than the instrument. The formula’s small-room readings are the ones this essay leans on hardest and the ones the point-source assumption serves worst.

Sabine’s decay time carries its own error into this, as the site has recorded, and it is largest in the small heavily damped rooms where the critical distance is most useful. Under a square root the error is halved, which is the only reason the formula is quoted with confidence at those sizes at all.

The reverberant field also masks. A direct sound eleven decibels under a diffuse field of the same spectrum is close to the region where one sound hides another, and the model computes a level ratio rather than an audibility.

And the direct sound is not a point. An orchestra is twenty metres wide and a piano is two and a half, so “distance from the source” is not a single number for either. For an ensemble the critical distance is best read as applying to each player separately, which means different members of the same ensemble cross it at different seats.

What the picture cannot show

It cannot show time. Both lines are steady-state levels, and the perceptual difference between direct and reverberant sound is mostly a matter of when each arrives. A listener at twenty metres receives the direct sound first, then a few strong early reflections, then the diffuse tail — and the early reflections are neither of the two things this figure plots.

It cannot show what is being played. The reverberant level depends on the total power a source has put into the room over the preceding second or two, so a sustained passage builds a reverberant field and a staccato one does not. The crossing is a good description of a held chord and a poor one of a fast passage.

It cannot show two sources at once. Every ensemble balance question is about the ratio between two players’ direct sounds and one shared reverberant field, and the figure has one source in it.

And it cannot show the seat. Every number here is a radius, and a hall is not a sphere. Under a balcony or against a rear wall the local field is nothing like the average, and the average is what the formula computes.

Where the ladder goes next

Two rungs have now been about a single number standing for a whole room — a crossover frequency and a crossover distance. The last one takes away the assumption both of them rest on: that a room has a reverberation time. It has six, they differ by a factor of two or three, and a chord left ringing in a room does not fade so much as change shape.

Part 6 of 9

One essay in the series on room acoustics. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 18.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AbsorptionDiffuse fieldLocalisationPrecedence effectReverberationSabine equationSpherical wave