Timbre and acoustics

The first eighty milliseconds are a different room

Draw a line across a room's decay and the energy on either side is two opposite verdicts about one building — clarity before it, reverberation after. The line is a property of the ear, not of the room, and putting it at 50 milliseconds and at 80 makes the two published design targets fall out — a room for speech at one second, a room for music at 1.6.

Assumes: The room is part of the instrument

The seventh rung of this ladder ended with a list of things its curves could not show, and the first item on it was this: “everything here is the reverberant tail, and the first eighty milliseconds — the early reflections that fuse with the direct sound — have their own spectrum and their own dependence on which surfaces are near the listener. A hall’s clarity is largely an early-energy quantity and none of it is in these curves.”

It is in this one, and getting it requires no new measurement of the room at all.

One decay, two verdicts, and the line is the listener's. The early-to-late energy ratio against reverberation time, for 50 millisecond and 80 millisecond windows. Nothing about the room differs between the curves; only where the line is drawn across its decay. Zero comes at 1.00 seconds for the 50 ms window and 1.59 seconds for the 80 ms window — which are, to two figures, the published rules of thumb for a room for speech and a room for music. The design targets were not put in; they came out.
Fig. 1 The ratio of early energy to late, in decibels, against reverberation time, for two positions of the line. Nothing about the room differs between the two curves. The 50 millisecond window crosses zero at 1.00 seconds and the 80 millisecond window at 1.59, which are the design targets for a room for speech and a room for music.

One curve, cut in two

Sabine’s decay is one exponential: energy falls 60 decibels in T seconds, so it falls by a factor of 10 every T/6.

Draw a line at t₀ and the curve is divided into the energy that arrived before it and the energy that arrived after. For a pure exponential the ratio of the two is 10 to the power of 6t₀/T, less one — a single expression with the decay time and the line position in it, and nothing else.

That is the whole computation. A room does not have to be measured twice to yield both quantities; the second falls out of the first as soon as a line is drawn, and the only new information is where to draw it.

And the line belongs to the ear

The line is not a property of the building.

Energy arriving very soon after the direct sound is not heard as a separate event. It fuses with it, arrives from the same apparent direction, and adds to its loudness — which is the precedence effect, and which is the mechanism by which a listener in a reverberant room hears one instrument in a place rather than a dozen copies.

Energy arriving late enough is heard as reverberation: a wash without a direction, following the sound rather than being part of it.

Between those there is a transition, and the two positions the literature has settled on are 50 milliseconds for speech and 80 for music. The difference is not arbitrary. Speech has to be parsed into syllables, which arrive several times a second, so late energy that would merely blur a sustained note obliterates a consonant. Music tolerates more.

Nothing has been measured that Sabine did not measure. His arithmetic gives one number per room — 1.9 seconds for a concert hall, 4.0 for a stone church, 0.4 for a studio live room — and everything in this essay is a second number read off the same decay curve, cut at a point the ear chooses rather than at the sixty-decibel mark the convention chose. Two rooms with the same reverberation time can have quite different clarities if their early reflections differ, and two rooms with the same clarity can decay for very different lengths of time. One decay, two readings, and the second is the one a listener parses syllables with.

The design targets were not put in

The result worth stopping on is that the two rules of thumb come out rather than going in.

A room for speech should have a reverberation time under about a second. A room for music should be around 1.6 to 2 seconds. Those numbers are in every handbook and are usually presented as accumulated experience.

Set clarity to zero — the point at which early energy exactly equals late — and solve. For a 50 millisecond window the answer is 0.997 seconds. For an 80 millisecond window it is 1.595.

The only inputs were Sabine’s exponential and the two window lengths, and the window lengths come from psychoacoustics rather than from architecture. The handbook numbers are a psychoacoustic constant divided by the logarithm of two, and the accumulated experience was accumulating a fact about ears.

How much of that survives the two fictions

The agreement is striking enough to be worth attacking, and the two caveats at the foot of this essay are the attacks. Both can be run.

A real decay is not one exponential. The early part falls faster, because the first reflections meet the absorbing surfaces before the field is diffuse. Model that as a first hundred milliseconds decaying at r times the later rate and solve for the zero crossings again:

early decay C50 zero C80 zero
1.0 — one exponential 0.997 s 1.595 s
0.8 1.151 1.747
0.6 1.389 1.986
0.5 1.567 2.166

A mild double slope leaves the music target inside its handbook band and pushes the speech target above a second. A strong one takes both out. So the result is robust to the first fiction being slightly wrong and not to its being badly wrong, which is the ordinary condition of a model and is worth saying rather than implying.

The direct sound is the harder one. The published measures put the direct sound in the early energy and this arithmetic has none in it. Add a direct component as a fraction D of the whole reverberant energy:

direct energy C50 zero C80 zero
none — as computed above 0.997 s 1.595 s
a quarter 1.470 2.352
a half 2.401 3.842
equal never never

At the critical distance there is no zero crossing at all. With the direct sound as large as everything the room adds, clarity is positive at every reverberation time, and a listener in the front third of a hall is above zero whatever the building does.

That is the honest boundary on the headline. The zero crossing is not a property of a room; it is a property of a room listened to from well beyond the critical distance, where the direct sound has fallen away and the ratio is between two parts of a reverberant field. The design targets falling out of the arithmetic is a real agreement and it is an agreement about a particular kind of seat — the far ones, which are also the seats a hall is judged by.

It also explains something the next section is about to say. The reason a room’s clarity varies across its seats by more than the difference between two rooms is that the direct term varies across the seats and the ratio it is added to does not.

Which is why one room cannot do both

The same arithmetic explains the oldest problem in hall design, which is that a room for speech and a room for music are different rooms.

At 2.0 seconds, a shoebox concert hall, the 80 millisecond clarity is −1.3 dB — inside the ±2 dB window halls are designed to — and the 50 millisecond clarity is −3.9, which is poor for speech. At 1.0 second, an ordinary lecture room, the numbers are 3.1 and 0.0: good for speech and thin for music.

One decay, two verdicts, and the line is the listener's. The early-to-late energy ratio against reverberation time, for 50 millisecond and 80 millisecond windows. Nothing about the room differs between the curves; only where the line is drawn across its decay. Zero comes at 1.00 seconds for the 50 ms window and 1.59 seconds for the 80 ms window — which are, to two figures, the published rules of thumb for a room for speech and a room for music. The design targets were not put in; they came out.
Fig. 2 The same two curves over the range real rooms occupy. The gap between them is about 3 decibels at every reverberation time — because the ratio of the two window lengths is fixed — so a room can be at zero on one curve or the other and never both.

The gap between the two curves is nearly constant, which follows from the exponential: moving the line from 50 ms to 80 always multiplies the early energy by the same factor. So the two verdicts are locked a fixed distance apart, and no reverberation time makes both zero.

An opera house at 1.4 seconds is the compromise made explicit: it is short because the words have to arrive, and it is short at a measurable cost to the orchestra.

The window is a listener’s constant, so it is not everybody’s

Making the line a property of the ear rather than of the room has a consequence that the design targets hide: the targets are averages over listeners, and the window is not the same for all of them.

Temporal integration slows with age and with hearing loss; the effective window widens, so more of what a room does counts as late energy and the same hall is less clear to some of its audience than to others. A hall at exactly −1.3 dB for a young listener is further down for an older one, and the shift is in the direction that makes a marginal room bad rather than a good room marginal.

One decay, two verdicts, and the line is the listener's. The early-to-late energy ratio against reverberation time, for 50 millisecond and 80 millisecond and 120 millisecond windows. Nothing about the room differs between the curves; only where the line is drawn across its decay. Zero comes at 1.00 seconds for the 50 ms window and 1.59 seconds for the 80 ms window and 2.39 seconds for the 120 ms window — which are, to two figures, the published rules of thumb for a room for speech and a room for music. The design targets were not put in; they came out.
Fig. 3 Three windows rather than two: 50 and 80 milliseconds, and 120 for a listener who integrates more slowly. The zero crossing moves from 1.00 to 1.59 to 2.39 seconds — so a hall designed at 1.9 seconds is above zero for one listener and below it for another, in the same seat, hearing the same music.

That is a real and slightly uncomfortable finding. Every quantity earlier in this ladder was a property of a building; this one is a property of a pairing, and a hall does not have a clarity so much as a clarity for somebody.

A room has six clarities

Rung seven established that absorption is a strong function of frequency and a room therefore has six decay times rather than one. Clarity inherits that immediately.

Clarity band by band in a large stone church. The early-to-late energy ratio at an 80 millisecond window, computed from a large stone church's own six decay times: -7.2 dB at 125 Hz, -7.0 dB at 250 Hz, -6.0 dB at 500 Hz, -5.0 dB at 1000 Hz, -4.5 dB at 2000 Hz, -2.3 dB at 4000 Hz. Absorption is a strong function of frequency, so a room has six clarities as well as six decay times — and a room can be clear in the treble and muddy in the bass by 4.9 decibels at once. A band centre below 140 hertz is sounded two octaves up, because it is below what most speakers reproduce.
Fig. 4 Clarity band by band in a stone church, computed from its own six decay times. It runs from −7.2 dB at 125 hertz to −2.3 at 4 kilohertz — the same room is nearly five decibels clearer in the treble than in the bass.

That is the same fact rung seven reported, restated in the quantity that decides intelligibility. A stone church is not uniformly muddy; it is muddy in the bass and merely reverberant in the treble, which is exactly the description anybody who has been in one gives.

And it explains a piece of practice. Reinforcing the treble of a voice in a large church improves intelligibility more than raising its level does, because level raises early and late energy together and leaves the ratio where it was.

Six reverberation times for one room. Sabine's arithmetic evaluated in each octave band from the published absorption coefficients of the surfaces. a large stone church runs from 6.3 seconds at 125 Hz to 2.4 at 4 kHz — a bass ratio of 1.38, where concert halls are specified between 1.1 and 1.25.
Fig. 5 The six decay times the clarities above are computed from. One measurement, two derived quantities, and the second one is the one a listener’s experience is actually about.

What this does to the harmonic-rhythm ceiling

Rung four argued that a chord in a long room is still sounding when the next arrives, so the reverberation time sets a ceiling on chords per second — and that the music written for the great stone rooms sits under it.

Clarity sharpens that and moves it. What matters for whether two chords are heard as two is not how long the first is still audible, but how much of the second chord’s energy arrives inside the window against how much of the first is still arriving. The ceiling is not a statement about audibility; it is a statement about a ratio.

How often the chord changes, and what a room allows. Chord changes a second implied by each style's stated rate and tempo, on a logarithmic axis, with the rate above which a room leaves more than one earlier chord above 40 dB marked for six rooms. The style rates are conventions rather than corpus measurements and the figure says so; the room rates are arithmetic from the reverberation time.
Fig. 6 The overlap of successive chords in a long room. The clarity picture says the relevant comparison is between the new chord’s early energy and the old one’s tail, which is a ratio rather than a duration — and it moves the ceiling down, because the old chord is still contributing late energy long after it has stopped being separately audible.

The direction of the correction is worth stating: clarity makes the ceiling lower than the audibility argument does. A chord whose predecessor is 40 decibels down is inaudible as a chord and is still adding to the late energy of everything that follows it.

The two windows and the two repertoires

Putting the numbers beside the rooms this site already carries makes the historical picture unusually tidy.

room decay clarity at 80 ms
a gothic cathedral 8.0 s −8.3 dB
a large stone church 4.0 s −5.0 dB
a shoebox concert hall 2.0 s −1.3 dB
an opera house 1.4 s +0.8 dB
a jazz club 0.8 s +4.7 dB
a recording studio 0.35 s +13.5 dB

That column spans nearly twenty-two decibels, which is a larger range than any other quantity this ladder has computed. And it lines up with what was written in each: the cathedral repertoire changes harmony perhaps once every four seconds, the opera house exists so that words arrive, the jazz club is where a rhythm section plays a chord every beat, and the studio is where the reverberation is added afterwards by choice.

One decay, two verdicts, and the line is the listener's. The early-to-late energy ratio against reverberation time, for 80 millisecond windows. Nothing about the room differs between the curves; only where the line is drawn across its decay. Zero comes at 1.59 seconds for the 80 ms window — which are, to two figures, the published rules of thumb for a room for speech and a room for music. The design targets were not put in; they came out.
Fig. 7 The 80 millisecond curve alone with the six rooms placed on it. The vertical spread is the point: a cathedral and a studio are not two settings of one variable, they are twenty-two decibels apart in the quantity that decides whether two events are heard as two.

Where the model stops

A pure exponential is a fiction and this one is a strong fiction. Real decays are not straight lines on a log scale — the early part falls faster than the late part in most rooms, because the first reflections meet the absorbing surfaces before the field is diffuse — so the early energy is systematically overestimated by this arithmetic and clarity with it.

The direct sound is not in it, and a section above prices what that costs: the zero crossing moves to 2.35 seconds with a quarter of the energy direct and disappears entirely at the critical distance. Beyond the critical distance the direct term contributes almost nothing and near the front of a hall it dominates, so clarity varies across the seats of one room by more than the difference between two rooms — and every number in this essay is a far-seat number.

The transition is not a step. Fusion does not stop at 50 or 80 milliseconds; it fades over tens of milliseconds and the fade depends on level, on direction, and on what the sound is. A step function at 80 is a modelling convenience of exactly the kind this ladder has been finding.

It has no directions in it. Early energy arriving from the side and the same energy arriving from the front do different things to a listener — the first adds spaciousness and the second adds loudness — and an energy ratio has no term for either.

What the picture cannot show

It cannot show the seat. Every number here is for a room. A listener is in a place in the room, and the early-to-late ratio there depends on which surfaces are near them, which is exactly what a whole-room quantity averages away.

It cannot show the source. Clarity is measured with an omnidirectional source, and an instrument points: a violin’s treble goes where it is aimed and its bass goes everywhere, so the early energy reaching a listener is a different spectrum from the one leaving the instrument.

It cannot show the modal region. Below the frequency at which a room stops being a set of resonances the decay is not one exponential at all — each mode decays at its own rate — so the 125 hertz clarity of a small room is a number the model should not really be asked for.

And it cannot show whether more clarity is better. The whole apparatus is a ratio with a zero in the middle, and nothing in it says a listener wants to be at zero. What listeners prefer is a matter of repertoire and of taste, and a hall at the top of its window is described as dry by some of the same people who call the bottom of it muddy.

Why raising the level does not help

There is a practical consequence that follows in one line and is routinely got wrong.

Clarity is a ratio of two energies that come from the same source. Turn the source up and both go up by the same amount, so the ratio does not move. Loudness and clarity are independent, and no amount of volume makes a reverberant room intelligible.

The same hall, empty and full. A hall of 18700 cubic metres with 900 square metres of audience, designed to 1.9 seconds occupied, with three kinds of seat under the audience. It is 2.76 s empty and 1.90 s full with hard wooden seats, a change of 31 per cent; 2.29 s empty and 1.90 s full with lightly padded, a change of 17 per cent; 1.96 s empty and 1.90 s full with heavily upholstered, a change of 3 per cent. The audience is 45 per cent of the total absorption when the hall is full, which is the largest single term in the equation — and how much the hall changes is decided entirely by what the seats were doing before anybody sat on them.
Fig. 8 And the room does not hold still, which is where the ratio argument bites hardest. An audience is the largest single absorber a hall contains, so the same building measured empty and full has two reverberation times and therefore two clarities — and the empty measurement is the one a rehearsal is held in. Turning the source up cannot recover the difference for the same reason it cannot recover any of the others: clarity is a ratio of two energies from one source, so raising the source raises both by the same amount and leaves every ratio exactly where it was. Loudness and clarity are independent, and no amount of volume makes a reverberant room intelligible.

What does move the ratio is anything that adds early energy without adding late: a reflector above the source, a listener nearer to it, a directional source aimed at the listener, or a loudspeaker close to the ear. All four are the same intervention, and it is the reason a public-address system in a large church is a distributed array of small speakers rather than one big one at the front — many nearby sources add direct energy to each listener and put no more energy into the room.

That is also why an instrument that points is loud in a way that a room cannot spoil. Directivity puts a larger share of the radiated energy into the direct path, which raises clarity at the listener rather than merely raising level.

Whose halls, and when

The 50 and 80 millisecond windows are conventions of twentieth-century room acoustics, arrived at from intelligibility testing in the first case and from listening tests in the second, mostly in Europe and North America.

The ±2 decibel target for C80 is a specification for the orchestral repertoire in purpose-built halls. It is not a fact about music. The rooms the earliest notated polyphony was written for have a C80 of about −7, and the music written for them is written to be heard that way — long notes, few simultaneous changes, and a harmonic language that does not require two chords to be told apart in a second.

So this rung’s numbers are a design tool for one tradition and a descriptive tool for all of them. Read descriptively, the interesting thing is that composers wrote to whatever their room’s number was.

What the ear is doing to earn the window

The window is a number taken from listening tests, and it is worth saying what mechanism it is a number about, because the site’s habit is to name the model rather than to imply one.

An onset is not an instant to the auditory system. Energy arriving over some tens of milliseconds is integrated into one percept, which is why a sound’s loudness grows with its duration up to a point and stops growing after it, and why two attacks closer than about thirty milliseconds fuse into one event. The clarity window is the same integration seen from the other side.

What the early window captures is the attack and the first part of the body; what falls outside it is the release the room has added. Drawn as two amplitude envelopes for the same note — one in a dry room, one in a long one — the two are identical for the first eighty milliseconds and then diverge by a factor of twenty in how long they take to disappear. A room does not touch a note’s onset, because no reflection can arrive before the direct sound; it changes only what the note leaves behind. That is the same asymmetry the clarity ratio is built on, seen in the time domain rather than as a number.

There is a second mechanism in the same window and it is the precedence effect: a reflection arriving inside it is not merely counted as early energy, it is assigned to the direct sound’s direction and heard as part of the same event. Clarity is an energy ratio and precedence is a spatial judgement, and they share a window because they share the integration that produces it.

The ladder from here

Two quantities have now been read off one decay curve — how long it takes and how it divides — and both belong to a room with nobody in it.

The last rung of this ladder puts people in. The audience is the largest single absorber in a hall and Sabine’s equation has one term for it, and the players close a loop the model has no term for at all.

Part 8 of 9

One essay in the series on room acoustics. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 16.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AbsorptionClarityIntegration timePrecedence effectReverberationRoom acoustics