Timbre and acoustics

The room is part of the instrument

A room has frequencies it supports and frequencies it will not. In a small one those frequencies are far apart, so some bass notes are loud in one corner and absent in another — and no equipment fixes it.

No sound has ever reached a listener unmodified. Between the instrument and the ear there is a room, and the room is not a neutral conduit — it is a resonator with modes of its own, and it imposes them on everything that passes through.

A small room's lowest modesThe first few axial standing waves of a room, drawn in plan, with the frequency of every mode below 160 hertz listed underneath. The low modes are far apart in frequency, so some bass notes are loud in one corner and absent in another. The sound buttons play these two octaves above their real pitch, because a room's lowest modes are below what most speakers reproduce.mode 1 · 40.8 Hzloudest at both endsmode 2 · 81.7 Hz2 quiet linesmode 3 · 122.5 Hz3 quiet lines4.2 × 3.4 × 2.5 m — every mode below 160 Hz:4150658094107118129140151gaps here are notes the room does not support
Fig. 1 The lowest standing waves of a small room, drawn in plan, with every mode below 160 hertz listed underneath. The dashed lines are the pressure nodes — places where that mode is silent. The gaps in the list at the bottom are frequencies the room does not support.

A rectangular room supports standing waves the same way a string does, and for the same reason: the boundaries impose conditions, and only certain wavelengths fit. For a room of dimensions L×W×HL \times W \times H, the modal frequencies are

fpqr=c2(pL)2+(qW)2+(rH)2,f_{pqr} = \frac{c}{2}\sqrt{\left(\frac{p}{L}\right)^2 + \left(\frac{q}{W}\right)^2 + \left(\frac{r}{H}\right)^2},

with p,q,rp, q, r non-negative integers not all zero. That is the whole calculation, and it is what the figure above is drawn from.

Why small rooms are worse

The problem with a small room is not that it has modes but that they are far apart.

The lowest axial mode of a room is set by its longest dimension: for a 4.2-metre room, about 41 hertz. The next is at 82, the next at 122. Between them the room supports nothing along that axis, and a note landing in a gap is quiet everywhere while a note landing on a mode is loud in some places and inaudible in others.

The number of modes below a given frequency grows as the cube of that frequency, so the density increases fast. At some point the modes are packed closely enough to overlap, and the room stops behaving as a set of discrete resonances and starts behaving statistically — the diffuse field, where reverberation is describable by a single number.

That crossover is the Schroeder frequency, roughly

fS2000T60V,f_S \approx 2000\sqrt{\frac{T_{60}}{V}},

with reverberation time in seconds and volume in cubic metres. For a small domestic room it lands around 150 to 250 hertz; for a concert hall it is around 20 hertz, which is to say below hearing.

That single number explains the difference between the two cases. A concert hall is statistical everywhere it matters. A small room is modal across the whole bass register. No amount of equipment changes it — it is a function of volume, and the only fixes are bass trapping, which is bulky, and moving, which works.

What a node means

The nodes are the striking part, and they are easy to verify.

Every mode has positions where the pressure is always zero. A listener at a node hears that frequency at greatly reduced level, however loud it is played. Move a metre and it returns.

For the lowest mode along a room’s length, the node is at the centre and the pressure maxima are at both ends. So a bass note at that frequency is loudest against the walls and quietest in the middle of the room — which is the opposite of what most people arranging a room would guess, and which is why bass sounds boomy in corners.

A small room's lowest modesThe first few axial standing waves of a room, drawn in plan, with the frequency of every mode below 160 hertz listed underneath. The low modes are far apart in frequency, so some bass notes are loud in one corner and absent in another. The sound buttons play these two octaves above their real pitch, because a room's lowest modes are below what most speakers reproduce.mode 1 · 26.4 Hzloudest at both endsmode 2 · 52.8 Hz2 quiet linesmode 3 · 79.2 Hz3 quiet lines6.5 × 5 × 3 m — every mode below 160 Hz:263443536372798798114126136154gaps here are notes the room does not support
Fig. 2 A larger room. The same calculation with different dimensions gives modes at different frequencies and more of them in the same span, which is the entire advantage of size — the gaps between supported frequencies close up.

The practical consequences are all of the same kind. A double bass placed in a corner sounds different from one in the middle. A recording microphone at a node misses a note that a microphone a metre away captures. Two listeners at different points in a small room are not hearing the same balance, and neither is wrong.

Reverberation, which is the statistical part

Above the Schroeder frequency the modes are too dense to consider individually, and what matters is how long the energy takes to die away.

Reverberation time, T60T_{60}, is the time for sound to decay by 60 decibels — a factor of a million in energy. Sabine derived the relationship in the 1890s, at Harvard, in the first genuinely quantitative work in architectural acoustics:

T600.161VA,T_{60} \approx 0.161 \frac{V}{A},

with VV the volume in cubic metres and AA the total absorption in square metres of open window equivalent.

The formula says something simple and consequential: reverberation is volume divided by absorption. A large room with hard surfaces rings; a small room with soft ones does not. Nothing about shape appears in it, which is both its power and its limitation.

Typical values, and the repertoire they belong to: a recording studio, 0.3 seconds. A domestic room, 0.5. A theatre, 1.0. A concert hall, 1.8 to 2.2. A large stone cathedral, 6 to 10.

Rooms shaped the music written in them

The last figures are large enough to have determined how music was written, and this is the most interesting thing in the subject.

A cathedral with an eight-second reverberation time cannot support fast harmonic change. By the time a chord has arrived, the previous three are still sounding, and any progression faster than a few seconds a chord turns into a wash. What such a space does support is slow-moving polyphony over long-held notes, where the accumulated sound is consonant with what is arriving.

That is exactly what was written for those spaces. Renaissance sacred polyphony, organum, Gregorian chant — slow harmonic rhythm, stepwise motion, consonant sonorities, and no fast bass lines — no room for a progression to be heard as a path. The style is usually explained theologically. The acoustics explain it at least as well, and the two are not in competition.

Move to an eighteenth-century room with a reverberation time near one second and fast harmonic rhythm becomes possible. Classical style has chord changes every beat, rapid bass figuration, and a texture that would be unintelligible in a cathedral. The music tracked the rooms.

The same argument runs forward. Amplified popular music is mixed for reproduction in small dead rooms and in headphones, where reverberation is added deliberately and controlled. A style built on a tight low-frequency groove requires a space that does not smear it — which no large stone building can provide, and which a recording can.

Three envelopesHow loudness changes over the life of a note, for a plucked, a bowed and a struck instrument. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.00.511.5200.20.40.60.81secondsamplitudeplucked — no sustainbowedstruck — no sustainkey released
Fig. 3 Three envelopes. A room adds its own decay on top of whatever the instrument does, and in a reverberant space the room’s decay dominates: the note’s own release becomes irrelevant because the room outlasts it.

The room is inside the instrument too

The same physics operates at smaller scales, inside the instruments themselves, and it is where instrument quality actually lives.

A violin body has modes. Its air cavity has a Helmholtz resonance around 270 hertz; its top and back plates have their own bending modes; the whole assembly has a response curve with peaks and troughs that varies enormously between instruments. That response is what makes one violin sound different from another built to the same dimensions.

The consequence is a wolf note — the string-player’s kind, unrelated to the tuning kind — where a played pitch coincides with a strong body resonance and the two fight, producing a stuttering, unstable tone. It is a property of that instrument at that note, it is unrelated to the wolf in tuning, and players work around it.

A guitar body, a piano soundboard, a drum shell, the human vocal tract: all resonators with modes, all imposing a fixed pattern of emphasis on whatever passes through. The vocal tract case has a name and a large literature — the resonances are formants, they sit at fixed frequencies regardless of the pitch being sung, and which frequencies they sit at is what makes a vowel a particular vowel.

Four spectra of the same noteThe amplitude of each partial for four timbres at the same pitch. These are the exact lists the sound buttons on this site synthesise from, so the picture and the sound are the same data.1pureone partial, nothing else12345678stringall partials, falling12345678clarineteven partials nearly absent123456789bellodd partials onlyamplitude
Fig. 4 Four partial lists. A body resonance boosts whichever partials fall near it, so the same excitation through two different bodies produces two different spectra — which is what distinguishes instruments of the same family.

Early reflections, which matter more than the number

Reverberation time is one number and it is not what distinguishes a good hall from a bad one. What does is the first eighty milliseconds.

Sound arriving directly is followed by a handful of discrete early reflections from the nearest surfaces, and only then by the dense statistical tail. Reflections arriving within about 80 milliseconds of the direct sound are not heard as separate events — the ear fuses them with the original, an effect known as the precedence effect — and they add to its perceived loudness and body without smearing it.

Reflections arriving later are heard as reverberation, and later still as echo.

So a hall’s job is to deliver strong early reflections and a smooth tail. The classic shoebox halls — the Musikverein, the Concertgebouw, Boston Symphony Hall — do it with narrow parallel side walls that put a lateral reflection at the listener within the fusion window. Lateral reflections specifically, because arriving from the side rather than overhead, they differ between the two ears and produce the sense of being surrounded that listeners describe as spaciousness.

Fan-shaped halls, which look more democratic on a plan, spread the side walls too far for early lateral reflections and are consistently rated worse. That relationship was not understood until the 1960s and a number of expensive mid-century halls were built without it.

Three envelopesHow loudness changes over the life of a note, for a plucked, a bowed and a struck instrument. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.00.511.5200.20.40.60.81secondsamplitudeplucked — no sustainbowedstruck — no sustainkey released
Fig. 5 Three envelopes. A room adds its own decay after the note ends, and in a hall the early part of that decay arrives soon enough to be fused with the note itself — which is why the same instrument sounds larger in a good hall and not merely more echoing.

What performers do about it

A room is not a fixed condition that musicians endure. It is a parameter they play into, and the adjustments are systematic.

Tempo. Musicians slow down in reverberant spaces, because fast passages smear. The effect is measurable and largely unconscious, and organists in large churches play markedly slower than the same music in a studio.

Articulation. In a dry room, notes must be joined by the player to sound connected. In a wet one, the room joins them, and the player must separate them — which is why a passage that sounds legato in a cathedral needs shortening to survive the recording of it.

Dynamics. A reverberant room builds up sound energy, so a loud passage in a cathedral goes further than the same passage in a studio, and the useful dynamic range at the top compresses.

Balance. Ensemble balance is room-dependent, because the room does not treat all frequencies equally — it is nearly always more reverberant in the bass, so low parts are boosted relative to high ones and must be played back.

Organ registration is the extreme case: an organist choosing stops is choosing a spectrum to suit a specific building’s response, and a registration that works in one church is wrong in another. The instrument was designed for the room it stands in, and moving one is not possible.

Absorption is frequency-dependent

Sabine’s formula uses a single absorption figure, and no real material has one.

Soft porous materials — curtains, carpet, people — absorb high frequencies efficiently and low ones hardly at all, because absorption requires air movement and a thin layer sits in a region where low-frequency air movement is small. A room treated entirely with foam becomes dead in the treble and remains reverberant in the bass, which is a common and unpleasant result.

Low frequencies need either great thickness or a different mechanism: panel absorbers that flex, or Helmholtz resonators tuned to a specific band. Both are bulky, both are frequency-selective, and neither is something that can be hung on a wall as an afterthought.

The consequence is that nearly every room is more reverberant in the bass than in the treble, and an audience makes it worse — people are excellent high-frequency absorbers and poor low-frequency ones, so a full hall is drier and proportionally bassier than an empty one. Rehearsing in an empty hall and performing in a full one is a change of instrument, and experienced players adjust for it.

Four spectra of the same noteThe amplitude of each partial for four timbres at the same pitch. These are the exact lists the sound buttons on this site synthesise from, so the picture and the sound are the same data.1pureone partial, nothing else12345678stringall partials, falling12345678clarineteven partials nearly absent123456789bellodd partials onlyamplitude
Fig. 6 Four partial lists. A room that absorbs unevenly across frequency reshapes every one of these before it reaches a listener — so the timbre that arrives is the instrument’s spectrum multiplied by the room’s response.

Where the model stops

Rectangular rooms with rigid walls. The modal formula assumes a box with perfectly reflecting boundaries. Real rooms have furniture, doors, absorbent surfaces and non-parallel walls, all of which shift and damp the modes.

No absorption in the modal calculation. Real modes have width, because they are damped. A damped mode is a bump rather than a spike, and the gaps between them are partly filled.

Sabine’s formula ignores shape. It gives a good answer for a diffuse field and nothing at all about early reflections, which is where most of a hall’s perceived quality lives. Two halls with identical reverberation times can be excellent and unusable.

A single number for a whole room. Reverberation time varies with frequency — rooms are almost always more reverberant in the bass — and one number hides that.

The figures draw plan views. Rooms are three-dimensional and the tangential and oblique modes are not shown, though they are in the frequency list. The pictures show the axial modes because those are the strong ones, which is a defensible simplification and is a simplification.

Measuring a room, and carrying it away

A room’s effect on sound is, to a good approximation, a single linear operation, and that fact has a consequence that changed recorded music.

If a room is linear and time-invariant — which it is, at ordinary levels — then everything it does is captured by its impulse response: what comes back when a single infinitely short click is played into it. Every other sound’s behaviour follows by convolution with that response.

So a room can be measured once, as a recording of a click or of a swept sine, and applied afterwards to anything. Convolution reverb does exactly this: the impulse responses of concert halls, cathedrals and stairwells are recorded, sold, and applied to material recorded in a dead studio.

That means a room is now a portable parameter rather than a fixed condition, and a mix can be assembled from sources that were never in the same building. It also means that the room’s contribution can be separated from the instrument’s, which was impossible when the two arrived together and inseparably.

What convolution does not capture is anything nonlinear or time-varying: an audience absorbing sound, air movement, and the fact that a real player adjusts to the room while a recording cannot. The captured hall is the hall as it was, empty, with nobody playing into it.

A small room's lowest modesThe first few axial standing waves of a room, drawn in plan, with the frequency of every mode below 160 hertz listed underneath. The low modes are far apart in frequency, so some bass notes are loud in one corner and absent in another. The sound buttons play these two octaves above their real pitch, because a room's lowest modes are below what most speakers reproduce.mode 1 · 53.6 Hzloudest at both endsmode 2 · 107.2 Hz2 quiet linesmode 3 · 160.8 Hz3 quiet lines3.2 × 2.6 × 2.4 m — every mode below 160 Hz:54668597107126142153gaps here are notes the room does not support
Fig. 7 A very small room. The modes are further apart and there are fewer of them below any given frequency, which is why a domestic space is the worst acoustic environment most listening happens in — and why a measured impulse response of a good hall is worth having.

The ladder from here

Later rungs: modal density and the Schroeder frequency derived. Sabine and Eyring formulas, and where each applies. Early reflections and why they matter more than reverberation time. Diffusion and absorption. Bass trapping. Concert hall design and the shoebox. Formants and vowel identity. Instrument body resonances. The wolf note on a cello. Convolution reverb, and capturing a room as a measurement. And the history of halls, where Sabine’s work at Harvard turned architectural acoustics from folklore into engineering.

Sabine was given an unusable lecture room to fix and no theory to fix it with. He borrowed seat cushions from a nearby theatre, moved them in and out at night, and timed the decay of an organ pipe with a stopwatch until the relationship between absorption and reverberation fell out. The formula still carries his name and is still what a first estimate uses.