The quietest thing audible, and why the volume knob is a tone control
A microphone answers one question about a sound: how much the air moved. An ear answers a different one, and the difference is not a refinement. Two tones that move the air by identical amounts are not equally loud, they are not close to equally loud, and how far apart they are depends on which two tones and on how loud they both are.
Read the 40-phon curve across. At 1 kHz it sits at 40 dB, by definition — the phon is defined as the level in decibels of an equally loud tone at 1 kHz, which is why that one point is exact and everything else on the curve is a measurement. At 50 Hz the same curve sits near 92 dB. A tone at the bottom of the bass has to move the air about a hundred and seventy times as hard to arrive at the same place.
That is a large number and it is not the surprising one. The surprising one is that the curves converge as they descend.
What non-parallel means
If the contours were parallel, loudness would be a relabelling of level: every frequency would have a fixed handicap, the handicap would be the same at every volume, and once the handicap was applied everything else about a spectrum would follow the arithmetic. There would be no reason to draw more than one curve.
They are not parallel. The 80-phon and 20-phon contours are about 60 dB apart at 1 kHz and about 45 dB apart at 50 Hz, so the same reduction in level takes away different amounts of loudness at different frequencies.
The figure is drawing a small everyday fact and giving it a size. Music played quietly sounds thin, and it sounds thin because the bass has genuinely gone further away than the rest of it has. The loudness contour is doing the removing, not the loudspeaker.
Hi-fi amplifiers used to carry a switch labelled loudness which applied a bass boost at low volume, and it was not a marketing invention: it was an attempt to undo this figure. The reason the switch has disappeared is not that the physics changed. It is that the correction depends on the absolute level at the listener’s ear, which the amplifier does not know, so the compensation was right at exactly one setting of the volume control and wrong everywhere else.
Where the numbers come from, and what they cost
ISO 226:2003 is a table of three parameters at twenty-nine third-octave centre frequencies, plus one equation that turns them into a level. A tone’s level is computed from an exponent that describes how loudness grows at that frequency, a transfer term from the free field to the eardrum, and the threshold of hearing there. The exponent is the interesting one. At 1 kHz it is 0.250. At 20 Hz it is 0.532 — more than twice as large — which is the formal statement of the same fact the previous figure drew: loudness grows faster with level in the bass than in the middle, so it also collapses faster when the level comes down.
That the numbers are a standard rather than a derivation is worth being explicit about. Nothing in this essay is computed from first principles the way a comma is. The contours are the result of listening experiments — the 2003 revision was assembled from twelve studies in several countries — and the standard’s own stated range is 20 to 90 phons, from 20 Hz to 12.5 kHz. Outside that it is silent, and so is this essay.
There is one check that comes for free and is worth more than it costs. The phon is defined by the 1 kHz point, so the table has to reproduce that definition: the contour for N phons must pass through exactly N dB at 1 kHz. Evaluated here it does, to within 0.012 dB at 20, 40, 60 and 80 phons. A mistyped digit anywhere in the 1 kHz column would show up immediately, and the site’s gate asserts it.
The dip at three kilohertz is a tube
The threshold curve has a distinct minimum between about 2 and 5 kHz, bottoming out around 3,150 Hz where it dips a few decibels below zero — the ear is more sensitive there than at the frequency the whole scale is referenced to. Something is amplifying.
The something is a tube. The ear canal is a pipe about 25 mm long, closed at one end by the eardrum and open at the other, and a closed pipe resonates at the frequency whose quarter wavelength is its length. At 343 metres a second, a quarter wavelength of 25 mm puts the resonance at 3,430 Hz, and a real canal, which is neither straight nor uniform, broadens that into a gain of 10 to 15 dB over roughly an octave and a half. The most sensitive region of human hearing is a piece of plumbing.
This connects two things that look unrelated. A clarinet overblows a twelfth because it is a closed pipe and produces only odd partials; the ear canal is a closed pipe of the same kind, and its first resonance is computed the same way. The instrument and the listener are doing the same physics, one to make a sound and the other to receive it.
The consequence for music is that the region where the ear is cheapest to reach — where a given amount of acoustic power buys the most loudness — is precisely the region of the upper partials of a soprano, the top of a violin’s range, the presence peak of a trumpet, and the consonants of speech. It is not a coincidence that this is the band engineers reach for when they want something to cut through, and it is not a coincidence that it is the band that becomes unbearable first when it is overdone.
Equally loud is not twice as loud
The contours answer one question — which tones match — and are silent about the other one, which is how much louder one thing is than another. Matching and magnitude are separate measurements and they need separate units.
Ten violins are twice one violin. This is not a curiosity, it is a fact orchestration has been organised around for two hundred years, and it explains a shape that would otherwise look arbitrary: the string sections of an orchestra are enormous and the wind sections are not. Sixteen first violins against two flutes is not a balance of forces, it is a balance of loudnesses, and the reason the ratio has to be that extreme is that adding players buys loudness so slowly.
It also explains something about the other direction. A single violin dropping out of a section of sixteen changes the level by 0.28 dB, which is inaudible. A single flute dropping out of two changes it by 3 dB, which is not. The vulnerability of a section to one absent player is a function of its size and it is steeply non-linear.
What the picture cannot show
Three things, and each of them limits what the essay above is entitled to claim.
The contours are for steady pure tones presented from the front, and music is none of those things. A tone shorter than about 200 ms is quieter than the same tone held — the ear integrates energy over a window, so a short note needs more level to match a long one. That window is the same one that makes the first fifty milliseconds of a note decide what instrument it is, and none of it is anywhere in the drawing. Nor is the room: what arrives at an ear has already been through one, and a reverberant room adds several decibels to a sustained note and almost nothing to a short one.
They are averages over listeners, and the spread is large. Individual thresholds at a given frequency vary by 20 dB among people with clinically normal hearing, and the average is not a description of anybody. Age moves the top of the range down, sharply and permanently, and there is no reason to expect two people in a room to be on the same contour at all.
A spectrum is not a loudness. The figure treats each frequency separately, and the ear does not: several tones inside one critical band are not as loud as their levels suggest, because the band is the unit that is being analysed rather than the individual tone. Loudness models that get real signals right are built on band energies for exactly that reason, and none of them is this figure.
The curves are not parallel — they crowd together in the bass and spread apart in the treble — and that single fact is what this essay is about. A quiet sound is not a loud sound scaled down, because the shape of the ear’s sensitivity changes with level, so the quietest thing audible has a different spectrum from the same thing played loudly.
Two decibel measurements that are equal can therefore be two audibly different loudnesses. A cluster spread over three octaves is louder than the same total power packed into a minor third, and no meter reports the difference. That the narrow one is also the rougher of the two is a separate fact arriving from the same geometry, and the two should not be run together: roughness is about beating between partials and this is about how energy is totalled.
The line is not flat, so turning a piece of music down does not turn all of it down equally. The bass loses far more than the middle, which is why the quietest audible version of a passage is not the passage — it is the passage with its bottom removed, and the removal was done by the ear rather than by anyone’s hand.
A note is not at one point on this graph
Every previous figure has plotted single tones, and a musical note is not a single tone. It is a whole harmonic series, spread across two or three octaves, and its partials sit at different places on the contours.
The consequence is that turning a note down changes its timbre, not merely its size, and the change has a direction: the fundamental recedes faster than the partials above it. A bass note played quietly is heard mostly through its upper partials, which is the same mechanism as the missing fundamental arriving by a different route — there the fundamental is deleted from the signal, and here it is merely pushed below where the ear can weigh it properly. The pitch survives both — which is worth holding onto, because it means the ratio between two notes survives a change of volume even though almost nothing else about the sound does.
This is the practical reason a double bass and a bass guitar are heard at all on a small loudspeaker that reproduces nothing below 100 Hz. What is coming out is the second partial upward, and the ear supplies the rest.
The head start is not a number, it is a curve
The paragraph under the harmonic series just quoted a single figure for how much of a start the fourth partial has over the first, and doing that is the exact mistake this essay is about. The contours are not parallel, so a difference read between two of their points is a different difference on every contour:
| the note is | head start of the fourth partial over the first |
|---|---|
| 20 phons | 21.8 dB |
| 40 phons | 18.6 dB |
| 60 phons | 14.5 dB |
| 80 phons | 10.1 dB |
Eleven and a half decibels of swing across the useful range, and it runs the way it has to: loud, the ear is comparatively even-handed and the fundamental nearly catches up; quiet, the fundamental falls away and the note is carried by everything above it. The single number that used to stand here was not merely imprecise but the wrong kind of quantity, since a figure that varies by a factor of two over the range a listener actually uses cannot be quoted without the level it belongs to. Every difference read between two points on these contours has that property, which is what non-parallel means when it is put to work rather than described.
Which makes the small-loudspeaker claim above quantitative rather than rhetorical. Taking the same low A with the string spectrum, adding its partials in critical bands rather than one at a time — the correction a later rung of this ladder built, because partials inside one band add as intensity and not as loudness — and then deleting everything below 200 Hz:
At 90 decibels the missing bottom costs 23 per cent of the note’s loudness. At 50 decibels it costs 11 per cent. The cheap loudspeaker is less wrong at low volume, and the reason is the reason for everything else here: at low volume the fundamental had almost stopped contributing before the loudspeaker failed to reproduce it. A bass line played quietly through a full-range system and the same line played quietly through a telephone are much closer together than the same comparison made loud.
That is a small result and it is the only quantitative thing this rung can say about a note rather than a tone. Everything past it — how the eight partials add up to one loudness, what a chord weighs, what happens when the bands overlap — needs a model of adding-up that this rung does not have, and the last section says so.
Where this bites in practice, and whose practice
The claim that a mix sounds different at different volumes is a claim about a specific technology and a specific period, and it is worth saying which.
It matters for recorded music, from roughly the middle of the twentieth century onward, because a recording is a fixed spectrum played back at an unknown level. A balance decided in a studio at 85 dB is a different balance in a car at 70 and a different one again on headphones at 95. This is why studio monitoring levels are conventionally standardised — the cinema standard fixes a specific reference level for exactly this reason — and why engineers check a mix at more than one volume rather than at their favourite one.
It matters much less for live acoustic music, because there the listener’s level is set by the physics of the room and the instruments and by nothing else — the same reason an absolute pitch standard had to be legislated rather than discovered does not apply to loudness, which nobody has ever tried to standardise at the point of performance, and a quiet passage is quiet in the way the composer heard it be quiet. An orchestral pianissimo has a spectral balance the composer could hear; a pianissimo on a recording turned down has a spectral balance nobody chose.
Two cuts of different sizes give differently shaped answers. If the contours were parallel these two figures would be the same picture at two scales.
The model, named
Three separate models have been used above and they should not be run together.
ISO 226:2003 gives the equal-loudness contours: which tones match. It is a standard, assembled from listening data, valid over a stated range, and silent about everything else.
Stevens’s power law, from 1957, gives the sone scale: how much louder. It is a fit to magnitude-estimation experiments — listeners asked to say that one sound is twice another — and the exponent that produces the ten-decibel doubling is an average over a method that is known to be sensitive to how the question is asked.
The threshold of hearing is the bottom contour, and it is a statistical construct: the level at which a tone is detected on half the trials by young adults with no hearing damage, in silence that is quieter than most rooms ever get.
None of the three is a physical law and all three are reproducible to within a few decibels, which is the useful sense in which they are true.
Where the ladder goes next
This essay has treated the ear as a device that weighs each frequency separately and reports a number. It does not. The next rung is what happens when two sounds arrive at once: one of them can remove the other from the record entirely, the removal is asymmetric in frequency, and the asymmetry is the reason an orchestrator can predict which line will disappear.
Further along the same ladder sits the question this one deliberately left alone. The contours say what a tone’s loudness is; they do not say what a chord’s loudness is, or a room’s, or an orchestra’s. That is a question about how the ear adds things up, and adding things up is where the critical band stops being a fact about roughness and starts being a fact about volume.
Part 1 of 8
One essay in the series on loudness. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 26.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
DecibelEqual-loudness contourHearing thresholdLoudnessPhonSoneSpectral balance
- One voice over ninety players hearing threshold, spectral balance
- The dynamics are in the score already loudness, sone
- Who plays what and how loud is one question loudness, spectral balance