Instruments and their design

The bass a small loudspeaker does not make

A three-inch cone at its excursion limit produces sixty-six decibels at forty hertz, which the ear converts to seventeen phons — barely above nothing. The note is heard anyway, because its harmonics are radiated and its fundamental is supplied by the listener. Computing what arrives turns the residue from a curiosity into a design decision, and finds that the fundamental of a low note is not the loudest part of it on any system a listener is likely to own.

Assumes: The note that is not there · Which harmonics carry the pitch

The first rung of this anchor opens with an observation and does not compute it: a laptop speaker reproduces essentially nothing below two hundred hertz, a bass guitar’s bottom string is at forty-one, and the bass line comes through at the right pitch with nobody noticing.

The observation is correct. What it leaves out is how far from nothing “essentially nothing” is, and the answer is a long way — far enough that the residue stops being a curiosity about perception and becomes the load-bearing assumption in a piece of consumer engineering that ships in billions of units a year.

The most a 6 cm cone can make of a low note. The maximum sound pressure level at one metre from a circular radiator of effective radius 3.2 cm, moving 1.5 mm at its limit, in a system resonating at 250 Hz. Below resonance the cone is already at that limit and the pressure a piston makes goes as the square of frequency, so the curve falls at twelve decibels an octave: 66 dB at 40 Hz, 78 dB at 80 Hz, 94 dB at 200 Hz. The 40 Hz figure is 28 decibels below the 200 Hz one, and that gap is arithmetic about a radius and a displacement rather than a property of any particular loudspeaker. The dots are the harmonics of a 41.2 Hz note with a one-over-n source spectrum; the loudest of them is the 6th.
Fig. 1 The most a circular radiator 6 cm across can produce at one metre, against frequency, with its cone moving 1.5 mm at the limit of its suspension and the system resonating at 250 Hz. Below resonance the cone is already travelling as far as it can and the pressure a piston makes goes as the square of frequency, so the curve falls at twelve decibels an octave: 94 dB at 200 Hz, 78 at 80, and 66 at 40. The dashed line is the threshold of hearing. The gap between the two at the bottom of the range is what this essay is about.

The argument

Two independent curves conspire against a low note and a third finishes the job, and all three are computable.

The source falls as one over the harmonic number for any ordinary plucked or bowed spectrum. The radiator rises as the square of frequency at a fixed cone displacement, which more than cancels that fall. And the ear is roughly fifty decibels less sensitive at forty hertz than at two hundred and fifty. Multiply the three and the loudest component of a forty-hertz note is nowhere near forty hertz.

The surprising part is not that this is true through a three-inch driver. It is that it is true through anything at all. The residue is not a repair for bad loudspeakers; it is how low bass is heard on any system, and the small speaker only widens a gap that was already there.

The most a 18 cm cone can make of a low note. The maximum sound pressure level at one metre from a circular radiator of effective radius 9.0 cm, moving 6.0 mm at its limit, in a system resonating at 45 Hz. Below resonance the cone is already at that limit and the pressure a piston makes goes as the square of frequency, so the curve falls at twelve decibels an octave: 96 dB at 40 Hz, 98 dB at 80 Hz, 98 dB at 200 Hz. The 40 Hz figure is 2 decibels below the 200 Hz one, and that gap is arithmetic about a radius and a displacement rather than a property of any particular loudspeaker. The dots are the harmonics of a 41.2 Hz note with a one-over-n source spectrum; the loudest of them is the 1st.
Fig. 2 The same arithmetic on a cone 18 cm across, travelling 6 mm, in a cabinet resonating at 45 Hz — which is what a loudspeaker built to reproduce bass is. It makes 96 decibels at forty hertz where the small one makes 66, and the whole gap comes from a radius and a travel: eight times the area and four times the excursion is thirty decibels. Nothing about the physics is different and nothing about the ear is different. The small driver is not badly designed; it is small, and the square law is not negotiable.

Which computation produced the number

On axis in a baffle the far-field pressure from a piston of area SS moving with peak displacement xx is

p=2πρf2Sxr,|p| = \frac{2\pi \rho f^2 S x}{r},

which is the whole of the first curve. It falls as f2f^2, so every octave downward costs twelve decibels before anything about the amplifier or the enclosure has been mentioned. Below the system’s resonance the displacement is already at its limit and the level goes with it; above resonance the displacement needed for a given output falls as fast as the pressure would rise, the limit stops binding, and the maximum flattens.

For a radiator of effective radius 3.2 cm at a 1.5 mm limit, one metre away, that gives 66 decibels at forty hertz. None of those three numbers describes a product: they are the dimensions of a cone of a stated size at a stated limit, and every figure here prints all three.

The second computation is ISO 226, the published equal-loudness contours, evaluated at each harmonic’s own frequency rather than only at the standard’s tabulated third-octaves. Sixty-six decibels at forty hertz is seventeen phons. The same sixty-six decibels at a kilohertz would be sixty-six phons.

Why the conversion is so brutal down there is the equal-loudness contours: every point on one of them is a level that sounds equally loud, and they crowd together at the bottom of the spectrum, so a decibel lost at 40 hertz costs several times what a decibel lost at a kilohertz does.

What actually arrives, harmonic by harmonic

How loud each harmonic of a 41.2 Hz note is. Loudness in phons of the first 8 harmonics of a 41.2 Hz note with a one-over-n source spectrum, computed through ISO 226. Through a system flat to 20 Hz the fundamental is 54 phon and the loudest harmonic is the 6th at 68. Through the small radiator the fundamental is 17 phon and the loudest harmonic is the 6th at 78. So the fundamental is not the loudest part of a low bass note on either — the difference is 14 phon on the unrestricted system and 60 on the small one. Harmonics 1, 2, 3 are the only ones the ear resolves, which is a separate question and is marked underneath.
Fig. 3 The first eight harmonics of a 41.2 Hz note, in phons, through a system with no low-frequency limit and through the small radiator. On the unrestricted system the fundamental is 54 phon and the sixth harmonic is 68 — the fundamental is fourteen phons down on the loudest thing in its own note. Through the small radiator the fundamental is 17 phon and the sixth is 78, a gap of sixty. The row underneath marks which harmonics the ear can separate at this pitch, which is a different question and comes out at three.
How loud each harmonic of a 82.4 Hz note is. Loudness in phons of the first 8 harmonics of a 82.4 Hz note with a one-over-n source spectrum, computed through ISO 226. Through a system flat to 20 Hz the fundamental is 73 phon and the loudest harmonic is the 2nd at 75. Through the small radiator the fundamental is 56 phon and the loudest harmonic is the 3rd at 85. So the fundamental is not the loudest part of a low bass note on either — the difference is 3 phon on the unrestricted system and 29 on the small one. Harmonics 1, 2, 3, 4, 5, 6 are the only ones the ear resolves, which is a separate question and is marked underneath.
Fig. 4 The same computation an octave higher, on the E below the bass clef. Every gap narrows: on the unrestricted system the fundamental is now 3 phon below the loudest harmonic rather than 14, and through the small radiator 29 rather than 60. Six harmonics are resolved here against three an octave down. So the whole effect this essay is about halves with every octave upward, and it is spent by the middle of the cello’s range — which is why the residue is a fact about bass lines and not about music.

The fourteen-phon figure on the unrestricted system is the result this rung was not slated to find, and it is the more interesting one. A forty-one-hertz note played through a system that reproduces forty-one hertz perfectly still has its loudness carried by its fourth to sixth harmonics. The fundamental is present, is audible, contributes body — and is not what the ear is measuring the note by.

That reframes the whole anchor. The missing fundamental is normally introduced as what happens when something goes wrong: a channel is band-limited, a partial is deleted, a bell has no prime. The arithmetic says that in the bottom octave and a half of the musical range the fundamental was never carrying the note in the first place, and the residue is the ordinary mechanism rather than the exceptional one.

How much of that depends on the source falling as one over n

The caveat below records that a real bass spectrum is not exactly one over n, and the two halves of the finding turn out to depend on it completely differently.

Recomputing with the source falling as n to the power −α, for α from 0 to 2, through the small radiator:

source loudest harmonic fundamental below it
flat 8th 80 phon
1/n^0.5 7th 70
1/n (as drawn) 6th 60
1/n^1.5 6th 51
1/n² 6th 41

Through a small driver nothing is at risk. The loudest component is the sixth to eighth harmonic across the whole plausible range of source spectra, and the fundamental is between forty and eighty phons below it. A margin of forty phons does not care what the source was doing.

The full-range case is where the assumption is load-bearing. Recomputed on the eighteen-centimetre driver of the second figure:

source loudest harmonic fundamental below it
1/n 4th 11 phon
1/n^1.5 2nd 5
1/n² 2nd 1

At a source falling as one over n squared, the fundamental is one phon below the loudest thing in its own note — which is to say, tied.

Where the claim flips, and it is close

Solving for the exponent at which the fundamental takes the lead: it happens at 1/n^2.07.

That is uncomfortably near the edge of what a real string does. A plucked string’s spectrum falls roughly as one over n at low harmonic numbers and steepens above the plucking point’s first null, and a bass with a soft attack, a heavy string or a tone control turned down can present something close to one over n squared into the low band. So the claim that the fundamental is never the loudest part of a low note holds for every source spectrum in the plausible range and stops holding just outside it, on a full-range system.

That is worth separating from the essay’s main argument rather than blurred into it, because the two claims now have very different standing. On a small device the residue is the mechanism, by a margin nothing could close. On a full-range system the fundamental is not the loudest component for ordinary sources and is within a phon or two of being so for dull ones, which is a much weaker statement than “never, on any system at all” — and it is the statement the arithmetic supports.

The reason the two behave so differently is the resonance. Below it the driver’s displacement limit binds and radiation rises as frequency squared, which multiplies the harmonics up by n² against the source’s fall; above it the limit stops binding and the source’s own slope takes over unopposed. A small driver resonating at 250 hertz has the whole first six harmonics of a low E inside the rising region. A large one resonating at 45 has none of them.

The three curves also explain why the loudest harmonic lands where it does. Radiation rises as f2f^2 and the source falls as 1/n1/n, so the radiated level of harmonic nn rises as nn — six decibels per doubling — until the system resonance, above which radiation is flat and the source’s fall takes over. The turning point is therefore at the resonance, which for this driver is the sixth harmonic of a forty-one-hertz note. On a driver resonating at 150 hertz it would be the fourth. In every case it lands close to the dominance region, and it lands there for reasons that have nothing to do with hearing.

What the listener has to work with

The most a 6 cm cone can make of a low note. The maximum sound pressure level at one metre from a circular radiator of effective radius 3.2 cm, moving 1.5 mm at its limit, in a system resonating at 250 Hz. Below resonance the cone is already at that limit and the pressure a piston makes goes as the square of frequency, so the curve falls at twelve decibels an octave: 66 dB at 40 Hz, 78 dB at 80 Hz, 94 dB at 200 Hz. The 40 Hz figure is 28 decibels below the 200 Hz one, and that gap is arithmetic about a radius and a displacement rather than a property of any particular loudspeaker. The dots are the harmonics of a 82.4 Hz note with a one-over-n source spectrum; the loudest of them is the 3rd.
Fig. 5 The same driver an octave up, which is the comparison that shows the slope rather than the number. Below its resonance a cone is already at its excursion limit and the pressure a piston makes goes as the square of frequency, so every octave down costs twelve decibels — and that is why the octave above a bass guitar’s low E is reproduced adequately by the very device that cannot touch the E itself. Nothing about the enhancement changes this curve; what the enhancement does is move the evidence for the note into the part of it that works.

So what reaches a listener from a bass line on a small device is a set of unresolved partials whose sum repeats forty-one times a second. That is precisely the object the previous rung built in a laboratory: no template to match, and a repeat rate for something to read.

So what reaches a listener from a bass line on a small device is a set of unresolved partials whose sum repeats forty-one times a second — the autocorrelation of harmonics two to five alone returns to one at 24.3 milliseconds, the period of 41.2 hertz, and that repeat is the entire evidence for the note. The price is determinacy rather than accuracy. Four consecutive harmonics from the fifth up are consistent with two fundamentals inside thirty cents: 41.2 hertz with nothing unaccounted for, and 20.6 which fits every partial exactly and predicts three slots that nothing occupies. With eight harmonics the octave candidate carries seven empty slots and is easy to reject; with two it carries one. A narrower passband does not shift the pitch, it makes the pitch less determined.

The trick, and it is not a repair

A device that cannot radiate a fundamental can synthesise its harmonics instead. Generate the second to fifth harmonics of whatever is in the low band, put them where the driver works, and let the listener’s own machinery supply the note. The category is called psychoacoustic bass enhancement, it is in phones, laptops, televisions and car doors, and it is a deliberate application of everything on this anchor.

It is worth being precise about what it substitutes. For an ordinary plucked bass the harmonics are already there, and the enhancement is adding little. For a synthesised bass or a kick drum whose low band is close to a pure tone, they are not there at all: a forty-hertz sine through the driver of the hero figure is seventeen phons and inaudible in any real listening situation, and there is nothing above it to imply anything. The enhancement manufactures the evidence.

Two consequences follow from the arithmetic and one of them is the opposite of what is usually claimed.

The substitution is not paid for in masking, which is where a cost would be expected. To be as loud as its own harmonics, a forty-hertz fundamental needs 104 decibels where the harmonics need 82 — because the ear’s contour is so steep down there. Masking spreads upward much further than downward and its upward reach grows with level, so the very loud low tone buries slightly more of the lower midrange than the substitute does: at 300 hertz the real fundamental raises the threshold to 82 decibels and the substitute to 78. Four decibels is not much either way, and the direction is the one nobody would guess.

How loud each harmonic of a 41.2 Hz note is. Loudness in phons of the first 8 harmonics of a 41.2 Hz note with a one-over-n source spectrum, computed through ISO 226. Through a system flat to 20 Hz the fundamental is 54 phon and the loudest harmonic is the 6th at 68. Through the small radiator the fundamental is 28 phon and the loudest harmonic is the 6th at 65. So the fundamental is not the loudest part of a low bass note on either — the difference is 14 phon on the unrestricted system and 38 on the small one. Harmonics 1, 2, 3 are the only ones the ear resolves, which is a separate question and is marked underneath.
Fig. 6 And the smallest driver of the three, which is where the trick stops being optional. Through a system flat to 20 hertz the fundamental of a 41-hertz note is 54 phons and the loudest harmonic is the sixth at 68; through a phone-sized radiator the fundamental is 28 phons — below the threshold in a room with any noise in it at all — while the harmonics above 200 hertz arrive almost untouched. The enhancement is not adding a fundamental. It is adding harmonics to a note whose own harmonics the driver was already delivering, and the honest description of what it buys is a stronger residue rather than a restored bass.

What it does cost is determinacy and headroom. The octave ambiguity above is real and it grows as the passband narrows. And the substituted harmonics have to be loud enough to be heard against the rest of the mix, in a band where the driver is also reproducing everything else, so a device that enhances aggressively is spending its excursion and its amplifier on synthetic content.

And the trick is worth more the quieter the listening is.

What 20 decibels off the volume takes away. The perceived loudness lost, in phons, when 20 dB is removed from a tone that was 80 phons loud, computed frequency by frequency from ISO 226:2003. The line is not flat, so turning a piece of music down does not turn all of it down equally: the bottom of the spectrum loses about 42 phons where the middle loses 20.
Fig. 7 What twenty decibels off the volume takes away, frequency by frequency, from a tone that was 80 phons loud. The line is not flat: the bottom of the spectrum loses 42 phons where the middle loses 20, because the equal-loudness contours crowd together down there. So a bass line played quietly loses its fundamental twice as fast as it loses its harmonics, and a device whose residue is carried by harmonics is losing the cheaper half. The trick is not a fixed benefit; it is one that grows as the volume comes down.

And the substituted harmonics are not resolved from each other either. Harmonics two to five of a forty-one-hertz note sit at 82, 124, 165 and 206 hertz, forty-one apart, against an analysis bandwidth that runs from 34 to 47 hertz across that span. Only the lowest two of them have a filter to themselves. The rest arrive together, and what a filter delivers from a pair of partials inside it is the fluctuation that this site computes as roughness — here at forty-one a second, which is fast enough to be a buzz rather than a beat. That is the honest form of the complaint that a small speaker’s bass sounds synthetic. The added partials are not heard as separate tones; they are heard as one rough thing whose roughness rate is the note.

The claim that the restored bass has a timbre the real note does not is the one the computation does not support in the form it is usually made. Through the same driver, an ordinary bass note’s harmonics arrive at very nearly the levels the enhancement supplies, because the driver has already thrown the fundamental away. The timbre difference is between the small device and a full-range system, and it belongs to the driver rather than to the enhancement. What the enhancement changes is which sources get the treatment: a note that had harmonics keeps its own, and a note that had none is given a set it never had.

What the picture cannot show

The room, the enclosure and the placement, all of which move the curve by more than the differences argued about here. A driver in a corner gains several decibels at the bottom for free; the same driver in free air on a small box loses several more than the figures show.

Distortion, which at the excursion limit is severe. The hero figure draws the maximum a cone can reach and says nothing about what it sounds like there, and a driver run to its limit at forty hertz generates harmonics of its own — which, by an irony this essay cannot resolve, are the same harmonics the enhancement would have added deliberately.

Directivity is not the problem, and it is worth showing that it is not.

How directional a source 6 cm across becomes. Directivity index against frequency for a circular radiator of radius 3 cm. It is flat and near zero while the radiator is small compared with the wavelength, and rises at six decibels per octave once it is not. The crossover is at ka = 1, which for this radius is 1706 Hz — below it the instrument fills the room and above it it points.
Fig. 8 The same radiator’s directivity index. It is omnidirectional until ka=1ka = 1, which for this radius is 1,706 Hz, and only beams above that — so at forty hertz the driver radiates equally in every direction and loses nothing to beaming at all. The directionality argument that governs an instrument’s spectrum in a hall has no purchase on this case. What limits the bass is displacement, and nothing else.

The spectrum is drawn as one over n and a real bass is not. A plucked string’s spectrum depends on where it was plucked, an electric bass adds a pickup’s own filtering on top, and a recorded one has usually been compressed and equalised before it reaches any loudspeaker. The section above sweeps the exponent and finds the small-driver result untouched across the whole range and the full-range result marginal at the steep end — so the one-over-n assumption is doing no work in half the essay and a great deal in the other half.

And loudness is not pitch. Every number here is about level and audibility. A note whose fundamental is seventeen phons and whose sixth harmonic is seventy-eight has the right pitch and the wrong weight, and the difference between a bass line that is audible and one that is felt is not in any of these figures.

The spectrum sweep varies one exponent and holds the rest of the model still. A real source departs from a power law in ways an exponent cannot express — a plucking point puts a null in the series, a pickup adds a resonance, a compressor moves the balance with the envelope — and the sweep says only that the conclusion is insensitive to the slope. What it establishes is where the sensitivity lives: not in the small-driver case, which is decided by the radiator, and squarely in the full-range case, which is decided by the source.

Whose recordings, and when

The engineering decision is old and the physics did not change.

The telephone channel of the 1920s excluded the fundamental of every adult voice on cost grounds, and the engineers had the psychoacoustic literature. Small transistor radios of the 1950s and 60s were voiced with a deliberate lower-midrange lift for exactly the reason the loudness figure gives. Recording practice in popular music since roughly the 1960s has routinely added harmonic distortion to bass parts — described in the trade as making the bass “translate” — which is the enhancement above, applied at the mix rather than at the device, and applied because most listeners were on small speakers.

The orchestral parallel runs the other way and was arrived at by ear. Doubling a bass line at the octave with cellos or bassoons, standard from the eighteenth century onward, puts a partial in the resolved range that belongs to the same series as the double bass’s note. It strengthens the fit without adding a pitch, and it is the same operation as an organ’s quint rank and the same operation as a bass-enhancement stage, arrived at three times independently by people who could not have agreed about the mechanism.

Where the ladder goes next

This anchor now has eight rungs and they divide in a way the table did not predict. The first four establish that pitch is inferred from a pattern — in a laboratory, in a bell, in a kettledrum. The last four price the inference, and the price turns out to be paid in the same currency every time: how many partials arrive separately, and what is left when none of them do.

Three things the last four rungs found that the first four did not have. The residue’s ceiling is not a resolvability ceiling, because resolvability improves with frequency and never fails at the top. The shifted-residue experiment does not separate the two surviving accounts, because a correlation predicts the shift as accurately as a template does. And the fundamental of a low note is not the loudest part of it on any system, so the residue is the normal mechanism in the bass rather than an artefact of bad equipment.

What the ladder has not done is decide between a template and a correlation, and it now looks as though the question is badly posed: the two agree wherever partials are resolved and only one of them applies where they are not. The rungs still open are the ones that would test the division rather than the models — pitch strength as a measured quantity across the resolved-to-unresolved boundary, what happens to a residue when a room adds a reflection at a fraction of the period, and whether a listener’s octave errors track the empty-slot count these figures keep printing.

Part 8 of 9

One essay in the series on missing fundamental. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Equal-loudness contourLoudnessMaskingMissing fundamentalRadiation efficiencyResidue pitchResolvabilitySpectral balance