The ear hears the list, not the shape
Here is an experiment with a counterintuitive result.
Take a complex tone and shift the phase of each of its partials by a random amount. The waveform changes completely — a shape that looked like a sawtooth becomes something with no resemblance to it at all.
Play the two. They sound the same.
This is Ohm’s acoustical law, stated by Georg Ohm in 1843 and confirmed exhaustively since. The ear decomposes a sound into frequency components and is very largely insensitive to the phase relationships between them.
It has a large consequence: the waveform is the wrong picture of a sound. The right picture is the spectrum.
The right picture
A spectrum is a list: which frequencies are present and how strong each is. That list is what the ear recovers, what a synthesiser needs, and what makes one instrument sound different from another at the same pitch.
Read the four in the figure and the differences are legible without hearing them. The pure tone has one component and nothing else, which sounds thin and characterless. The string-like spectrum has all partials with amplitudes falling off gradually, which is rich and bright. The clarinet has the odd partials strong and the even ones nearly absent, which is the hollow, woody quality. The bell has odd partials only and widely spaced.
The clarinet case is the most satisfying because the spectrum is a direct consequence of the geometry. A pipe closed at one end and open at the other supports only odd modes, so the even partials have nowhere to live. That same fact makes the clarinet overblow to a twelfth — partial 3 — where a flute overblows to an octave, which is why the clarinet’s fingering system is more complicated than any other woodwind’s.
How the decomposition happens
The mechanism is physical and worth knowing, because it explains both the successes and the failures.
The cochlea is a coiled tube containing the basilar membrane, whose stiffness and width vary continuously along its length. A travelling wave in the fluid peaks at a position determined by its frequency: high frequencies near the entrance, low frequencies at the far end. Hair cells along the membrane report activity at their own position.
That is a mechanical frequency analyser. Each point on the membrane responds to a band of frequencies around its characteristic one, and the width of that band is the critical bandwidth — roughly a minor third in the middle of the range, proportionally much wider at the bottom.
Two consequences follow immediately, and both are load-bearing elsewhere on this site.
Components further apart than a critical band are resolved separately, and contribute independently. Components closer than a critical band excite overlapping regions, interact, and produce roughness. That single distinction explains the shape of every consonance curve there is.
And because the analysis is done by position, phase is discarded — the membrane’s response depends on where the energy is, not on when each component started.
Where phase deafness fails
Ohm’s law is an approximation and it has known limits, which is more interesting than a clean rule would be.
Phase matters within a critical band. Two partials close enough to interact do so in a way that depends on their relative phase, because they are physically summing before the analysis rather than after. For densely packed high partials this is audible.
Phase matters for transients. The attack of a note is a burst in which the relative timing of components is exactly what makes it sound like a pluck or a bow or a strike, and timing is phase.
Phase matters for localisation. The difference in arrival time between two ears is a phase difference, and it is the primary cue for locating a low-frequency sound.
So the honest version is: for steady tones, phase relationships among resolved components are largely inaudible. That is still a strong claim, still surprising, and still enough to make the spectrum the right representation for sustained sound.
The pitch that is not there
The best demonstration that hearing is inference rather than measurement is the missing fundamental.
Take a tone with partials at 400, 600, 800 and 1000 hertz. There is no energy at 200 hertz anywhere in the signal. The perceived pitch is 200 hertz.
The auditory system observes components spaced 200 hertz apart, infers the fundamental they would be partials of, and reports that as the pitch. Filter the low end out entirely and the pitch is unchanged. This is not a subtle laboratory effect — it is why a telephone, which transmits nothing below about 300 hertz, conveys a male speaking voice at its correct pitch, and why a small speaker reproduces a bass line it physically cannot produce — an inference of exactly the kind the harmonic series invites.
The consequence for a theory of hearing is substantial. Pitch is not a physical property being read off; it is a hypothesis about what would produce the observed pattern. And a hypothesis can be wrong, which is why ambiguous and paradoxical pitch phenomena exist at all.
Why an inharmonic instrument is a different animal
If pitch is inferred from a pattern of spacings, then a sound whose components have no regular spacing gives the inference nothing to work with.
A bell’s partials are not whole-number multiples of anything. There is no fundamental they would all be partials of, so the auditory system’s hypothesis is weak or multiple, and the resulting pitch is ambiguous — bells famously have a “strike note” that may not correspond to any partial present.
That is also why such instruments generate different scales. Consonance is a property of spectra, so an inharmonic spectrum has its low-roughness intervals somewhere other than the simple ratios — and a tradition built on gongs and metallophones tunes to those instead. Indonesian sléndro and pélog are the working example, and they fit their instruments in a way that no harmonic-series theory predicts.
Timbre is more than a spectrum
Having argued that the spectrum is the right picture, the argument needs its limit stated, because timbre is not a static list.
A steady spectrum played through a synthesiser sounds like an organ, whatever list is used. Real instruments are identifiable because their spectra change: the attack has a different balance from the sustain, upper partials decay faster than lower ones, and the whole thing varies with dynamic level.
Remove the first hundred milliseconds of a recorded piano note and listeners can no longer reliably identify it as a piano. The steady part carries much less identity than the transient does, and the transient is exactly what a spectrum plot omits.
So the honest summary is: the spectrum is the right picture of a steady sound, and a note is not steady. The full description is a spectrum that varies over time, which is what a spectrogram draws and what any convincing synthesis has to reproduce.
Formants, and why a vowel is a vowel
The most consequential application of spectral thinking is to the voice, and it introduces a distinction the harmonic series alone cannot make.
The vocal folds produce a buzz with a harmonic spectrum whose fundamental is the pitch being sung. That buzz then passes through the vocal tract — a tube of variable shape — which has resonances of its own. Those resonances boost whichever partials fall near them, and they are called formants.
The crucial property is that formants sit at frequencies determined by the shape of the tract, not by the pitch. Sing “ah” at any pitch and the first two formants stay near 700 and 1200 hertz; sing “ee” and they move to about 300 and 2300. The vowel is the formant pattern, and it is independent of the note.
So the voice carries two independent streams of information in one signal: pitch, in the spacing of the partials, and vowel, in which partials are emphasised. Speech uses the second and largely ignores the first; song uses both.
There is a limit case that singers live with. Above about 700 hertz, a soprano’s fundamental is higher than the first formant of most vowels, so there is no partial in the region that would identify the vowel. Vowels become genuinely indistinguishable at the top of the range, which is why high soprano text is unintelligible — a physical constraint, not a fault of diction.
Instruments have formants too
Once the idea is available it turns up everywhere, and it separates two things that are easy to confuse.
An instrument’s spectrum has two contributions: what the excitation produces, and what the body does to it. The excitation — a bowed string, a reed, a buzzing lip — generates a harmonic series whose frequencies move with the note. The body — a violin’s plates and air cavity, a guitar’s box, a piano’s soundboard — has fixed resonances that do not move.
So an instrument imposes a fixed pattern of emphasis on a moving series, and that fixed pattern is a large part of what makes one violin sound different from another built to the same measurements. Two instruments of the same design differ in their body resonances and in nothing else that matters.
It also explains the wolf note that string players complain about — a played pitch coinciding with a strong body resonance, where the string and the body exchange energy and the tone stutters. It is a property of one instrument at one note, it has nothing to do with the wolf in tuning, and the two share a name because both beat.
Whose music, and when
The analysis is physics and the uses of it are historical.
Helmholtz’s On the Sensations of Tone (1863) is where the spectral account of timbre is established, using resonators to isolate individual partials by ear. It is one of the great works of nineteenth-century science and it got the main claim right: instrument identity lives in the partial content.
It also got a substantial thing wrong, and the correction is instructive. Helmholtz treated timbre as essentially a matter of steady-state spectrum, and the twentieth century established that transients matter at least as much. The error was reasonable — his instruments could isolate a steady tone far more easily than a fifty-millisecond attack.
Additive synthesis, the technique of building a sound by summing sinusoids, is the direct application, and it is what every sound on this site uses. It is also, notably, a poor way to make convincing instrument sounds, precisely because getting the time-varying behaviour right requires an enormous number of parameters. The sounds here are deliberately synthetic for that reason: an honest sine stack rather than a bad imitation of an oboe.
Separating what arrives together
A listener in a room receives one pressure signal containing every source at once, and separating it is the basic problem of hearing rather than a refinement of it.
The auditory system solves it with a set of grouping rules, all of which operate on the spectrum over time.
Harmonicity. Components in whole-number ratios are grouped as one source. Mistune a single partial of a complex tone by a few per cent and it pops out as a separate whistle — a striking demonstration, and evidence that the grouping is active rather than passive.
Common onset. Components starting together are grouped. This is the strongest cue there is, and it is why an orchestra is heard as instruments rather than as a wall of sound.
Common modulation. Components sharing a vibrato move together and are grouped. This is why vibrato helps a singer stand out from an accompaniment: it labels a set of partials as belonging together.
Proximity and continuity. A sequence of tones close in pitch is heard as one melodic line; a sequence alternating between two registers splits into two, which is the illusion Bach exploits in unaccompanied string writing to imply two voices with one instrument.
Where the model stops
Steady state. Everything above concerns sustained tones. Transients are the largest omission and they matter most.
Linear analysis. The cochlea is nonlinear, generating combination tones at frequencies not present in the signal. At high levels these are easy to hear.
One sound at a time. Real listening involves separating simultaneous sources, which the auditory system does using onset timing, common modulation and harmonicity — none of which is in a single spectrum.
Amplitudes are illustrative. The partial lists drawn here are plausible rather than measured. A real clarinet’s spectrum varies with register, dynamic and player, and the essay’s clarinet is a schematic.
Phase is discarded in the figures too. Every waveform drawn here uses zero phase for all partials, which is why they look tidy. Real waveforms look far messier and sound identical, which is the essay’s point and is not something the pictures can demonstrate.
Why the spectrum is a choice of coordinates
A last framing, because it is easy to treat the spectrum as the truth and the waveform as an illusion.
Both are complete descriptions. A waveform and its spectrum contain identical information — the Fourier transform is invertible, and nothing is lost going either way. Neither is more real than the other.
What makes the spectrum the useful description is that the ear performs approximately that transform, so spectral coordinates align with perceptual ones. Two sounds close in spectrum sound alike; two sounds close in waveform may not. A representation is good when small changes in it correspond to small changes in what is being described, and for hearing, the spectrum has that property and the waveform does not.
That is worth generalising. Choosing coordinates that match how a thing is perceived, rather than how it is generated, is what makes cents the right unit for pitch and a cycle the right picture for a rhythm. The same move, three times.
The ladder from here
Later rungs: Fourier analysis proper. Critical bands measured. The missing fundamental and pitch models. Formants, and why a vowel is a vowel regardless of pitch. Spectrograms and time-varying spectra. Transients, and what removing them costs. Inharmonic spectra and their scales. Additive, subtractive and FM synthesis. Auditory scene analysis. And the phase question in full, including the cases where it does matter.
Helmholtz identified individual partials in a played note using a set of tuned glass spheres, one per partial, each with a nipple to fit in the ear. It is a spectrum analyser made of blown glass, it works, and it predates any means of recording sound at all.