Timbre and acoustics

The first fifty milliseconds

A spectrum is supposed to be what makes a trumpet a trumpet. Cut the first fifty milliseconds off a recorded note and listeners stop being able to name the instrument — while the spectrum they are hearing is unchanged. Identity is in the part of the sound that ends before the note has properly started.

Assumes: The shape of a note, which is most of what an instrument is · The ear hears the list, not the shape

The standard account of timbre is spectral. Two instruments playing the same note differ in their partials, the ear hears the list of partials rather than the shape of the wave, and the list is therefore what identifies the instrument. It is a good account, it is nearly a century old in its modern form, and it is missing the part that does most of the work.

Three attacks, the first 50 ms. How loudness changes over the life of a note, for plucked, bowed and struck, drawn over the first 50 milliseconds. By the right-hand edge the plucked note is at 91%, the bowed note is at 36%, the struck note is at 97% — attack times of 4 ms, 140 ms, 2 ms, a spread of 70 to one, and the part a listener uses to tell them apart. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.
Fig. 1 The first fifty milliseconds of three notes, at forty times the resolution of an ordinary envelope plot. A struck note is at nearly full amplitude before a bowed one has begun to speak. Everything that identifies these instruments to a listener happens inside this window, and an ordinary two-second envelope plot draws it as a vertical line.

Fifty milliseconds is a twentieth of a second. At 261 Hz it is thirteen cycles of the fundamental — barely enough for the ear to establish a pitch at all — and it is where the information is.

The experiment

The demonstration is an editing experiment and it has been run in several forms since the 1960s, the best-known being by Jean-Claude Risset and Max Mathews at Bell Labs and by Robert Erickson and others through the following decade.

Take recordings of orchestral instruments playing the same note at the same loudness. Cut off the attack — the first few tens of milliseconds — and present what remains. Ask listeners to name the instrument.

Identification collapses. Instruments that are trivially distinguishable in full become confusable, and the confusions are systematic: the sustained portions of a flute, a violin and a trumpet at the same pitch and level are far more alike than anybody expects.

Three envelopes with the first 50 ms removed. How loudness changes over the life of a note, for plucked, bowed and struck. With the first 50 milliseconds cut away they are all far harder to tell apart, which is the experiment this figure is about.
Fig. 2 The same three envelopes with the first fifty milliseconds removed, which is what the edited stimuli sound like: a note that begins at full amplitude, with no onset. The steady state and the decay are untouched, and the spectral content of what remains is exactly what it was.

Then run it the other way. Present only the attack — the first fifty milliseconds and nothing else — and identification recovers substantially. A fragment far too short to have a musical pitch is enough to name the instrument.

That pair of results is the whole finding, and its form matters: it is not that the attack helps, it is that the attack carries information the steady state does not have.

What is actually in there

Three things happen during an attack, and all of them are excluded from a spectrum measured over a steady note.

The partials arrive at different times. A plucked or struck string starts every partial at once; a bowed one does not, and the higher partials take longer to establish because the bow has to settle into its stick-slip regime. A flute’s fundamental takes tens of milliseconds while its upper partials are present almost immediately. The order in which the components arrive is a property of the excitation mechanism, and it is discarded by any measurement that averages over time.

A note gets duller as it dies. Each partial of a string note against time, with the loss rising as the partial number to the power 2 — so the fundamental takes 6 seconds to fall sixty decibels and the 8th takes 0.09. The heavy line is the power-weighted centroid, falling from partial 1.77 toward the fundamental; it is halfway there after 0.04 seconds. A single-rate envelope would draw all of these as parallel lines and the centroid as a horizontal one, and a struck string does neither: what is left at the end of a long note is very nearly a sine.
Fig. 3 Each partial of a string note against time, with the loss rising as the square of the partial number: the fundamental takes six seconds to fall sixty decibels and the eighth takes 0.09.

A steady-state spectrum is a fiction with a half-life. The list of amplitudes a spectral account works from is only true at one instant, and it is untrue almost immediately afterwards — the upper partials are gone within a tenth of a second while the fundamental is still sounding, so the note’s colour is changing throughout the part a listener is using to identify it.

There is inharmonic noise. A bow scrapes, a hammer thuds, a reed chiffs, a plectrum clicks, a valve clatters. None of it is periodic, none of it is in the harmonic series, and all of it is gone by the time the note is sustaining. It is also the most instrument-specific part of the sound: the mechanism makes a noise, and the noise is a signature of the mechanism.

The pitch is not yet stable. Most instruments arrive at their pitch from somewhere else. A brass player’s note settles from below, a bowed string wanders during the first cycles, and a plucked string is sharp at the moment of release because the displacement raises the tension. Those excursions are small, brief, and characteristic.

Why it fooled everyone for so long

The spectral account is not wrong; it is incomplete, and the reason the incompleteness went unnoticed is worth stating because it is a general hazard.

Fourier analysis needs a window. To resolve partials to within a few hertz the window must be a substantial fraction of a second, and a window that long has averaged the attack into nothing before the analysis begins. The measurement technique that established the spectral account was structurally incapable of seeing the thing it was leaving out.

A note gets duller as it dies. Each partial of a clarinet note against time, with the loss rising as the partial number to the power 0.5 — so the fundamental takes 4 seconds to fall sixty decibels and the 8th takes 1.41. The heavy line is the power-weighted centroid, falling from partial 1.93 toward the fundamental; it is halfway there after 0.25 seconds. A single-rate envelope would draw all of these as parallel lines and the centroid as a horizontal one, and a struck string does neither: what is left at the end of a long note is very nearly a sine.
Fig. 4 The same computation for a clarinet, where the loss rises only as the square root of the partial number: the fundamental takes four seconds and the eighth partial 1.41.

So “steady state” is a property of an instrument rather than of notes in general. A clarinet very nearly has one and a struck string does not, and the exponent is what separates them — which means the attack this essay is about is a larger share of the whole note on some instruments than on others.

There is a second reason, and it is about what synthesis made obvious. Additive synthesis from a measured spectrum produces a tone that is recognisably the right family and unmistakably artificial, and for decades that was blamed on not having enough partials. Risset’s demonstration was that adding time-varying behaviour — partials with individual attack times and independent envelopes — produced a convincing trumpet with no more spectral detail than before.

The synthesiser was the experiment. What has to be put in to get a convincing tone out is a direct measure of what the ear is using, and what had to be put in was time.

How short is short enough

The fifty milliseconds in this essay’s title is a round number, and the real quantity varies by instrument in a way that is itself informative.

The attack time — the interval from the first audible sound to the peak — runs from about two milliseconds for a struck or plucked note to well over a hundred for a bowed or blown one. A piano hammer is in contact with the string for a couple of milliseconds and the note is at full amplitude almost immediately. A cello note played at the frog with a slow bow may take two hundred milliseconds to speak.

That range has a consequence for the editing experiment, and it is worth naming because the experiment is usually described as though one operation were being applied to every instrument. Against the attack times this collection carries, a fixed fifty-millisecond cut is:

instrument attack what a fifty-millisecond cut takes
marimba 3 ms the whole attack and 47 ms of steady state
piano 8 the whole attack and 42 ms of steady state
trumpet 30 the whole attack and 20 ms
clarinet 45 the whole attack and 5 ms
flute 60 83% of the attack
bowed violin 90 56% of the attack
a sung vowel 110 45% of the attack

A fixed cut removes the whole of a struck instrument’s onset and half of a bowed one’s, so the stimuli in the two conditions are not the same manipulation. The struck instruments in such an experiment are heard entirely without their attacks and the bowed ones with a substantial part of theirs intact — which is the direction that would understate the effect, since the instruments most likely to survive the edit are the ones the edit least affected.

So the finding is robust against the objection and the experiment is cruder than the finding. A version matched to each instrument’s own attack would cut 3 milliseconds from a marimba and 110 from a voice, and would be the version that tests the claim rather than a fixed window that tests it unevenly.

The other half of the title has a number too. Fifty milliseconds is thirteen cycles of middle C, and a fragment that long specifies its own frequency to no better than 65 cents — two thirds of a semitone, by the Fourier bound alone and before any question about the ear. So “far too short to have a musical pitch” is exact rather than rhetorical: a listener naming an instrument from such a fragment is doing it while being unable to name the note to within a semitone.

The bound falls quickly, which sets the scale of the whole subject. Ten milliseconds gives 303 cents — a quarter of an octave, and the fragment could be almost any note. Twenty gives 158, fifty gives 65, a hundred gives 33 and two hundred gives 16, which is the first length at which the frequency is pinned more tightly than the ear’s own steady-tone limen. So there is a window between about ten and two hundred milliseconds in which a sound has an identity and does not yet have a pitch, and every result in this essay lives inside it. That is the sense in which the attack is a different kind of information from the spectrum: it is available before the quantity a spectrum is a description of exists.

Three attacks, the first 200 ms. How loudness changes over the life of a note, for plucked, bowed and struck, drawn over the first 200 milliseconds. By the right-hand edge the plucked note is at 61%, the bowed note is at 91%, the struck note is at 86% — attack times of 4 ms, 140 ms, 2 ms, a spread of 70 to one, and the part a listener uses to tell them apart. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.
Fig. 5 The first two hundred milliseconds, which is four times the window of the opening figure. At this scale the plucked and struck notes have finished their attacks and begun to decay while the bowed one is still climbing. The spread between the fastest and slowest attack here is a factor of seventy, and it is the single largest difference between these three sounds.

That spread is why “the attack” is not one duration. The useful statement is that identification needs the interval from the onset until the sound has settled into a periodic steady state, and that interval is short for percussive instruments and long for sustained ones.

The asymmetry has a consequence for the editing experiment. Removing fifty milliseconds from a struck note removes the whole attack and a chunk of the decay; removing fifty from a bowed note removes only the beginning of the attack, and the note still has an onset of a sort. That is one reason bowed strings survive the edit better than the raw finding suggests, and it is a confound worth knowing about before reading too much into any single number.

What survives the removal

The negative result is not total, and the exceptions locate the effect precisely.

Instruments with a strongly distinctive steady spectrum survive better. The clarinet’s near-absence of even partials is a large and easily heard feature that persists into the sustain, and clarinets remain identifiable with their onsets removed rather better than flutes or violins do.

Vibrato survives, and it is a second time-varying cue. A note with vibrato has partials moving coherently, which is a strong grouping cue and is instrument-specific in rate and depth.

And the confusions are not random. Edited notes are confused within families — bowed strings with each other, brass with brass — which says the steady spectrum carries family information and the attack carries the rest. A spectral account is correct about what it describes and describes less than it claims.

Where the model stops

The envelope drawn in these figures is an idealisation with four segments — attack, decay, sustain, release — and it is a synthesiser’s model rather than an instrument’s.

Real envelopes are not four straight lines. A piano note has a fast initial decay and a much slower one behind it, from the two polarisations of the string’s motion coupling to the bridge at different rates, and no ADSR model produces that. A bowed note has no sustain level in any fixed sense, because the player is continuously varying it. A note’s envelope also depends on how hard it was played, and not proportionally: a loud piano note is not a quiet one scaled up, it has different partial content, which is why sample-based instruments record several dynamic levels rather than adjusting the gain.

More seriously, the four-segment model describes the amplitude of the whole note. What the attack experiments show is that the useful information is in the amplitudes of the individual partials, which have their own onsets and are not synchronised. A single envelope curve, however accurately drawn, has already averaged away the thing being described.

The first eight partials of a string. A string vibrating in one, two, three and more equal parts, with the frequency ratio and the nearest named note beside each. The seventh partial is a third of a semitone flat of anything on a keyboard, which is a fact about strings rather than about tuning.
Fig. 6 Eight partials of one note. During a steady state their amplitudes hold the relationship a spectrum records. During an attack each has its own history — arriving at its own time, at its own rate — and it is that set of individual histories rather than their eventual balance that a listener uses to name the instrument.

Backwards, which is the other demonstration

There is a second and cheaper demonstration of the same point, and it needs no editing at all: play a recording backwards.

A reversed piano note is instantly recognisable as not a piano, and what it sounds like is a wind or bowed instrument with a strange ending. The spectrum is identical — reversal does not change which frequencies are present, or in what proportion — and every partial is exactly where it was. What has changed is only the direction of the envelope, and the identity is gone.

This is the same finding as the editing experiment with the manipulation moved. Cutting the attack removes the information; reversing the note supplies the wrong information, an onset shaped like a decay, and the auditory system takes the shape at face value.

It is also the reason a reversed reverberation tail is such a recognisable studio effect. A hall’s decay played backwards is a swell, and it reads as a sound approaching rather than a room responding — because what a room does is entirely in the tail, and putting the tail first turns a room into a gesture.

Two consequences follow that are worth keeping.

The auditory system is not time-symmetric, which is a stronger claim than it looks. A Fourier magnitude spectrum is time-symmetric: reversal leaves it unchanged. So any account of timbre stated purely in terms of the magnitude spectrum predicts that a reversed note sounds like the original, and every listener knows it does not. That is a one-line refutation of the strong spectral account, available to anybody with a recording.

And the effect is not about musical training. Untrained listeners are as good at spotting a reversed piano as trained ones, which places the mechanism well below the level at which interval categories are learned.

Whose instruments, and where the result generalises

The identification experiments were run on Western orchestral instruments, with Western listeners, on isolated notes. All three restrictions matter.

Isolated notes are the strongest restriction. In real music a note arrives in a context — after another note, with an instrument already established, in a texture — and the attack is correspondingly less critical because identification has already happened. The experiments measure the information available in a fragment, not the information used in listening.

The result nevertheless generalises well outside music, which is the best evidence that it is about hearing rather than about instruments. Consonants are attack transients. The distinction between ba and pa is a matter of tens of milliseconds at the onset of a syllable, the distinction between da and ga is in the direction the formants move during the first few tens of milliseconds, and a vowel — which is a steady state — is comparatively robust. Speech and instrument identification are using the same part of the signal.

The practical corollary appeared in audio engineering long before the psychoacoustics did. A gate or compressor with a slow attack destroys instrument identity while leaving the level curve looking correct, and every engineer knows that the first milliseconds of a drum hit are not to be touched.

What this does to the word “timbre”

Timbre is defined, in the standards documents, by what it is not: the attribute by which a listener judges two sounds of the same loudness, pitch and duration to be dissimilar. That is a definition by subtraction and it has been criticised for a century, usually on the grounds that it is unhelpfully negative.

The finding here suggests a sharper complaint. The definition holds pitch and duration constant, and the attack is neither — it is a property of the beginning, largely independent of how long the note lasts and of what pitch it settles on. So the standard definition does not exclude it, but nothing in the way the definition is phrased points at it either, and generations of textbooks have moved directly from “not pitch, not loudness” to “therefore spectrum”.

A more useful decomposition has three parts, and each answers a different question.

The steady spectrum says what the resonating system is. It is what the source–filter account of a vowel describes, it survives a change of pitch, and it places an instrument in a family.

The attack says what the excitation mechanism is: struck, plucked, bowed, blown. It is the part this essay is about, and it is what identifies the individual instrument within its family.

The time-varying behaviour of both — vibrato, the way a spectrum changes with dynamic level, the way a note’s brightness falls as it decays — carries the rest, and is the part that separates a real instrument from an accurate synthesis of one.

None of the three is dispensable, and the reason “timbre” resists definition is probably that it names three things at once rather than one thing badly.

What the picture cannot show

An envelope plot has one curve per note, and the finding is that the note does not have one envelope — it has one per partial, and their asynchrony is the content. Drawing that requires a different figure and, honestly, a different medium: the phenomenon is thirty milliseconds long, and the reason it took a synthesiser to discover it is that thirty milliseconds is not a duration a picture is good at.

The window itself is worth a second look, because it is not only an interval over which an instrument reveals itself. It is the length of the auditory present — the span inside which a later sound can hide an earlier one and inside which a reflection is heard as part of the note rather than after it. That three separate literatures arrived at a few tens of milliseconds is not a coincidence about instruments; it is a fact about the listener they were all measuring.

The figures also cannot show the noise. Every plot here is of a periodic signal with a smooth envelope, and the scrape, chiff and thud that carry much of the identity are aperiodic and have no place in a partial list or an amplitude curve. The most instrument-specific part of an instrument’s sound is the part this site’s whole apparatus is least able to draw.

The ladder from here climbs back towards the steady state that this essay has spent its length qualifying — the source and the filter, which is the other half of what makes a sound identifiable and which does survive the removal of the attack.

Part 2 of 10

One essay in the series on envelope. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 24.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientEnvelopeOnsetSource-filterTimbre