The first fifty milliseconds
Assumes: The shape of a note, which is most of what an instrument is · The ear hears the list, not the shape
The standard account of timbre is spectral. Two instruments playing the same note differ in their partials, the ear hears the list of partials rather than the shape of the wave, and the list is therefore what identifies the instrument. It is a good account, it is nearly a century old in its modern form, and it is missing the part that does most of the work.
Fifty milliseconds is a twentieth of a second. At 261 Hz it is thirteen cycles of the fundamental — barely enough for the ear to establish a pitch at all — and it is where the information is.
The experiment
The demonstration is an editing experiment and it has been run in several forms since the 1960s, the best-known being by Jean-Claude Risset and Max Mathews at Bell Labs and by Robert Erickson and others through the following decade.
Take recordings of orchestral instruments playing the same note at the same loudness. Cut off the attack — the first few tens of milliseconds — and present what remains. Ask listeners to name the instrument.
Identification collapses. Instruments that are trivially distinguishable in full become confusable, and the confusions are systematic: the sustained portions of a flute, a violin and a trumpet at the same pitch and level are far more alike than anybody expects.
Then run it the other way. Present only the attack — the first fifty milliseconds and nothing else — and identification recovers substantially. A fragment far too short to have a musical pitch is enough to name the instrument.
That pair of results is the whole finding, and its form matters: it is not that the attack helps, it is that the attack carries information the steady state does not have.
What is actually in there
Three things happen during an attack, and all of them are excluded from a spectrum measured over a steady note.
The partials arrive at different times. A plucked or struck string starts every partial at once; a bowed one does not, and the higher partials take longer to establish because the bow has to settle into its stick-slip regime. A flute’s fundamental takes tens of milliseconds while its upper partials are present almost immediately. The order in which the components arrive is a property of the excitation mechanism, and it is discarded by any measurement that averages over time.
A steady-state spectrum is a fiction with a half-life. The list of amplitudes a spectral account works from is only true at one instant, and it is untrue almost immediately afterwards — the upper partials are gone within a tenth of a second while the fundamental is still sounding, so the note’s colour is changing throughout the part a listener is using to identify it.
There is inharmonic noise. A bow scrapes, a hammer thuds, a reed chiffs, a plectrum clicks, a valve clatters. None of it is periodic, none of it is in the harmonic series, and all of it is gone by the time the note is sustaining. It is also the most instrument-specific part of the sound: the mechanism makes a noise, and the noise is a signature of the mechanism.
The pitch is not yet stable. Most instruments arrive at their pitch from somewhere else. A brass player’s note settles from below, a bowed string wanders during the first cycles, and a plucked string is sharp at the moment of release because the displacement raises the tension. Those excursions are small, brief, and characteristic.
Why it fooled everyone for so long
The spectral account is not wrong; it is incomplete, and the reason the incompleteness went unnoticed is worth stating because it is a general hazard.
Fourier analysis needs a window. To resolve partials to within a few hertz the window must be a substantial fraction of a second, and a window that long has averaged the attack into nothing before the analysis begins. The measurement technique that established the spectral account was structurally incapable of seeing the thing it was leaving out.
So “steady state” is a property of an instrument rather than of notes in general. A clarinet very nearly has one and a struck string does not, and the exponent is what separates them — which means the attack this essay is about is a larger share of the whole note on some instruments than on others.
There is a second reason, and it is about what synthesis made obvious. Additive synthesis from a measured spectrum produces a tone that is recognisably the right family and unmistakably artificial, and for decades that was blamed on not having enough partials. Risset’s demonstration was that adding time-varying behaviour — partials with individual attack times and independent envelopes — produced a convincing trumpet with no more spectral detail than before.
The synthesiser was the experiment. What has to be put in to get a convincing tone out is a direct measure of what the ear is using, and what had to be put in was time.
How short is short enough
The fifty milliseconds in this essay’s title is a round number, and the real quantity varies by instrument in a way that is itself informative.
The attack time — the interval from the first audible sound to the peak — runs from about two milliseconds for a struck or plucked note to well over a hundred for a bowed or blown one. A piano hammer is in contact with the string for a couple of milliseconds and the note is at full amplitude almost immediately. A cello note played at the frog with a slow bow may take two hundred milliseconds to speak.
That range has a consequence for the editing experiment, and it is worth naming because the experiment is usually described as though one operation were being applied to every instrument. Against the attack times this collection carries, a fixed fifty-millisecond cut is:
| instrument | attack | what a fifty-millisecond cut takes |
|---|---|---|
| marimba | 3 ms | the whole attack and 47 ms of steady state |
| piano | 8 | the whole attack and 42 ms of steady state |
| trumpet | 30 | the whole attack and 20 ms |
| clarinet | 45 | the whole attack and 5 ms |
| flute | 60 | 83% of the attack |
| bowed violin | 90 | 56% of the attack |
| a sung vowel | 110 | 45% of the attack |
A fixed cut removes the whole of a struck instrument’s onset and half of a bowed one’s, so the stimuli in the two conditions are not the same manipulation. The struck instruments in such an experiment are heard entirely without their attacks and the bowed ones with a substantial part of theirs intact — which is the direction that would understate the effect, since the instruments most likely to survive the edit are the ones the edit least affected.
So the finding is robust against the objection and the experiment is cruder than the finding. A version matched to each instrument’s own attack would cut 3 milliseconds from a marimba and 110 from a voice, and would be the version that tests the claim rather than a fixed window that tests it unevenly.
The other half of the title has a number too. Fifty milliseconds is thirteen cycles of middle C, and a fragment that long specifies its own frequency to no better than 65 cents — two thirds of a semitone, by the Fourier bound alone and before any question about the ear. So “far too short to have a musical pitch” is exact rather than rhetorical: a listener naming an instrument from such a fragment is doing it while being unable to name the note to within a semitone.
The bound falls quickly, which sets the scale of the whole subject. Ten milliseconds gives 303 cents — a quarter of an octave, and the fragment could be almost any note. Twenty gives 158, fifty gives 65, a hundred gives 33 and two hundred gives 16, which is the first length at which the frequency is pinned more tightly than the ear’s own steady-tone limen. So there is a window between about ten and two hundred milliseconds in which a sound has an identity and does not yet have a pitch, and every result in this essay lives inside it. That is the sense in which the attack is a different kind of information from the spectrum: it is available before the quantity a spectrum is a description of exists.
That spread is why “the attack” is not one duration. The useful statement is that identification needs the interval from the onset until the sound has settled into a periodic steady state, and that interval is short for percussive instruments and long for sustained ones.
The asymmetry has a consequence for the editing experiment. Removing fifty milliseconds from a struck note removes the whole attack and a chunk of the decay; removing fifty from a bowed note removes only the beginning of the attack, and the note still has an onset of a sort. That is one reason bowed strings survive the edit better than the raw finding suggests, and it is a confound worth knowing about before reading too much into any single number.
What survives the removal
The negative result is not total, and the exceptions locate the effect precisely.
Instruments with a strongly distinctive steady spectrum survive better. The clarinet’s near-absence of even partials is a large and easily heard feature that persists into the sustain, and clarinets remain identifiable with their onsets removed rather better than flutes or violins do.
Vibrato survives, and it is a second time-varying cue. A note with vibrato has partials moving coherently, which is a strong grouping cue and is instrument-specific in rate and depth.
And the confusions are not random. Edited notes are confused within families — bowed strings with each other, brass with brass — which says the steady spectrum carries family information and the attack carries the rest. A spectral account is correct about what it describes and describes less than it claims.
Where the model stops
The envelope drawn in these figures is an idealisation with four segments — attack, decay, sustain, release — and it is a synthesiser’s model rather than an instrument’s.
Real envelopes are not four straight lines. A piano note has a fast initial decay and a much slower one behind it, from the two polarisations of the string’s motion coupling to the bridge at different rates, and no ADSR model produces that. A bowed note has no sustain level in any fixed sense, because the player is continuously varying it. A note’s envelope also depends on how hard it was played, and not proportionally: a loud piano note is not a quiet one scaled up, it has different partial content, which is why sample-based instruments record several dynamic levels rather than adjusting the gain.
More seriously, the four-segment model describes the amplitude of the whole note. What the attack experiments show is that the useful information is in the amplitudes of the individual partials, which have their own onsets and are not synchronised. A single envelope curve, however accurately drawn, has already averaged away the thing being described.
Backwards, which is the other demonstration
There is a second and cheaper demonstration of the same point, and it needs no editing at all: play a recording backwards.
A reversed piano note is instantly recognisable as not a piano, and what it sounds like is a wind or bowed instrument with a strange ending. The spectrum is identical — reversal does not change which frequencies are present, or in what proportion — and every partial is exactly where it was. What has changed is only the direction of the envelope, and the identity is gone.
This is the same finding as the editing experiment with the manipulation moved. Cutting the attack removes the information; reversing the note supplies the wrong information, an onset shaped like a decay, and the auditory system takes the shape at face value.
It is also the reason a reversed reverberation tail is such a recognisable studio effect. A hall’s decay played backwards is a swell, and it reads as a sound approaching rather than a room responding — because what a room does is entirely in the tail, and putting the tail first turns a room into a gesture.
Two consequences follow that are worth keeping.
The auditory system is not time-symmetric, which is a stronger claim than it looks. A Fourier magnitude spectrum is time-symmetric: reversal leaves it unchanged. So any account of timbre stated purely in terms of the magnitude spectrum predicts that a reversed note sounds like the original, and every listener knows it does not. That is a one-line refutation of the strong spectral account, available to anybody with a recording.
And the effect is not about musical training. Untrained listeners are as good at spotting a reversed piano as trained ones, which places the mechanism well below the level at which interval categories are learned.
Whose instruments, and where the result generalises
The identification experiments were run on Western orchestral instruments, with Western listeners, on isolated notes. All three restrictions matter.
Isolated notes are the strongest restriction. In real music a note arrives in a context — after another note, with an instrument already established, in a texture — and the attack is correspondingly less critical because identification has already happened. The experiments measure the information available in a fragment, not the information used in listening.
The result nevertheless generalises well outside music, which is the best evidence that it is about hearing rather than about instruments. Consonants are attack transients. The distinction between ba and pa is a matter of tens of milliseconds at the onset of a syllable, the distinction between da and ga is in the direction the formants move during the first few tens of milliseconds, and a vowel — which is a steady state — is comparatively robust. Speech and instrument identification are using the same part of the signal.
The practical corollary appeared in audio engineering long before the psychoacoustics did. A gate or compressor with a slow attack destroys instrument identity while leaving the level curve looking correct, and every engineer knows that the first milliseconds of a drum hit are not to be touched.
What this does to the word “timbre”
Timbre is defined, in the standards documents, by what it is not: the attribute by which a listener judges two sounds of the same loudness, pitch and duration to be dissimilar. That is a definition by subtraction and it has been criticised for a century, usually on the grounds that it is unhelpfully negative.
The finding here suggests a sharper complaint. The definition holds pitch and duration constant, and the attack is neither — it is a property of the beginning, largely independent of how long the note lasts and of what pitch it settles on. So the standard definition does not exclude it, but nothing in the way the definition is phrased points at it either, and generations of textbooks have moved directly from “not pitch, not loudness” to “therefore spectrum”.
A more useful decomposition has three parts, and each answers a different question.
The steady spectrum says what the resonating system is. It is what the source–filter account of a vowel describes, it survives a change of pitch, and it places an instrument in a family.
The attack says what the excitation mechanism is: struck, plucked, bowed, blown. It is the part this essay is about, and it is what identifies the individual instrument within its family.
The time-varying behaviour of both — vibrato, the way a spectrum changes with dynamic level, the way a note’s brightness falls as it decays — carries the rest, and is the part that separates a real instrument from an accurate synthesis of one.
None of the three is dispensable, and the reason “timbre” resists definition is probably that it names three things at once rather than one thing badly.
What the picture cannot show
An envelope plot has one curve per note, and the finding is that the note does not have one envelope — it has one per partial, and their asynchrony is the content. Drawing that requires a different figure and, honestly, a different medium: the phenomenon is thirty milliseconds long, and the reason it took a synthesiser to discover it is that thirty milliseconds is not a duration a picture is good at.
The window itself is worth a second look, because it is not only an interval over which an instrument reveals itself. It is the length of the auditory present — the span inside which a later sound can hide an earlier one and inside which a reflection is heard as part of the note rather than after it. That three separate literatures arrived at a few tens of milliseconds is not a coincidence about instruments; it is a fact about the listener they were all measuring.
The figures also cannot show the noise. Every plot here is of a periodic signal with a smooth envelope, and the scrape, chiff and thud that carry much of the identity are aperiodic and have no place in a partial list or an amplitude curve. The most instrument-specific part of an instrument’s sound is the part this site’s whole apparatus is least able to draw.
The ladder from here climbs back towards the steady state that this essay has spent its length qualifying — the source and the filter, which is the other half of what makes a sound identifiable and which does survive the removal of the attack.
Part 2 of 10
One essay in the series on envelope. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 24.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Attack transientEnvelopeOnsetSource-filterTimbre
- The blend arrives before the note does attack transient, envelope, onset, source-filter, timbre
- A note is heard after it starts attack transient, envelope, onset
- An attack time is not an attack attack transient, envelope, onset
- Read at two different heights attack transient, envelope, onset
- What the tongue actually removes attack transient, envelope, onset
- A blown note does not start late, it starts slowly attack transient, onset