Timbre and acoustics

The shape of a note, which is most of what an instrument is

Cut the first fifty milliseconds off a recorded piano and listeners stop calling it a piano. The attack carries more identity than the steady tone it leads into, and it is the part every spectrum plot leaves out.

A note is not a steady thing. It starts, it changes, it stops, and the changing is not incidental to what it sounds like.

Three envelopesHow loudness changes over the life of a note, for a plucked, a bowed and a struck instrument. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.00.511.5200.20.40.60.81secondsamplitudeplucked — no sustainbowedstruck — no sustainkey released
Fig. 1 Three envelopes: how loudness changes over the life of a note for a plucked, a bowed and a struck instrument. The plucked and struck notes have no sustain at all — they begin decaying immediately — and the bowed one holds until the bow stops.

The classic demonstration is a recording of a piano note with its attack removed. What remains is a decaying tone with a piano’s exact spectrum, and listeners identify it as an organ, or a flute, or something they cannot name. It is not identifiable as a piano.

The reverse works too. Take the attack transient of one instrument, splice it onto the steady tone of another, and listeners tend to name the instrument the attack came from.

Whatever carries the identity of an instrument, a great deal of it is in the first fifty milliseconds.

What the envelope describes

The envelope is the outline of a note’s amplitude over time, and synthesiser convention breaks it into four segments.

Attack — the rise from silence to peak. A plucked string reaches peak in a few milliseconds; a bowed note takes tens to hundreds; an organ pipe takes tens.

Decay — the fall from the initial peak to whatever level is sustained.

Sustain — the level held while the note continues. For a bowed or blown instrument this is a real level maintained by continuous energy input. For a plucked or struck instrument there is no sustain: the note decays continuously from its peak, because nothing is putting energy back in.

Release — what happens after the player stops. A damped piano string stops fast; an undamped one rings; a note in a reverberant room continues in the room.

That division is a synthesiser abstraction rather than a physical law, and it fits some instruments well and others badly. It is a good description of a plucked string and a poor one of a trumpet, whose envelope depends continuously on what the player’s air is doing.

The excitation decides everything

The envelope is not an arbitrary property. It follows from how energy gets into the resonator, and instruments sort into two classes on exactly that question.

Impulsive excitation — plucked, struck, hammered. Energy is delivered in one burst and the note decays from there. Piano, guitar, harp, marimba, drum. The player has control over the beginning and essentially none afterwards.

Continuous excitation — bowed, blown, driven. Energy is supplied throughout, so the note can be sustained, swelled and shaped. Violin, trumpet, flute, organ, voice.

The distinction runs deeper into the music than it looks. Continuous-excitation instruments can hold a note and shape it, so they carry melody with an unbroken line, and repertoire written for them can rely on sustain and crescendo. Impulsive instruments cannot: a pianist creating a crescendo across a held note is doing something impossible, and the illusion is created by the notes around it.

It also explains a fact about keyboard writing. Harpsichord music is full of ornaments and repeated notes, because the instrument cannot make a note louder or longer once struck, and the only way to sustain interest across a long value is to re-strike. Those ornaments are usually described as stylistic. They are at least as much a workaround for an envelope.

Spectra move too

Describing a note as a fixed spectrum with an amplitude envelope on top is still a simplification, and the simplification is where synthetic sound gives itself away.

Real partials decay at different rates. On a piano, the upper partials die much faster than the lower ones, so the spectrum darkens continuously through the note. On a struck bell the reverse can happen. On a bowed string the balance shifts with bow pressure and speed, continuously, under the player’s control.

Four spectra of the same noteThe amplitude of each partial for four timbres at the same pitch. These are the exact lists the sound buttons on this site synthesise from, so the picture and the sound are the same data.1pureone partial, nothing else12345678stringall partials, falling12345678clarineteven partials nearly absent123456789bellodd partials onlyamplitude
Fig. 2 Four partial lists, drawn as though they were fixed. They are not: on a real instrument each of these bars has its own envelope, and the whole picture reshapes itself over the course of a single note.

So a full description of a note is a time-varying spectrum — a surface rather than a list plus a curve. That is what a spectrogram draws, and it is why additive synthesis with static partials sounds synthetic no matter how carefully the amplitudes are chosen. The sounds on this site are exactly that: honest sine stacks with a simple envelope, which is why they sound like a synthesiser rather than like an oboe. Attempting the imitation and failing would be worse than not attempting it.

What the attack contains

The attack is not simply a fast rise in amplitude. It is a burst of noise and inharmonic energy that the steady tone does not contain at all.

A bowed string’s attack includes the scrape of rosin catching. A flute’s includes the breath noise before the air column locks into oscillation. A piano’s includes the mechanical thump of the hammer and the key bed. A trumpet’s includes the moment before the lip and the air column agree on a frequency.

None of that is periodic, none of it is in the harmonic series, and all of it is highly characteristic. Listeners identify instruments from attacks of a few tens of milliseconds — far too short for the ear to have performed a useful frequency analysis of the steady part.

The first eight partials of a stringA string vibrating in one, two, three and more equal parts, with the frequency ratio and the nearest named note beside each. The seventh partial is a third of a semitone flat of anything on a keyboard, which is a fact about strings rather than about tuning.1130.8 HzC2261.6 HzC3392.4 HzG4523.3 HzC5654.1 HzE -14¢6784.9 HzG7915.7 HzB♭ -31¢81046.5 HzCpartialthe dots are the nodes — the places that do not move
Fig. 3 The steady-state partials of a string. Everything in this picture arrives after the attack is over, and the attack is what a listener uses to name the instrument.

There is a general point here about what the spectrum leaves out. Phase is largely inaudible in a steady tone; in a transient, the relative timing of components is exactly what makes it sound like a pluck rather than a strike. Phase deafness is a property of steady sounds, and attacks are not steady.

Onset is what separates sources

The envelope does one more job, and it is arguably the most important thing hearing does.

A listener in a room with several instruments playing receives one pressure signal containing everything. Separating it into sources is not optional — it is the basic problem of hearing — and common onset is the strongest cue available. Components that start at the same moment are grouped as one sound; components that start at different moments are assigned to different sources.

This is why an orchestra is heard as instruments rather than as a wall. It is why a chord played exactly together fuses into one sonority and the same chord spread by twenty milliseconds is heard as separate notes. And it is why a synthesised chord in which every partial starts at exactly the same instant sounds artificially fused — real ensembles never achieve that, and the ear expects the slop.

Two further grouping cues work the same way: components that share a vibrato move together and are grouped, and components in a harmonic relationship are grouped. All three are the auditory system deciding what belongs to what, and all three are about behaviour over time rather than about a snapshot.

Notation has nothing for this

Staff notation records pitch precisely and duration precisely, and it has almost nothing for the shape of a note.

A phrase in ordinary notationA phrase written on a stave. Notation records what a player should do rather than what the air does, so it shows the note names exactly and the pitches only by convention — which is the reason this site draws so much of its evidence some other way.four notes, with no way to say what shape any of them has
Fig. 4 Four quarter notes. The page specifies which pitches and for how long. Whether they are plucked or bowed, how sharply each begins, and how the sound behaves after the attack are all left to the instrument and the player.

What exists is a small vocabulary of articulation marks — staccato, tenuto, accent, marcato — which are relative and instrument-dependent rather than specified. A staccato quarter note on a violin and on a marimba are not the same instruction, and neither is a measurement.

That is not a defect exactly. Notation was designed to instruct a player who already knows their instrument, not to describe a sound. But it means that a great deal of what makes a performance sound like something is not on the page, and that transcription between instruments loses more than the pitches suggest it should — notation records instructions, not sounds.

Why an instrument sustains or does not

The two classes of excitation come from different physics, and the mechanism is worth having because it explains the envelope rather than describing it.

An impulsively excited instrument receives a fixed amount of energy and then loses it — to the air as sound, to the bridge and body as vibration, and to internal friction. The loss is proportional to the energy present, so the decay is exponential, and the rate is set by how efficiently the instrument radiates. That produces an uncomfortable trade: an instrument that radiates efficiently is loud and decays fast, and one that decays slowly is quiet. A banjo is loud and short; a well-damped classical guitar is quieter and longer.

A continuously excited instrument sustains because a nonlinear feedback loop keeps putting energy in. A bowed string is caught by the bow’s rosin, dragged, released when the restoring force exceeds friction, and caught again — a stick-slip cycle that locks to the string’s own frequency. A reed and an air column do the same with a pressure-controlled valve. In both cases the oscillation is self-sustaining and the player controls the energy input continuously.

The consequence is that sustain is not a property that could be added to a piano by better engineering. It requires a mechanism that puts energy in continuously, which is why every attempt at a sustaining keyboard — the Hurdy-gurdy, the Geigenwerk, the modern electromagnetic sustain systems — replaces the hammer with something that keeps acting.

The same partials, drawn as pressureEach spectrum summed into the wave it actually produces, over two cycles. The shapes are strikingly different and the ear has almost no access to that difference — what it hears is the list of partials, not the shape they add up to.pure1 partialstring8 partialsclarinet8 partials2 cycles · each normalised by its own peak
Fig. 5 Three steady waveforms. Every one of these presupposes an oscillation already established and maintained, which for half the instruments in an orchestra is a condition that only lasts as long as the player keeps working.

Loudness is not amplitude

One more axis the envelope figures quietly hold constant, and it does not behave the way the drawings imply.

Perceived loudness grows roughly as the cube root of intensity, so a tenfold increase in power is heard as roughly a doubling. It also depends strongly on frequency: the ear is far less sensitive at low frequencies, and much more so quietly than loudly, which is why a quiet mix sounds thin and why “loudness compensation” controls exist.

For an instrument, playing louder is not the same as playing the same note bigger. Brass instruments get dramatically brighter with level, because the air column’s nonlinearity transfers energy upward into the higher partials — a fortissimo trumpet has a completely different spectrum from a piano one, not merely a larger amplitude. Bowed strings do the same to a lesser degree.

So dynamics and timbre are entangled, and the envelope drawn as a single amplitude curve treats them as separable. For a plucked string that is nearly true. For brass it is quite wrong, and it is why synthesised brass that only scales amplitude sounds like nothing at all.

The attack decides what gets heard together

The envelope’s most consequential job is not describing one note but separating several, and it happens in the first few milliseconds.

Components that begin together are grouped as one source; components that begin apart are assigned to different ones. That single rule is why an orchestra resolves into instruments, why a chord played precisely together fuses into a sonority while the same chord spread by twenty milliseconds is heard as separate notes, and why a synthesised chord whose partials all start on the same sample sounds unnaturally welded.

The window is narrow — differences above about thirty milliseconds are heard as separate events — and it is why ensemble precision matters in a way that is perceptual rather than aesthetic. A slightly ragged entry is not merely untidy; it changes what the listener assigns to which instrument.

Four spectra of the same noteThe amplitude of each partial for four timbres at the same pitch. These are the exact lists the sound buttons on this site synthesise from, so the picture and the sound are the same data.1pureone partial, nothing else12345678stringall partials, falling12345678clarineteven partials nearly absent123456789bellodd partials onlyamplitude
Fig. 6 Four partial lists. Whether a listener hears these as four instruments or as one composite sound is decided almost entirely by whether their onsets coincide — a temporal question, settled before the spectra are established.

Vibrato does the same job later in the note: partials sharing a common frequency modulation are grouped, which is why a singer with vibrato separates from an accompaniment that has none. Both cues are about behaviour over time, and neither is in a spectrum.

Where the model stops

Four segments is a synthesiser’s abstraction. Real envelopes are continuous curves shaped by physical processes, and the attack-decay-sustain-release division fits a plucked string reasonably and a wind instrument poorly.

One envelope for the whole note. Every partial has its own, and the difference between them is what makes the spectrum evolve.

No noise. The figures draw smooth amplitude curves. Real attacks contain broadband noise, which is characteristic and which none of these pictures represents — one of several things a spectrum leaves out.

Dynamic level is a separate axis. A loudly played note is not a scaled-up quiet one — brass in particular gets dramatically brighter with level, because the nonlinearity of the air column pumps energy upward. The envelope figures hold level constant.

The room is not in it. Everything here describes a note in isolation. What a room does to the decay is comparable in magnitude to what the instrument does.

What a synthesiser has to get right

Additive synthesis — summing sinusoids — can in principle produce any sound at all, and in practice it is the hardest way to imitate an instrument. The reason is a good summary of everything above.

To reproduce a sustained tone, a synthesiser needs a partial list, which is a handful of numbers. To reproduce a note, it needs a separate amplitude envelope for every partial, a frequency envelope for each to capture the small pitch drift that real instruments have, a noise component for the attack, and all of that varying with dynamic level and with register.

For a single instrument that is thousands of parameters, and they have to be measured from recordings rather than reasoned about. This is why additive synthesis, despite being the most general method, lost to techniques that get the behaviour approximately right with far fewer controls: subtractive synthesis, which filters a rich source; FM, which produces complex evolving spectra from two oscillators; and sampling, which sidesteps the problem entirely by recording the answer.

The sounds on this site are deliberately at the simple end of that scale — a fixed partial list and one envelope, which produces something that is honestly a synthesiser rather than a poor imitation of an oboe. The spectra are drawn exactly as they are synthesised, and the point is the correspondence rather than the realism.

Notation, and the gap it leaves

The envelope is the clearest case of something musically decisive that the page does not record, and it is worth putting alongside the others.

Staff notation specifies pitch exactly, duration exactly, and everything about the sound approximately or not at all. Articulation marks are relative and instrument-dependent; dynamics are a six-step ordinal scale with no absolute meaning; timbre is specified only by naming the instrument.

That is not a failure of design. Notation was built to instruct a player who already knows their instrument, and a player supplied the rest. It does mean that the notated score and the sounding music are related the way a recipe is related to a meal, and that transcription between instruments loses more than the pitches suggest.

It also means that a great deal of the twentieth century’s notational experimentation was an attempt to close exactly this gap — and that electronic music, where the composer specifies the sound directly, sidesteps it by removing the performer.

The ladder from here

Later rungs: attack transients analysed. Time-varying spectra and the spectrogram. The physics of impulsive versus continuous excitation. Bowed-string motion and the Helmholtz kink. Brass nonlinearity and why loud is bright. Auditory scene analysis and the grouping cues. Onset asynchrony as an ensemble parameter. Synthesis methods and why additive is the hardest way to imitate anything. And articulation notation, and the gap between what is written and what is played.

Cutting the attack off a recording and playing the remainder to musicians is a demonstration that has been repeated in lecture halls since the 1960s, and it never fails to produce the same reaction: an audience confidently naming the wrong instrument, twice in a row.