Timbre and acoustics

The middle nobody could have guessed

A struck note has no steady state, only a slide from one spectrum to another — so the question is what the middle carries that the ends do not. The answer is exact rather than statistical: every loss law in the family leaves the strike with the same spectrum and ends in the same silence, so both endpoints carry precisely nothing about which of them it is. The whole difference is 41.3 decibels, and it peaks 0.38 seconds in, seven per cent of the way through the note.

Assumes: The note that gets duller as it dies · The first fifty milliseconds

Two rungs of this ladder are about the beginning of a note and one is about the end. The first fifty milliseconds found the identity of an instrument in its attack: cut the onset off a recording and listeners stop naming it, while the spectrum they hear is unchanged. The note that gets duller as it dies found a second cue at the other end, in how fast the colour drains, and pointed at the obvious next question.

If the attack carries the identity and the decay carries a second cue, what does the middle carry? For a struck instrument the honest suspicion is that there is no middle to carry anything: there is no steady state at all, only a continuous slide from one spectrum to another, and a listener may be using nothing but its two ends.

That is half a listening experiment and half a calculation, and the calculation comes first, because it decides what the experiment could possibly find. If the trajectory between the endpoints is fixed by them — if knowing where a note starts and where it finishes determines what it did in between — then a listener could not use the middle even in principle, and the experiment would be measuring nothing. So the question to settle first is how much information is in the slide at all.

A struck note's two ends are the same for every loss law. The partial levels of a string spectrum struck at 80 decibels on 130.8 hertz, and what is left of it when the fundamental itself falls under the threshold of hearing, for three laws relating a partial's decay rate to its number. The left panel is every one of them: a loss law cannot change the spectrum at the instant of the strike, because no time has passed. The other three are every one of them too: whatever the law, the note ends with nothing above the threshold. So both ends of the slide are shared, and everything that distinguishes an exponent of 0.5 from an exponent of 1 from an exponent of 2 is in the middle.
Fig. 1 The audible spectrum of a struck string at the strike and at the moment its fundamental falls under the threshold of hearing, for three laws relating a partial’s decay rate to its number. The left panel is all three of them. So are the other three.

The endpoints are the same object

The answer is not a small number. It is exactly zero, and it is zero for a reason that can be stated in one line each way.

At the strike, no time has passed. A loss law is a statement about how amplitude changes with time, so at t=0t=0 it has had no opportunity to do anything. The spectrum of a struck string at the instant of the strike is 78, 72, 69, 66, 64, 63, 61 and 60 decibels at its first eight partials, and that is the same list whether the loss rises as the square root of partial number, in proportion to it, or as its square.

At the end, the note is over. Every partial has fallen under the threshold of hearing, including the fundamental, and a spectrum with nothing above threshold in it is silence. That is also the same for every loss law, because every one of them takes the fundamental under eventually and the note’s life is defined as the moment it does.

So the family of struck notes generated by varying the one parameter that governs the whole spectral collapse is a family that shares both endpoints exactly. Not approximately, not to within a measurement error: identically, by construction. Whatever distinguishes an exponent of 0.5 from an exponent of 2 is in the middle, because there is nowhere else for it to be.

That is the strongest form the answer to this ladder’s question could have taken, and it is stronger than the question expected. The suspicion the third rung recorded was that the middle might be redundant. It is the opposite of redundant: it is the only part of the note that carries the parameter at all.

How much is in there

Zero information at the endpoints is a statement about the ends. It says nothing about whether the difference in between is large enough to matter, and that is a separate number.

Two notes that agree at both ends and nowhere in between. The distance in decibels between the audible spectra of two string notes struck identically on 130.8 hertz at 80 decibels, differing only in how fast the loss rises with partial number — an exponent of 0.5 against 2. The curve is pinned to zero at both ends by construction: no time has passed at the strike, and both notes end in silence. It reaches 41.3 decibels after 0.38 seconds, which is 7 per cent of the way through a note lasting 5.68 seconds. A listening test that presents only the attack, or only the tail, is presenting the part of the note on which these two agree exactly.
Fig. 2 The distance in decibels between the audible spectra of two struck strings differing only in their loss law. The curve is pinned to zero at both ends by the argument above. In between it reaches 41.3 decibels.

Forty-one decibels is an enormous separation for two sounds that are the same at both ends. For comparison, the whole distance a single note’s spectrum travels from strike to silence is sixty decibels, so two notes in this family get two-thirds of the way to being as different as a note is from its own death — and then converge again.

The peak is at 0.38 seconds, which is 6.7 per cent of the way through a note lasting 5.68 seconds. That is a specific and slightly awkward moment. It is well past the attack — the fifty milliseconds the second rung of this ladder is about are long gone — and it is nowhere near the middle in any ordinary sense of the word. The informative part of a struck note’s slide is the first tenth of it, immediately after the transient and before anything a musician would call the note’s body.

Why the middle is not a line

There is a second sense in which the endpoints might have determined the middle, weaker than the first and worth checking separately. Even if the two ends do not say which trajectory a note is on, they might still say roughly what any trajectory does — if every route between them is close to the straight line joining them, then interpolating would be nearly right whichever note it was.

The slide is not a straight line between its ends. How far the spectrum still has to travel to reach its terminus, in decibels, from the strike to the moment the fundamental goes under. The straight line is what a listener interpolating between the two endpoints would assume. Every real trajectory sits well below it: the spectrum does most of its moving early and then keeps a nearly fixed shape while the level falls. At an exponent of 0.5 the route is 91 decibels long against a 60-decibel gap between the ends; At an exponent of 1 the route is 117 decibels long against a 60-decibel gap between the ends; At an exponent of 2 the route is 153 decibels long against a 60-decibel gap between the ends. Half the journey is over in the first 27, 21, 26 per cent of the note.
Fig. 3 How far the spectrum still has to travel to reach its terminus, against time, with the straight line an interpolating listener would assume. Every real trajectory sits well below it, and the route is longer than the gap between the ends by half again to two and a half times.

It is not close. The straight-line distance between the two ends is 60 decibels and the route actually taken is 91 decibels at an exponent of 0.5, 117 at 1, and 153 at 2 — so the journey is 1.5 to 2.6 times longer than the gap it covers. A trajectory that wanders that far off the chord between its ends is not a trajectory a listener could reconstruct from them.

The shape of the wandering is consistent across the family and it is the same in every timbre tried. The spectrum does most of its moving early: half the distance to the terminus is covered in the first 20 to 28 per cent of the note, and the first half of the note carries 63 to 81 per cent of the total path length. What happens afterwards is that the shape stops changing much and the level goes on falling — which is very nearly the steady state the ladder said a struck note does not have, arriving late and lasting for most of the note.

That is a mild correction to the framing of the question rather than to its answer. A struck note does have something like a steady state; it simply arrives after the interesting part is over.

The other timbres, which agree

None of this is special to a string spectrum, and it is worth checking because the argument leans on the endpoints being shared, which is a claim about a construction rather than about a sound.

Two notes that agree at both ends and nowhere in between. The distance in decibels between the audible spectra of two bell notes struck identically on 130.8 hertz at 80 decibels, differing only in how fast the loss rises with partial number — an exponent of 0.5 against 2. The curve is pinned to zero at both ends by construction: no time has passed at the strike, and both notes end in silence. It reaches 45.1 decibels after 0.27 seconds, which is 5 per cent of the way through a note lasting 5.65 seconds. A listening test that presents only the attack, or only the tail, is presenting the part of the note on which these two agree exactly.
Fig. 4 The same computation on a bell’s spectrum, which has no even partials at all. The endpoints are shared for the same two reasons and the divergence in between is larger, at 45.1 decibels, because there is further between the surviving partials.

A bell reaches 45.1 decibels of separation at 0.27 seconds; a clarinet spectrum reaches 37.3 at the same moment. Both are the same shape of curve and both peak in the first five per cent of the note.

The ordering of those three numbers is the one thing in this section that is not obvious in advance, and it turns out to follow from how far apart the surviving partials are. A bell’s spectrum has no even partials, so what is left when the top of it has gone is more widely spaced than a string’s, and a difference measured over the partials present in either note is larger when there are fewer of them and they are further apart. A clarinet’s spectrum is the opposite case: its even partials are present but tiny, so a great deal of what it starts with is already close to the threshold and leaves early, which compresses the range the two notes have to differ over.

None of that moves the argument, and it is worth saying why not. The argument is about the ends, and the ends do not depend on the spectrum at all.

A struck note's two ends are the same for every loss law. The partial levels of a clarinet spectrum struck at 80 decibels on 130.8 hertz, and what is left of it when the fundamental itself falls under the threshold of hearing, for three laws relating a partial's decay rate to its number. The left panel is every one of them: a loss law cannot change the spectrum at the instant of the strike, because no time has passed. The other three are every one of them too: whatever the law, the note ends with nothing above the threshold. So both ends of the slide are shared, and everything that distinguishes an exponent of 0.5 from an exponent of 1 from an exponent of 2 is in the middle.
Fig. 5 The endpoint panels for a spectrum with almost no even partials. The attack is 78, 51, 72, 48, 69, 44, 66 and 44 decibels — a very different object from a string’s — and it is still identical across the three loss laws, and the note still ends in silence.

The differences between timbres are differences in how far apart the family gets, not in whether the endpoints are shared. The endpoints are shared for every spectrum, because the argument for it does not mention the spectrum.

Where the forty-one decibels comes from

A separation that large between two sounds sharing both ends is worth taking apart, because the obvious worry is that it is an artefact of how the distance is defined rather than a fact about the sounds.

It is not, and the reason is visible in what each note is doing at 0.38 seconds. At an exponent of 2 the upper partials are gone: the eighth partial’s decay rate is sixty-four times the fundamental’s, so it has fallen sixty-four times as far, and the note is already very nearly a sine. At an exponent of 0.5 the eighth partial’s rate is only 2.8 times the fundamental’s, so nearly the whole spectrum the note started with is still there. One note is a fundamental with a trace of colour on it and the other is still the struck string it was at the strike, and the level of the two is within a decibel — because the fundamental carries most of the power and the fundamental’s own decay does not depend on the exponent at all.

So the two notes are equally loud and completely different in colour, which is exactly the pair of properties that makes a difference audible rather than merely present. A difference that appeared as a level difference could be attributed to how hard the note was struck; this one cannot be.

The top of a struck note's series falls while the note lasts. The highest partial still above the threshold of hearing, against time, for a note struck at 80 decibels on a fundamental of 130.8 hertz with a 1/n spectrum and a loss rising as the partial number to the power 1. It starts at partial 152 and it is falling from the first millisecond. The horizontal lines are the three tops computed from frequency alone, all of which assume a note that never ends: the note's own top drops past the difference-limen top after 0.03 seconds, past the semitone top after 0.32, and past the resolvable top after 0.64. After that the series is shorter than the ear could have resolved, and what stops it is the clock.
Fig. 6 The same collapse counted rather than measured: the highest partial still above the threshold of hearing, moment by moment, on the same C3. The number falls fastest exactly where the two loss laws are furthest apart, and by six-tenths of a second there is almost nothing left to differ about.

Counting partials instead of measuring decibels gives the same window from the other direction. A note struck at 80 decibels on C3 has a hundred and fifty-two partials above threshold at the strike and eight after two-thirds of a second, and the rate at which that number falls is what the loss law is. Two exponents are two different curves through that count, and they can only differ while the count is large: once both notes are down to a handful of low partials there is nothing left to be different about, which is the analytical reason the divergence curve comes back to zero rather than merely tending to it.

That also explains why the peak sits so early. The count falls steeply and then flattens, so the interval during which two loss laws can produce visibly different spectra is the interval during which there are still partials to lose — and on a struck string at this pitch that interval is a few hundred milliseconds.

What follows for the listening experiment

The third rung recorded the next question as a listening experiment, and it was right to. What the arithmetic above does is constrain the experiment rather than replace it, and the constraints are specific enough to be worth writing down.

A test that presents the endpoints cannot work. Two notes from this family, played to a listener as an onset plus a tail with the middle removed, are the same stimulus. Not similar: the same, in the model. Any protocol that gates a note down to its first fifty milliseconds and its last half second is measuring nothing about the loss law, and would report a null that means only that the experiment was blind.

A test that gates the note at a fixed time should peak near 0.38 seconds. That is where the family is furthest apart, and it is the same for every timbre tried to within a tenth of a second. A gate at 100 milliseconds captures a small fraction of the available difference; a gate at two seconds captures a shrinking one.

And the attack is in the way. The peak sits just outside the fifty-millisecond window that carries an instrument’s identity, so a stimulus long enough to contain the informative part of the slide also contains the transient, and the two cues cannot be presented separately by truncation. Separating them needs a synthetic stimulus: the same attack on two different loss laws, which this collection can construct and a recording cannot supply.

A note gets duller as it dies. Each partial of a string note against time, with the loss rising as the partial number to the power 0.5 — so the fundamental takes 6 seconds to fall sixty decibels and the 8th takes 2.12. The heavy line is the power-weighted centroid, falling from partial 1.77 toward the fundamental; it is halfway there after 0.41 seconds. A single-rate envelope would draw all of these as parallel lines and the centroid as a horizontal one, and a struck string does neither: what is left at the end of a long note is very nearly a sine.
Fig. 7 The trajectory itself, in the drawing an earlier essay used: each partial with its own decay, at the shallow end of the plausible range. Everything on this page is a statement about the space between the left-hand edge of this figure and the right-hand one.

Which computation produced the numbers

The spectrum a listener has is the set of partials above the threshold of hearing at the stated playing level — this collection’s own ISO 226 curve, evaluated at each partial’s frequency. A partial under it is dropped rather than kept small, which is the convention the roughness of an audible spectrum uses and the reason a note has a finite life here at all.

The distance between two such spectra is the root-mean-square difference in decibels over the partials present in either, with an absent partial scored at its own threshold rather than at minus infinity. That is a compromise and it is the one place a different convention would move the numbers: scoring absences at minus infinity makes every distance infinite, and scoring them at zero makes a partial’s disappearance free. Scoring them at threshold says a partial that has just gone is worth what it was worth when it left.

The note is a struck string at C3, 80 decibels at the fundamental, with a sixty-decibel time of six seconds. Its life on those numbers is 5.68 seconds, which is the moment the fundamental itself goes under.

Where the model stops

One parameter is varied and a real instrument varies several. The family here differs only in the loss exponent. Two real instruments differ in that, in their attack, in their initial spectrum and in their decay time, and the endpoint argument covers only the first. Two notes with different attack spectra are distinguishable at the strike, trivially.

And the terminus is a convention. Ending the note when the fundamental crosses the threshold is a choice; ending it forty decibels down instead would leave a little of the fundamental audible and would make the second endpoint slightly informative. The choice was made to be conservative in the direction that would weaken the argument, and the argument survives it: at any terminus late enough for the upper partials to have gone, the surviving spectrum is a fundamental and nothing else, whatever the exponent.

There is no room in it. A hall’s own decay is faster in the treble and it acts on what the string sends it, so a listener in a hall receives a trajectory that is the composition of two. The two act on different timescales, so the composition is nearly sequential, but the room’s contribution is not zero and it is not in these numbers.

What the picture cannot show

It cannot show what a listener attends to. A difference of 41 decibels between two spectra is a difference in a physical description, and the mapping from that to what somebody notices is exactly the thing this page cannot supply. The smallest change in a spectrum a listener can detect is a separate ladder and its units are not these.

Nor a note in a texture. Every note here is struck into silence and left alone. In music a struck note is one of several, all at different stages of their own collapse, and the 0.38-second window this page identifies is short enough that a following note lands inside it. At a moderate tempo with the dampers up, four or five notes arrive inside the window that carries the difference this page is about.

It cannot show the exciter either. Where a hammer lands and how long it stays decide the spectrum at the strike, and the strike point silences a partial outright while a hammer’s own compliance rolls the top off. Both of those act on the left-hand endpoint, which is the one the argument here says is shared — so a change of exciter moves the shared endpoint rather than breaking the sharing. The family is still a family; it starts somewhere else.

And it cannot show the pedal. Lifting a piano’s dampers puts every other string in sympathy with the one struck, which adds partials rather than removing them and does so on a different timescale again. The endpoint argument survives it — a sympathetic string decays too — but the trajectory in between is not the one drawn here.

It also cannot show a note whose partials do not decay independently. A piano’s two or three strings to a note are joined at the bridge, and the pair has normal modes rather than two separate lives, one of which drives the bridge and dies while the other rings on. That puts a second decay rate into every partial, which is a structure the family drawn here does not have. It does not touch the endpoint argument — a note with two rates still starts at its own spectrum and still ends in silence — but it means the real trajectory has a corner in it somewhere around the moment the aftersound takes over, and the family here is smooth.

Where this ladder goes next

Four rungs. The shape of a note is most of what an instrument is; the identity is in the first fifty milliseconds; the colour drains at a rate that is a second cue; and now the middle, which turns out to be not merely informative but the only part of the note that carries the loss law at all, by 41 decibels, with both endpoints identical by construction.

What is owed after this is the listening experiment itself, stated precisely enough to run: two synthetic notes on the same attack and the same terminus, differing only in the exponent, presented gated at a series of durations from 50 milliseconds to two seconds, with the prediction that identification rises steeply through the first four hundred milliseconds and is flat afterwards. That needs listeners, and this collection has none — which is a real limit and not a formality, because the shape of that curve is the whole result and no amount of arithmetic produces it.

What can be paid without listeners is the parameter every figure on this ladder has held at one value by having none of it at all: the note’s pitch. Nothing in the decay model has a fundamental in it — the partials are numbered rather than measured in hertz — and yet whether a partial exists at all and whether it can be heard are both facts about absolute frequency. Putting a pitch in is arithmetic, and it would say whether the collapse this ladder has been describing is a fact about struck strings or a fact about the bottom half of a keyboard.

Part 4 of 10

One essay in the series on envelope. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientBrightnessDecayEnvelopeIdentificationSpectral centroidSpectral envelopeTimbre