The middle nobody could have guessed
Assumes: The note that gets duller as it dies · The first fifty milliseconds
Two rungs of this ladder are about the beginning of a note and one is about the end. The first fifty milliseconds found the identity of an instrument in its attack: cut the onset off a recording and listeners stop naming it, while the spectrum they hear is unchanged. The note that gets duller as it dies found a second cue at the other end, in how fast the colour drains, and pointed at the obvious next question.
If the attack carries the identity and the decay carries a second cue, what does the middle carry? For a struck instrument the honest suspicion is that there is no middle to carry anything: there is no steady state at all, only a continuous slide from one spectrum to another, and a listener may be using nothing but its two ends.
That is half a listening experiment and half a calculation, and the calculation comes first, because it decides what the experiment could possibly find. If the trajectory between the endpoints is fixed by them — if knowing where a note starts and where it finishes determines what it did in between — then a listener could not use the middle even in principle, and the experiment would be measuring nothing. So the question to settle first is how much information is in the slide at all.
The endpoints are the same object
The answer is not a small number. It is exactly zero, and it is zero for a reason that can be stated in one line each way.
At the strike, no time has passed. A loss law is a statement about how amplitude changes with time, so at it has had no opportunity to do anything. The spectrum of a struck string at the instant of the strike is 78, 72, 69, 66, 64, 63, 61 and 60 decibels at its first eight partials, and that is the same list whether the loss rises as the square root of partial number, in proportion to it, or as its square.
At the end, the note is over. Every partial has fallen under the threshold of hearing, including the fundamental, and a spectrum with nothing above threshold in it is silence. That is also the same for every loss law, because every one of them takes the fundamental under eventually and the note’s life is defined as the moment it does.
So the family of struck notes generated by varying the one parameter that governs the whole spectral collapse is a family that shares both endpoints exactly. Not approximately, not to within a measurement error: identically, by construction. Whatever distinguishes an exponent of 0.5 from an exponent of 2 is in the middle, because there is nowhere else for it to be.
That is the strongest form the answer to this ladder’s question could have taken, and it is stronger than the question expected. The suspicion the third rung recorded was that the middle might be redundant. It is the opposite of redundant: it is the only part of the note that carries the parameter at all.
How much is in there
Zero information at the endpoints is a statement about the ends. It says nothing about whether the difference in between is large enough to matter, and that is a separate number.
Forty-one decibels is an enormous separation for two sounds that are the same at both ends. For comparison, the whole distance a single note’s spectrum travels from strike to silence is sixty decibels, so two notes in this family get two-thirds of the way to being as different as a note is from its own death — and then converge again.
The peak is at 0.38 seconds, which is 6.7 per cent of the way through a note lasting 5.68 seconds. That is a specific and slightly awkward moment. It is well past the attack — the fifty milliseconds the second rung of this ladder is about are long gone — and it is nowhere near the middle in any ordinary sense of the word. The informative part of a struck note’s slide is the first tenth of it, immediately after the transient and before anything a musician would call the note’s body.
Why the middle is not a line
There is a second sense in which the endpoints might have determined the middle, weaker than the first and worth checking separately. Even if the two ends do not say which trajectory a note is on, they might still say roughly what any trajectory does — if every route between them is close to the straight line joining them, then interpolating would be nearly right whichever note it was.
It is not close. The straight-line distance between the two ends is 60 decibels and the route actually taken is 91 decibels at an exponent of 0.5, 117 at 1, and 153 at 2 — so the journey is 1.5 to 2.6 times longer than the gap it covers. A trajectory that wanders that far off the chord between its ends is not a trajectory a listener could reconstruct from them.
The shape of the wandering is consistent across the family and it is the same in every timbre tried. The spectrum does most of its moving early: half the distance to the terminus is covered in the first 20 to 28 per cent of the note, and the first half of the note carries 63 to 81 per cent of the total path length. What happens afterwards is that the shape stops changing much and the level goes on falling — which is very nearly the steady state the ladder said a struck note does not have, arriving late and lasting for most of the note.
That is a mild correction to the framing of the question rather than to its answer. A struck note does have something like a steady state; it simply arrives after the interesting part is over.
The other timbres, which agree
None of this is special to a string spectrum, and it is worth checking because the argument leans on the endpoints being shared, which is a claim about a construction rather than about a sound.
A bell reaches 45.1 decibels of separation at 0.27 seconds; a clarinet spectrum reaches 37.3 at the same moment. Both are the same shape of curve and both peak in the first five per cent of the note.
The ordering of those three numbers is the one thing in this section that is not obvious in advance, and it turns out to follow from how far apart the surviving partials are. A bell’s spectrum has no even partials, so what is left when the top of it has gone is more widely spaced than a string’s, and a difference measured over the partials present in either note is larger when there are fewer of them and they are further apart. A clarinet’s spectrum is the opposite case: its even partials are present but tiny, so a great deal of what it starts with is already close to the threshold and leaves early, which compresses the range the two notes have to differ over.
None of that moves the argument, and it is worth saying why not. The argument is about the ends, and the ends do not depend on the spectrum at all.
The differences between timbres are differences in how far apart the family gets, not in whether the endpoints are shared. The endpoints are shared for every spectrum, because the argument for it does not mention the spectrum.
Where the forty-one decibels comes from
A separation that large between two sounds sharing both ends is worth taking apart, because the obvious worry is that it is an artefact of how the distance is defined rather than a fact about the sounds.
It is not, and the reason is visible in what each note is doing at 0.38 seconds. At an exponent of 2 the upper partials are gone: the eighth partial’s decay rate is sixty-four times the fundamental’s, so it has fallen sixty-four times as far, and the note is already very nearly a sine. At an exponent of 0.5 the eighth partial’s rate is only 2.8 times the fundamental’s, so nearly the whole spectrum the note started with is still there. One note is a fundamental with a trace of colour on it and the other is still the struck string it was at the strike, and the level of the two is within a decibel — because the fundamental carries most of the power and the fundamental’s own decay does not depend on the exponent at all.
So the two notes are equally loud and completely different in colour, which is exactly the pair of properties that makes a difference audible rather than merely present. A difference that appeared as a level difference could be attributed to how hard the note was struck; this one cannot be.
Counting partials instead of measuring decibels gives the same window from the other direction. A note struck at 80 decibels on C3 has a hundred and fifty-two partials above threshold at the strike and eight after two-thirds of a second, and the rate at which that number falls is what the loss law is. Two exponents are two different curves through that count, and they can only differ while the count is large: once both notes are down to a handful of low partials there is nothing left to be different about, which is the analytical reason the divergence curve comes back to zero rather than merely tending to it.
That also explains why the peak sits so early. The count falls steeply and then flattens, so the interval during which two loss laws can produce visibly different spectra is the interval during which there are still partials to lose — and on a struck string at this pitch that interval is a few hundred milliseconds.
What follows for the listening experiment
The third rung recorded the next question as a listening experiment, and it was right to. What the arithmetic above does is constrain the experiment rather than replace it, and the constraints are specific enough to be worth writing down.
A test that presents the endpoints cannot work. Two notes from this family, played to a listener as an onset plus a tail with the middle removed, are the same stimulus. Not similar: the same, in the model. Any protocol that gates a note down to its first fifty milliseconds and its last half second is measuring nothing about the loss law, and would report a null that means only that the experiment was blind.
A test that gates the note at a fixed time should peak near 0.38 seconds. That is where the family is furthest apart, and it is the same for every timbre tried to within a tenth of a second. A gate at 100 milliseconds captures a small fraction of the available difference; a gate at two seconds captures a shrinking one.
And the attack is in the way. The peak sits just outside the fifty-millisecond window that carries an instrument’s identity, so a stimulus long enough to contain the informative part of the slide also contains the transient, and the two cues cannot be presented separately by truncation. Separating them needs a synthetic stimulus: the same attack on two different loss laws, which this collection can construct and a recording cannot supply.
Which computation produced the numbers
The spectrum a listener has is the set of partials above the threshold of hearing at the stated playing level — this collection’s own ISO 226 curve, evaluated at each partial’s frequency. A partial under it is dropped rather than kept small, which is the convention the roughness of an audible spectrum uses and the reason a note has a finite life here at all.
The distance between two such spectra is the root-mean-square difference in decibels over the partials present in either, with an absent partial scored at its own threshold rather than at minus infinity. That is a compromise and it is the one place a different convention would move the numbers: scoring absences at minus infinity makes every distance infinite, and scoring them at zero makes a partial’s disappearance free. Scoring them at threshold says a partial that has just gone is worth what it was worth when it left.
The note is a struck string at C3, 80 decibels at the fundamental, with a sixty-decibel time of six seconds. Its life on those numbers is 5.68 seconds, which is the moment the fundamental itself goes under.
Where the model stops
One parameter is varied and a real instrument varies several. The family here differs only in the loss exponent. Two real instruments differ in that, in their attack, in their initial spectrum and in their decay time, and the endpoint argument covers only the first. Two notes with different attack spectra are distinguishable at the strike, trivially.
And the terminus is a convention. Ending the note when the fundamental crosses the threshold is a choice; ending it forty decibels down instead would leave a little of the fundamental audible and would make the second endpoint slightly informative. The choice was made to be conservative in the direction that would weaken the argument, and the argument survives it: at any terminus late enough for the upper partials to have gone, the surviving spectrum is a fundamental and nothing else, whatever the exponent.
There is no room in it. A hall’s own decay is faster in the treble and it acts on what the string sends it, so a listener in a hall receives a trajectory that is the composition of two. The two act on different timescales, so the composition is nearly sequential, but the room’s contribution is not zero and it is not in these numbers.
What the picture cannot show
It cannot show what a listener attends to. A difference of 41 decibels between two spectra is a difference in a physical description, and the mapping from that to what somebody notices is exactly the thing this page cannot supply. The smallest change in a spectrum a listener can detect is a separate ladder and its units are not these.
Nor a note in a texture. Every note here is struck into silence and left alone. In music a struck note is one of several, all at different stages of their own collapse, and the 0.38-second window this page identifies is short enough that a following note lands inside it. At a moderate tempo with the dampers up, four or five notes arrive inside the window that carries the difference this page is about.
It cannot show the exciter either. Where a hammer lands and how long it stays decide the spectrum at the strike, and the strike point silences a partial outright while a hammer’s own compliance rolls the top off. Both of those act on the left-hand endpoint, which is the one the argument here says is shared — so a change of exciter moves the shared endpoint rather than breaking the sharing. The family is still a family; it starts somewhere else.
And it cannot show the pedal. Lifting a piano’s dampers puts every other string in sympathy with the one struck, which adds partials rather than removing them and does so on a different timescale again. The endpoint argument survives it — a sympathetic string decays too — but the trajectory in between is not the one drawn here.
It also cannot show a note whose partials do not decay independently. A piano’s two or three strings to a note are joined at the bridge, and the pair has normal modes rather than two separate lives, one of which drives the bridge and dies while the other rings on. That puts a second decay rate into every partial, which is a structure the family drawn here does not have. It does not touch the endpoint argument — a note with two rates still starts at its own spectrum and still ends in silence — but it means the real trajectory has a corner in it somewhere around the moment the aftersound takes over, and the family here is smooth.
Where this ladder goes next
Four rungs. The shape of a note is most of what an instrument is; the identity is in the first fifty milliseconds; the colour drains at a rate that is a second cue; and now the middle, which turns out to be not merely informative but the only part of the note that carries the loss law at all, by 41 decibels, with both endpoints identical by construction.
What is owed after this is the listening experiment itself, stated precisely enough to run: two synthetic notes on the same attack and the same terminus, differing only in the exponent, presented gated at a series of durations from 50 milliseconds to two seconds, with the prediction that identification rises steeply through the first four hundred milliseconds and is flat afterwards. That needs listeners, and this collection has none — which is a real limit and not a formality, because the shape of that curve is the whole result and no amount of arithmetic produces it.
What can be paid without listeners is the parameter every figure on this ladder has held at one value by having none of it at all: the note’s pitch. Nothing in the decay model has a fundamental in it — the partials are numbered rather than measured in hertz — and yet whether a partial exists at all and whether it can be heard are both facts about absolute frequency. Putting a pitch in is arithmetic, and it would say whether the collapse this ladder has been describing is a fact about struck strings or a fact about the bottom half of a keyboard.
Part 4 of 10
One essay in the series on envelope. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Attack transientBrightnessDecayEnvelopeIdentificationSpectral centroidSpectral envelopeTimbre
- A damper cannot reach into the room brightness, decay, envelope, identification, spectral centroid
- A damper changes the clock, not the colour brightness, decay, envelope, identification, spectral centroid
- The collapse belongs to the bass brightness, decay, envelope, spectral centroid
- A doubled pizzicato gives its note away early decay, envelope, timbre
- A hammer is not an impulse attack transient, envelope, timbre
- A section has a loudest member, not a colour brightness, spectral centroid, timbre