Timbre and acoustics

The blend arrives before the note does

Nine essays on spectrum draw a steady state, and the strongest cue that two instruments are two instruments is that they do not start together. Two envelopes rising at different rates turn out to be a balance — the same dial an earlier essay swept — so the attack is that dial moved by the clock, and its whole travel is fixed at twenty times the log of the two attack times. It is six decibels for a clarinet with a violin against a crossing twelve to twenty-two decibels out, so one pair in ten changes hands during its own attack, and which one depends on a convention rather than on the instruments.

Assumes: The blend table has a row for every note · A note is heard after it starts

Nine rungs of this ladder have put instruments together and every one of them has drawn a steady state. A spectrum is a list; a mouth filters it; a body filters it; a fixed filter under a moving note makes the list a function of pitch; two lists do not commute; three lists over three notes have six arrangements; two lists on one note make a composite one of them owns; three lists make one it owns more emphatically; and the whole of that asked at every note is a different table at each end of the range.

Not one of those figures has any time in it. And the single strongest cue that two instruments are two instruments is that they do not start together. The onset ladder has the numbers — a clarinet’s attack is 45 milliseconds and a bowed violin’s is 90 — and this ladder has the composite, and neither has looked at the other.

The join turns out to be exact rather than approximate, and it is the whole of this rung. Two players on one note are two envelopes rising at different rates. At time tt the pair is at some balance — twenty times the log of one envelope over the other — and a balance is the dial the eighth rung swept when it asked at what level a composite changes owner. So the attack does not need a new quantity. The attack is that dial, turned by the clock, at every note.

The attack is the balance dial, turned by the clock. The level of a violin against a clarinet on one note at 392 hertz, moment by moment through the attack, with both players starting together. Two envelopes rising at different rates are a balance, so this axis is the same dial a conductor turns — and its whole travel is 6.02 decibels, which is twenty times the log of the ratio of the two attack times, 45 against 90 milliseconds, and nothing else. The pair does not begin as one player alone: both envelopes leave zero at the same slope ratio, so the dial starts at a finite offset rather than at silence. The dashed line is the balance at which the composite changes owner, -3.48 decibels — inside the travel, so the note belongs to a clarinet for its first 29 milliseconds and to a violin for the rest of its life.
Fig. 1 The level of a violin against a clarinet on one note, moment by moment through the attack, with both starting at the same instant. It is the same axis a conductor’s balance runs on, and its whole travel is six decibels — twenty times the log of forty-five milliseconds against ninety.

The pair never starts as one player alone

The first thing the arithmetic says is not what the debt this rung pays assumed, and it is worth getting straight before anything else.

An attack is modelled here as 1et/τ1 - e^{-t/\tau} from rest, which is what a resonator driven from rest does, with τ\tau set so that the ten-to-ninety per cent rise takes the measured attack time. Near zero that function is t/τt/\tau — a straight line out of the origin. So the ratio of two such envelopes near zero is τa/τb\tau_a/\tau_b, a finite number, and not infinity.

A doubled note does not begin as one player and acquire the other. It begins as both, at a fixed offset, and the offset closes. For a clarinet doubled by a violin the note starts with the violin six decibels under the clarinet and ends with the two level, and the whole of the attack’s contribution to the composite is that six decibels going away.

That bound is the useful thing, because it is a bound. The travel of the dial is 20log10(risea/riseb)20\log_{10}(\text{rise}_a/\text{rise}_b) and contains nothing else — not the pitch, not the spectra, not the dynamic. The widest pair this collection can price is a sung vowel with a trumpet, at 11.3 decibels. The narrowest is a violin with an organ flue pipe, at 1.6.

How late each instrument is heard. Nine measured attack times converted to a heard moment at the 6 dB below peak criterion. The bar is the lag for the family's typical attack and the line through it is the range a player can produce on that instrument — which for the bowed and sung rows is wider than the gap between several of the other rows, so the ordering is a claim about typical playing and not about any single note. The fastest here is the marimba at 0.9 ms and the slowest the sung vowel at 35 ms.
Fig. 2 The measured attack times this essay reads, converted to the moment each instrument is heard. The bar is the family’s typical attack and the line through it is the range a player can produce, which for the bowed and sung rows is wider than the gap between several of the other rows.

Those nine rows are the whole input, and five of them belong to instruments this ladder also has a spectrum for. There is no oboe in the attack table and no marimba in the spectrum table: the two were built by different ladders for different arguments, and the overlap is a violin, a clarinet, an organ flue pipe, a voice and a trumpet. Every figure below is restricted to those five, because a doubling one of whose members has no measured attack is not a doubling that can be put on a clock, and inferring the missing row from its family would put a number in a picture that no measurement supports.

Which computation produced the number

At each instant the two players’ envelopes are evaluated, their ratio in decibels is handed to doubling as a balance, and the composite comes back as the root of the sum of the squares partial by partial — because two players are independent sources and their partials add in power. That composite is compared with the one the pair settles at, in the log-spectral distance this ladder has used since its seventh rung, and it is compared with the pair’s ownership crossing: the balance at which the composite stops resembling one player more than the other.

Nothing is fitted and nothing is new. The spectra are the five radiators this ladder built out of parts it already had: a string through the measured violin body, a glottal pulse through a published vowel’s formants, an odd-dominant list cut off above a computed tone-hole cutoff, a sawtooth through a bell’s high pass, and an organ list with no filter on it at all. The envelopes are the p-centre ladder’s own measured attack times. The two halves meet in one line of arithmetic and the rest is reading the result.

One property of those spectra decides how much the tilt can do, and it is the property an instrument is not one timbre established: every filter here is fixed in frequency, so the composite of two players is a fixed shape at a given pitch and the only thing the attack varies is the weight of the two halves. A balance cannot invent a partial. Whatever the dial does, the composite’s support is the union of the two partial lists from the first millisecond, and what changes is only how the energy is divided between them — which is why the whole effect is bounded by a few decibels rather than by the distance between the two instruments.

A doubling has its colour within a few tens of milliseconds. For each of the 10 pairs both halves can be priced for, how far the composite's spectrum is from the spectrum it settles at, in log-spectral decibels, through the attack of one note at 392 hertz. Every curve starts at a finite height and falls: the largest is a voice on “hod” with a trumpet at 7.0 decibels, the smallest under half a decibel, and the average pair is within one decibel of its final colour after 22 milliseconds. A doubling is therefore not a spectrum sweeping through the note — it is a spectrum that arrives almost at once and then holds, and what a listener has for all but the first fiftieth of a second is the steady-state composite these nine essays have been drawing.
Fig. 3 How far each pair’s composite still is from the colour it settles at, through the attack. Every curve starts at a finite height and falls; the average pair is within a decibel of its final spectrum after twenty-two milliseconds.

The largest displacement at the first instant is 7.0 decibels, for a voice doubled by a trumpet — two instruments whose attacks differ by a factor of nearly four and whose spectra have almost nothing in common. The smallest is under half a decibel. The average pair is within one decibel of its settled colour after 22 milliseconds, the slowest after 82.

So the picture the debt described — a spectrum sweeping from one player’s shape to the pair’s over some tens of milliseconds — is half right. The tens of milliseconds are real. The sweep is not from one player’s shape; it is from the pair’s shape at a slightly wrong balance to the pair’s shape at the right one, and the difference is a few decibels of tilt.

Nine pairs in ten never change hands

Whether that tilt matters has a sharp test, and the test is an inequality between two numbers from two different ladders that happen to be in the same unit.

One pair in ten changes hands during its own attack. Each pair's two decibel quantities on one axis: the bar is the whole travel of the balance dial through the attack, which is twenty times the log of the ratio of the two attack times, and the ring is the balance at which that pair's composite changes owner — an earlier essay's number, computed from two spectra with no clock in it. A pair changes hands during its attack exactly when the ring falls inside the bar, and 1 of the 10 do: a violin with a clarinet, at 29 milliseconds. Every other crossing sits 12 decibels or more out, which is further than any pair of attack times in this collection can travel. Rings past the right edge are pairs with no crossing at any balance at all.
Fig. 4 Each pair’s two decibel quantities together: the bar is the whole travel of the dial through the attack, the ring is the balance at which that pair’s composite changes owner. A pair changes hands during its attack exactly when the ring is inside the bar, and one of the ten is.

The travel is an onset-ladder quantity with no spectra in it. The crossing is a spectrum-ladder quantity with no clock in it. A pair changes owner during its own attack exactly when the second is inside the first, and one of the ten pairs is: a violin with a clarinet, whose crossing sits at 3.5 decibels inside a travel of 6.0, and whose note therefore belongs to the clarinet for its first 29 milliseconds and to the violin afterwards.

Every other crossing sits twelve decibels or more from level, and one pair — an organ flue pipe with a trumpet — has no crossing at any balance at all, meaning there is no level at which the quieter instrument is doing anything but colouring the louder. Twelve decibels is more than the widest travel any pair of these attack times can produce, and that pair is a sung vowel with a trumpet, whose attacks differ by a factor of nearly four. So for nine pairs in ten, the answer to whose note is this is the same answer at one millisecond and at half a second, and every steady-state figure on this ladder is telling the truth about the attack as well.

That is a refusal, and it is worth stating as one. The attack does not make a doubling into two instruments and then one. It makes it into one instrument slightly mis-balanced and then correctly balanced, in nine cases out of ten.

Two notes started together, heard 14 ms apart. Two amplitude envelopes rising from the same instant: a clarinet with a 45 millisecond attack and a bowed violin with 90. The horizontal line is the criterion — 6 dB below peak, from Vos & Rasch 1981 — and the two dots are where each envelope crosses it. Nothing about the onsets differs; the heard moments differ by 14 ms, which is why the bowed violin has to start early to be heard on the beat. The buttons play the pair as written and then with the clarinet delayed by that amount.
Fig. 5 The one pair that does change hands, drawn as the onset figures draw it: two envelopes from the same instant, and the moment each crosses six decibels below its own peak. The gap between those two moments is what a violinist plays early to remove.

Which is a fact about the convention, not about the pair

And that last picture contains the reason the whole result has to be stated twice.

Both figures above start the two players at the same physical instant, which is what a notated simultaneity says. That instant is itself a modelling choice, and a note takes a number of periods to speak is the rung that shows why: a physical onset is not an event a listener has access to, only a moment at which energy begins to arrive. It is not what an ensemble does. The p-centre ladder is entirely about the fact that a note is heard after it starts and that players compensate: the bowed and sung parts are played early so that the heard moments coincide. Under that convention the two physical onsets are staggered by the difference in the two lags, and during the stagger the slower player is sounding alone.

Whose note it is at the start depends on the convention, not the pair. Every pair of the 5 radiators both halves can be priced for, with the instrument its composite belongs to at the first instant and how long it keeps it — read twice. The upper bar of each pair starts both players at the same physical instant, which is what a notated simultaneity says; the lower one lines up their perceptual centres, which is what an ensemble does, and which staggers the two physical onsets by the difference in their lags — 4.7 to 25.3 milliseconds here. The two readings name a different first owner for 5 of the 10 pairs. Started together, 1 of them change hands during the attack; aligned on centres, 4 do, and they are a different set. Aligned on centres the first owner is the slower player every time, without exception, because it is the one that had to start early and is alone while it does.
Fig. 6 Every pair read twice: started together, and with their perceptual centres aligned. The two readings name a different first owner for five of the ten pairs, and the number that change hands during the attack goes from one to four.

Aligned on centres the composite belongs to the slower player at the first instant, for every pair without exception, because that is the player who had to start early and is alone while it does. Started together it belongs to whichever player the pair’s own crossing favours, and that comes out five and five.

So the two conventions name a different first owner for five of the ten pairs, the count that change hands during the attack goes from one to four, and they are a different set. The clarinet-and-violin pair — the only one to change hands when the two start together — does not change hands at all when their centres are aligned, because the violin starts 14 milliseconds early and never gives the note up.

Which instrument a doubled note sounds like during its attack is therefore a property of how the ensemble reads a simultaneity, and not a property of the two instruments. That is a stronger statement than the one this rung set out to make and it is the honest one: the spectrum ladder alone cannot answer the question, because the answer is not in the spectra.

In the bass there is nothing to give away

One more limit, and it comes from the onset ladder too.

In the bass a doubling has no attack to give it away. The whole travel of the balance dial for a clarinet doubled by a violin, against the pitch of the note. High up it is 6.02 decibels, the ratio of the two players' own attack times. Below 89 hertz the faster of the two can no longer be that fast — an amplitude cannot be established in fewer than 4 periods of the note itself — and below 44 hertz both are held at the same floor and the travel is exactly nought. So the register in which the fewest pairs were found to blend is the register in which the attack distinguishes them least: a low doubling is a single spectrum from its first millisecond, and a high one is two instruments for a few tens of milliseconds before it is one.
Fig. 7 The travel of the dial for a clarinet doubled by a violin, against the pitch of the note. Below 89 hertz the clarinet’s own attack is no longer available to it, and below 44 hertz both players are held at the same floor and the travel is exactly nought.

A low note cannot start on time: an amplitude cannot be established in fewer than about four periods of the note itself, so there is a floor under every attack time that rises as the pitch falls. Below 89 hertz a clarinet cannot rise in 45 milliseconds because 45 milliseconds is not four periods of the note it is playing. Below 44 hertz neither can the violin, and both are held at the same value.

A doubling in the bottom octave has a travel of exactly zero. The two players are the same instrument as far as the attack is concerned, and the composite is at its settled colour from the first millisecond.

That lands on top of the ninth rung’s own finding and points the same way. Down there only six of fifteen pairs blend by the steady-state definition and twelve do at the top of the range — so the register where blend is hardest to achieve in the steady state is the register where the attack gives least away, and the register where blending is easiest is the one where the attack is loudest about the seam. Two effects that would have been supposed to reinforce each other in fact cancel, and neither had been computed.

What the pictures cannot show

These are decibels against milliseconds and there is no listener in any of them.

The largest omission is that the attack carries cues these figures do not contain at all. A real attack has noise in it — bow scratch, breath, the thud of a hammer’s felt — and that noise is not in any spectrum here, which is a partial list at whole multiples of a fundamental. It has inharmonic transients that die before the steady state begins. It has the attack’s order: on many instruments the low partials arrive before the high ones, so the spectrum is not merely mis-balanced during the attack, it is incomplete. Any of those could be a stronger separation cue than the balance measured here, and none of them is in the model.

The envelope model itself is one exponential per instrument. An attack time is not an attack showed that the shape matters as much as the duration and that a bowed note’s own envelope is not exponential — so the travel computed here is right about its size and approximate about its course. And the measured attack times carry ranges wider than several of the gaps between them: a sffz violin attack and a dolce one differ by a factor of four, which is more than the whole travel of most of these pairs. The ordering in the table is a claim about typical playing, not about any single note.

Finally, the ownership crossing is a log-spectral distance, which is a measure of resemblance and not a measurement of anything anybody has been asked. What it licenses is a comparison — this composite is nearer that player’s colour than the other’s — and a claim that the comparison does or does not change during the attack. Whether a listener notices the change is a question for a listener.

Whose music, and when

The practice this is about is orchestral unison doubling in the European repertoire from Berlioz onward, and the specific advice it prices is the one every treatise gives: that a doubling blends better if the two players match their attacks.

The arithmetic says that advice is aimed at the right thing and is worth less than it sounds, in a way that has a number. Matching the attacks shrinks the travel of the dial, and the travel is already small — six decibels for a clarinet with a violin, against a crossing three and a half decibels out. Halving the difference in the attack times buys six decibels of travel down to three, which takes that one pair out of the changing set and leaves the other nine exactly where they were.

What the advice is really aimed at, on this reading, is the second figure of the pair above: the stagger. Under the alignment an ensemble actually uses, the slower player is alone for four to twenty-five milliseconds, and matching the attacks removes that window rather than shrinking a balance. A conductor asking two players to match their articulation is asking them to remove a stagger they created by being punctual, and the size of what they remove is the difference in their perceptual centres — which is a quantity the onset ladder measures and this one does not.

And in the bottom octave the advice is unnecessary. Two players on a low note have the same attack whatever they do, because the note’s own period has taken the choice away from them.

There is a scale on which none of this is worth acting, and the collection supplies it. Twelve violins are more punctual than one measured what an ensemble’s own scatter does to a section’s arrival, and a real orchestral section’s onsets are spread by more than the differences priced here. So the stagger is a genuine quantity for two soloists and is under the noise for two sections, and a doubling in a full orchestra is a doubling whose attack window is set by how well the two desks are together rather than by what instruments they are holding. The claim these figures license is about a chamber texture and a pair of players; it is not a claim about a tutti.

That places the result beside the one an orchestrator doubles a line reached from the other side. There the doubling was found to be nearly free to hold across a phrase because the objective is carried by the bass assignment; here it is found to be nearly free in time because the composite settles in twenty milliseconds. Both are the same shape of answer: the doubling is a small decision made once, and the arithmetic keeps finding it smaller than the treatises do.

Where this ladder goes next

Ten rungs. A spectrum is a list; a mouth filters it; a body filters it; a fixed filter under a moving note makes the list a function of pitch; two lists do not commute; three lists over three notes have six arrangements; two lists on one note make a composite one of them owns; three lists make one it owns more emphatically; the whole of that asked at every note is a different table at each end of the range; and now the clock under it, which turns the same dial the eighth rung swept and turns it about six decibels.

What is owed after this is the release. Every rung of this ladder now has an attack and a steady state and no ending, and an ending is not the attack run backwards: a bowed note stops when the bow does, an organ pipe stops when the valve shuts and the room takes over, and a struck note does not stop at all — it decays with a different rate for every partial, which is a thing this collection has already computed for one instrument. So a doubling of a struck instrument by a sustained one is a composite whose balance walks for the whole length of the note rather than settling in twenty milliseconds, and it walks in the direction that ends with the sustained player owning it outright. What that would settle is whether the pizzicato-and-flute or harp-and-clarinet doublings that fill orchestral scores have an ownership that is stable at all, or whether they are a handover with a duration — which is the same question this rung asked about the attack, asked over a window a hundred times longer, where the travel of the dial is not bounded by anything. It needs arithmetic and nothing else, and the first number to compute is the travel: for a struck note against a held one it is not a ratio of attack times but a decay in decibels per second, so the bound this rung found does not apply and there is no reason to expect the answer to be small.

Part 10 of 13

One essay in the series on spectrum. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientDoublingEnvelopeOnsetPerceptual-centreSource-filterSpectrumTimbre