Form and structure

Who plays what and how loud is one question

Two lines of argument, one about spectrum and one about loudness, each stopped at the same wall and each said so. One of them can choose who plays which note and has every player at the same level; the other can choose how loud each part is and has nobody assigned to anything. Put together they are a single problem with two kinds of variable, and solving it in stages picks a different answer from solving it at once — nine per cent rougher, at the same loudness, on an ordinary triad.

Assumes: Which player on which note · The dynamics are in the score already

Two essays on this site end by naming the same missing thing, from opposite sides of it, four essays apart.

Which player on which note enumerates the six ways of giving three players three notes and ranks them by how rough each arrangement is. Every player in it is at the same level, because the roughness model it uses takes a spectrum normalised to one and has nowhere to put a decibel. Its last paragraph says so: the balance question is the other half, and joining the two would make an orchestration model rather than an enumeration.

The dynamics are in the score already computes how loud a passage is from the notes on the page — count the parts, put every partial in its critical band, sum, smooth. Nobody in it is playing anything in particular. Its last paragraph says so: the joint problem is a discrete choice and a continuous one at once, the spectrum ladder has the discrete half, this rung has the continuous half, and neither has ever been handed to the other.

This is the two of them handed to each other.

Every assignment, at equal levels and at its own best balance. The 6 ways of putting 3 players on a chord, each drawn twice: hollow at equal levels, which is what an assignment ranking sees, and filled at the levels that balance the parts and then minimise roughness. Every scoring here is at the same total loudness, 23.3 sones, so two points are comparable. Solving the discrete problem first picks violin · clarinet · oboe; solving both at once picks violin · oboe · clarinet, and the two-stage answer costs 9.0 per cent more roughness. The orderings do not keep their places between the two columns, which is the whole of the argument: a ranking taken at equal levels is not a ranking.
Fig. 1 Every way of putting a clarinet, a violin and an oboe on a C major triad, drawn twice. The hollow point is where the arrangement sits with all three players at one level, which is what an assignment ranking sees. The filled point is where it sits once each player has been given the level that balances the parts. Every scoring in the figure is at the same total loudness, so the two columns are comparable — and the orderings do not keep their places between them.

The two halves, and why they are different kinds of thing

The discrete half is a permutation. Three players and three notes admit six arrangements, and there is nothing between them: an oboe is either on the top note or it is not. The quantity that ranks them is the roughness of the resulting sound, computed with the Plomp–Levelt model this site has used since roughness can be computed — a sum over every pair of partials of a function of how far apart they are in critical bands.

The continuous half is a vector of levels. Each player can be louder or quieter by any amount, and the quantity that constrains them is loudness in sones, which is not a level: it is compressive, it groups partials that fall inside one critical band, and it is what a chord is not as loud as its notes is about.

These do not combine by adding. A roughness is in pascals squared and a loudness is in sones, and any weighted sum of the two smuggles in an exchange rate that nobody has measured. The progression ladder met exactly this problem putting a cadence count beside a key-profile correlation, and it recorded that the commensuration had been moved rather than removed.

So the shape used here is a constrained optimisation rather than a weighted one, and the constraint is the thing a conductor actually says. Balance the parts is a requirement: each part is to carry its share of the loudness, within a stated tolerance. Make it less harsh is what is left over. There is no exchange rate because there is no sum.

variables    a permutation, and a level for each player
constraint   each part within 4.5 percentage points of an equal
             share of the loudness, and the whole scoring at one
             total loudness
objective    roughness

The second half of the constraint is what makes the figure above readable. Roughness is quadratic in pressure, so any comparison that let one candidate be quieter than another would simply find the quiet one. Every scoring in every figure here is solved for the common gain that puts it at one total loudness — by bisection, because sones are not a power of pressure and there is no formula for it.

Six ways to put three players on three notesThe same chord — C4, E4, G4 — played by clarinet, violin, oboe in all 6 possible assignments, scored by the roughness each produces. Every bar is the same pitches and the same instruments; only who is on which note changes. The worst is 1.05 times the best, which is a factor a score can control and a chord symbol cannot express at all. Each row is labelled from the bottom note upward.C4 clarinet · E4 oboe · G4 violin×1.00C4 clarinet · E4 violin · G4 oboe×1.01C4 violin · E4 clarinet · G4 oboe×1.01C4 violin · E4 oboe · G4 clarinet×1.02C4 oboe · E4 violin · G4 clarinet×1.03C4 oboe · E4 clarinet · G4 violin×1.0500.050.10.150.20.25total roughness over the three pairs
Fig. 2 The discrete half by itself, which is an earlier essay on spectra: the same six arrangements ranked with every player at one level. This is the picture the joint solve has to be compared against, and the comparison is the essay. Nothing here is wrong; it is answering a question with one of its variables held down.

The continuous half by itself is the loudness ladder’s: the same chord at five registers, with what its notes are worth in sones against what they would be worth if loudness added like power. It does not add like power, and the shortfall is largest at the bottom — which is the term the joint solve is trading against the discrete one.

What the answer is

For a C major triad with a clarinet, a violin and an oboe, solving the discrete problem first gives violin, clarinet, oboe from the bottom up. Solving both at once gives violin, oboe, clarinet, and the two-stage answer is nine per cent rougher at the same loudness.

Nine per cent is not a large number and it is not meant to be. What matters is that it is not zero, because zero is what a stage-by-stage method assumes. The reason it is not zero is visible in the second figure below: the three pairs of players do not contribute equally to the roughness, and they do not contribute in proportion to their loudness either. The violin against the oboe supplies about a third of it, the oboe against the clarinet about two fifths, and the violin against the clarinet the rest — so the level that makes the sound smoothest is the level that quietens the pair that is doing the damage, and which pair that is depends on who is where.

An equal-level ranking cannot see that, because at equal levels every pair is weighted by its spectra alone. Give the levels back and each pair acquires a second weight that the assignment itself decides.

One scoring, in decibels, in sones and in pairs. The joint answer for this chord: a violin on C4 at 71.1 dB, carrying 34.3 per cent of the loudness; an oboe on E4 at 63.6 dB, carrying 29.9 per cent of the loudness; a clarinet on G4 at 69.6 dB, carrying 35.8 per cent of the loudness. The balance requirement is that each part carry 33.3 per cent of the loudness within 4.5 percentage points, and this scoring meets it to 3.46. The three pairs contribute very unequally to the roughness — a violin against an oboe 36 per cent, a violin against a clarinet 23 per cent, an oboe against a clarinet 41 per cent — which is why the levels come out unequal: the loudness is shared evenly and the roughness is not shareable at all.
Fig. 3 The joint answer taken apart. Above: what each player is given, in decibels, and what that buys in loudness — the three shares are within a few points of a third each, which is the constraint being met. Below: where the roughness actually is, by pair. The loudness is shared out almost evenly and the roughness is not shareable at all, and that mismatch is the whole reason the levels come out unequal.

The answer that is wrong is also the answer that will not stay still

Running the same solve at five pitches produces the finding this rung is really for.

The joint answer is the same ordering at four of the five registers and changes once. The two-stage answer changes three times. And where the two agree — at the bottom of the range and at the top — the penalty is exactly zero, so a method that is checked only at the extremes looks perfect.

At the top register it is worse than perfect and worse than wrong. The assignment the equal-level ranking chooses at C6 cannot be balanced at all: its closest achievable share error is 4.8 percentage points against a tolerance of 4.5, so five of the six orderings admit a balance and the one the first stage picks is the sixth. A two-stage method there does not produce a rougher scoring; it produces a scoring the conductor’s own instruction cannot be carried out on, and it does so having reported success at the stage that made the choice.

That is the sharpest form of the argument this rung is making. The two halves do not merely interact — the discrete stage can select an assignment for which the continuous stage has no solution, and nothing in the discrete stage can tell.

The tolerance, which is asserted

Four and a half percentage points is a stated number with no measurement behind it, and the penalty depends on it in a shape worth having:

balance tolerance the two-stage penalty
1.0 points 0% — the two agree
2.0 7.6%
4.5 9.0%
6.0 12.8%
8.0 10.5%
12.0 6.7%
20.0 1.3%

The penalty is non-monotone and it peaks in the middle. At a very tight tolerance neither method has any freedom, so both are forced to the same levels and there is nothing to lose by staging. At a very loose one the two-stage method rescues itself, because with a slack constraint almost any assignment can be re-levelled close to the joint optimum. The cost of solving in stages exists only where the constraint is tight enough to bind and loose enough to leave choices — which is where a real balance instruction lives, and is why the effect is not an artefact of the number chosen.

The joint answer, meanwhile, is the same ordering at every tolerance from two points to twenty. So what the tolerance moves is how much the two-stage method loses, and not what the right answer is.

That is a worse property than being wrong by nine per cent. A rule of thumb that is wrong by a fixed amount can be corrected; one that is right at the edges and wrong in the middle of the range, by an amount that changes sign nowhere and simply appears, cannot be. It is also the shape a chord is a register found for roughness itself: the quantity moves by a factor of five across the compass, so any claim about a chord that does not say where it is has left out most of its own variance.

What the two-stage answer costs, register by register. The same chord at 5 pitches. Up the page is the roughness of the assignment-first answer over the roughness of the joint one, both at their own best balance and both at one total loudness, so a value of one means the two agree. They disagree at 2 of 5: C3 0.0%, G3 0.0%, C4 9.0%, G4 2.4%, C5 0.0%. The joint answer is 2 distinct orderings across the range and the two-stage answer is 3, which is the finding: the answer that is wrong is also the answer that keeps changing.
Fig. 4 The cost of solving in stages, at five pitches. A value of one means the two methods pick the same scoring. They agree at the outside of the range and disagree in the middle, and the letters under the axis are the two answers: the joint one settles on two orderings across the whole range while the two-stage one uses three. The instability is in the method rather than in the music.

The doubling question can be asked at every dynamic at once, which is the form in which it stops being a rule of thumb.

The doubling ranking, asked again at every dynamic. Every four-part voicing of the chord inside the four standard compasses — 480 of them — grouped by which chord member is doubled, and each group's mean roughness drawn as a multiple of the smoothest group's at that level. The absolute roughness rises by a factor of 1.0e+6 across the range drawn and the ranking does not move at all: root 1.000, fifth 1.013, third 1.090 at the quietest, and 1.000, 1.013, 1.090 at the loudest. There are no inversions: not one pair of doublings changes places.
Fig. 5 Every four-part voicing of the chord inside the four standard compasses — 480 of them — grouped by which chord member is doubled, with each group’s mean roughness as a multiple of the smoothest group’s at that level.

The absolute roughness rises by a factor of a million across the dynamic range and the ranking between the doubling choices barely moves, which is the same shape of answer this ladder keeps producing: the quantity is enormously level-dependent and the comparison is not. A rule about which note to double survives being stated without a dynamic; a rule about how rough a chord is does not.

What a tolerance is doing in the constraint

The balance requirement has a number in it — four and a half percentage points — and a number like that in a model is usually a place where the answer has been chosen rather than computed. It is worth saying exactly what it does here, because it does something unusual.

It is not a fudge factor traded against the objective. It is the width of the feasible set, and the feasible set is what the objective is minimised over. Tighten it to nothing and there is exactly one level vector per ordering that meets it, so the discrete problem is all that is left and the two-stage method becomes correct by construction. Widen it and eventually every ordering can be balanced any way it likes, at which point the constraint stops constraining and the solver quietens whatever is rough until only the total-loudness requirement holds it up.

So the tolerance is the amount of latitude a player has, and the interesting fact about it is that the two-stage penalty exists across a broad middle of that range rather than at one setting of it. A conductor who insists on an exact balance and a conductor who does not care are both solving a problem in which the assignment can be chosen first. Everybody in between is not.

The same solve on a minor triad with a different trio gives a different winner and the same shape of result, which is the check that none of this is a property of one chord.

Every assignment, at equal levels and at its own best balance. The 6 ways of putting 3 players on a chord, each drawn twice: hollow at equal levels, which is what an assignment ranking sees, and filled at the levels that balance the parts and then minimise roughness. Every scoring here is at the same total loudness, 22.6 sones, so two points are comparable. Solving the discrete problem first picks voice on “hod” · violin · flue pipe; solving both at once picks voice on “hod” · violin · flue pipe, and the two-stage answer costs 0.0 per cent more roughness. The orderings do not keep their places between the two columns, which is the whole of the argument: a ranking taken at equal levels is not a ranking.
Fig. 6 The same question asked of a G minor triad with a voice, a flue pipe and a violin. The players are deliberately unalike — one has formants, one has a fixed dynamic, one has a measured body resonance — and the arrangement that wins at equal levels is again not the arrangement that wins once the levels are free.

Which computation produced the numbers

The spectra are radiatedPartials, which is the spectrum ladder’s own machinery and has been on this site since which instrument is underneath: a source this collection already carries through a filter it already carries. The string’s 1/n list through the measured violin body; the site’s own odd-dominant clarinet list cut off above the note where the holes stop working; a reed list cut off higher for the oboe. Nothing is fitted and nothing is tabulated from a recording.

One thing had to change, and it is the kind of change that is easy to get wrong. radiatedPartials normalises each spectrum to its own strongest partial, which is exactly right for asking which of two spectra is rougher and exactly wrong for adding two instruments together — it would make every player equally loud by construction, which is the assumption being removed. So each spectrum is renormalised to unit energy before its level is applied, in the way partial levels have always been computed here, and only then is each partial tested against the threshold of hearing.

The roughness is dissonancePair with pressures instead of normalised amplitudes, summed over every pair of players rather than one pair. The unit is pascals squared. The loudness is bandLoudness, which groups components inside a critical band before converting to sones, and it is the same function the loudness ladder’s fourth rung uses to read a dynamic curve off a page.

The continuous search is over relative offsets from minus nine to plus nine decibels in steps of a decibel and a half, with the overall gain solved for rather than searched: for each relative vector, twenty bisection steps find the common level that hits the target loudness. That is what makes the total-loudness constraint exact instead of another axis to search, and it means every candidate in every figure is at the target loudness by construction — only the shares can fail.

What each of them radiates at 262 hertz. The partial amplitudes each radiator actually puts into the air at 262 hertz, normalised to its own strongest partial. Every one is a source already to hand through a filter already to hand — the string through the measured violin body, the glottal pulse through a published vowel, the odd-dominant list cut off above a woodwind's computed tone-hole cutoff. The filters do not move with the note, which is why these pictures are different at every pitch and why a duet is not symmetric.
Fig. 7 The spectra the whole calculation runs on, at middle C. The filters do not move with the note, which is why the answer changes with register: transposing a chord slides three fixed filters along three moving sets of partials, and which pairs nearly coincide is different at every pitch.

Where the model stops

There are three players and one chord. A real orchestra has sixteen sections and a passage, and the number of assignments grows as a factorial. Nothing here says the search scales; what it says is that the two halves interact, which is a statement about the problem rather than about the algorithm, and it is a statement that does not get better with more players.

The balance requirement is equal shares, and no repertoire asks for that. A melody is meant to be on top, and what a dynamic mark actually specifies is a place in an ordering rather than a level. The constraint takes a target vector, so an unequal target is one argument away, and the arithmetic is unchanged — but every number quoted here is for the flattest requirement there is, which is the one with the fewest assumptions in it and not the one a score would state.

Nothing in it is about time. The scoring is one chord held steadily. The loudness ladder’s third rung is entirely about the fact that loudness is a running impression with a two-second release, and none of that is here. A passage’s balance is a trajectory and this is a point on one.

And the spectra are shapes with levels in front of them. That is the assumption the third rung of this anchor removes, and it is not a small one: a brass player blowing harder does not scale a waveform, and neither does a singer. Every number above is computed as though they did.

What the picture cannot show

It cannot show the levels as a surface. The figure draws the best feasible point for each ordering, which is a minimum over a two-dimensional set, and the set has a shape — some orderings have a wide floor where several balances are nearly as good, and some have a narrow one. A conductor’s experience of an arrangement being forgiving is presumably about that shape, and one point per ordering discards it.

Nor can it show what is infeasible. Every ordering here can be balanced, because three instruments of comparable power on three notes of one triad always can be. The interesting orchestration problems are the ones where a part cannot be brought up to its share without being played harder than its instrument goes, and that is a constraint on the level, not on the loudness, and it is not in this model.

And it cannot say that roughness is what should be minimised. Roughness is a sensory quantity with a published model and known limits, and it is the only thing in this calculation with a number attached. Blend, clarity and the sense that one line is carrying are all real properties of a scoring and none of them appears — an instrument is not one timbre across its compass either, and the spectra here are taken at one pitch each; a solver that minimised roughness alone would double everything in octaves and sit still. What the constrained form buys is that the objective is only ever asked to break ties between arrangements that already meet a requirement, which is a much weaker use of it than a weighted sum would be.

Whose music, and when

The claim that the two problems interact is arithmetic and holds for any three sources. What is a claim about a repertoire is that anybody solves them in stages, and that claim is about how orchestration is taught rather than about how it is done.

The teaching order in the standard nineteenth-century treatises — Berlioz, and Rimsky-Korsakov after him — is instrument by instrument, then combination by combination, then balance as a chapter near the end. A student is given the discrete question first and the continuous one as a correction to it. That is exactly the two-stage method, and the figure above says what the correction cannot fix: the arrangement was already chosen.

Whether experienced orchestrators do anything of the kind is not something this collection can measure, and the honest reading of the finding is smaller and more useful. It is that an assignment ranking published without its levels is not a ranking, and this site published one four essays ago.

Where this ladder goes next

One rung. The anchor exists because two ladders named the same object, and the object turns out to be a constrained optimisation whose two halves do not separate.

The rung after it is the one the constraint’s own units name. Roughness is quadratic in level and loudness is compressive, so the two quantities in this problem do not rise together — and everything above holds the whole scoring at one loudness precisely so that they do not have to be compared. What happens to the balance when the dynamic itself moves is a question this machinery can now ask and this rung deliberately did not, and the answer decides whether a scoring that works at one dynamic is a scoring at all.

Part 1 of 14

One essay in the series on orchestration. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 18.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

LoudnessOptimisationOrchestrationRoughnessSensory dissonanceSpectral balanceTimbreVoicing