Concept

Orchestration — where it appears

Which instruments play what, which is the variable a cyclic form uses to change where a tonal form would change harmony. It is also the main control over loudness, since sources add in power rather than in amplitude.

Named by 40 essays across 7 fields — each of them below, with the objects they name alongside it.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward.

One sound hides another, and it hides upward

A tone can be made completely inaudible by a second tone that is nowhere near it in frequency, and the region it disappears into is lopsided. Masking spreads up the spectrum and barely down it, and the reach grows with the masker's level — so which line in an arrangement vanishes is a prediction, not a matter of taste.

perception · Masking
6 chords in a gothic cathedral. Each chord's reverberant decay in a room with a 8 second reverberation time, at 1 chord a second. Decay is linear in decibels, so each line is straight with a slope of -7.5 dB a second. When a chord arrives, 2 earlier ones are still above 20 dB down.

The room chooses the harmonic rhythm

A chord in a cathedral is still sounding, seven decibels down, when the next one arrives — and the one after that, and the one after that. Reverberation is linear in decibels, so the number of chords audible at once is one number divided by another, and it puts a hard ceiling on how fast a composer writing for that building can change harmony. The ceiling is computable, and the music written for those rooms sits under it.

timbre · Room acoustics
Every voicing of a major triad, least rough first. All 27 arrangements of the same three pitch classes within 3 octaves from 131 Hz, scored for roughness. The best is spaced 19 then 9 semitones — wide below, close above — and the worst is the chord in close position at the bottom of the range, 6.6 times rougher with exactly the same notes in it.

Where to put the third

Take three pitch classes, four octaves to put them in, and score all twenty-seven arrangements. The smoothest is root, fifth an octave up, third two octaves up — which is partials one, three and five of the harmonic series — and the roughest, at every register tried, is the chord in close root position at the bottom of the range. Every orchestration manual states that rule and none of them derives it.

harmony · Consonance
Six reverberation times for one room. Sabine's arithmetic evaluated in each octave band from the published absorption coefficients of the surfaces. a large stone church runs from 6.3 seconds at 125 Hz to 2.4 at 4 kHz — a bass ratio of 1.38, where concert halls are specified between 1.1 and 1.25.

A room does not decay evenly

Sabine's arithmetic gives one number and absorption is a strong function of frequency, so a room has six reverberation times rather than one. A stone church rings for 6.3 seconds at 125 hertz and 2.4 at 4 kilohertz, which means a chord left in it does not fade — it changes shape, losing its top before it loses its bottom, and arriving at the listener as a different sonority from the one played.

timbre · Room acoustics
Long-term average spectra: an orchestra, playing forte against a trained operatic soloist. Each source's mean spectrum over a long passage, in decibels below its own strongest region, on a logarithmic frequency axis. An orchestra, playing forte peaks at 250 Hz and is 30 dB down by 3,150 Hz; a trained operatic soloist peaks at 250 Hz and is 11 dB down by 3,150 Hz. The shapes are the same until about 1 kHz and separate above it: at 3153 Hz the difference is 19.0 decibels, which is the largest anywhere in the range. Nothing here is about level. Both curves are drawn against their own peaks, so what is being compared is shape.

One voice over ninety players

A soloist heard over a full orchestra is not louder than it and could not be. What the trained voice does instead is put a peak of energy at three kilohertz, which is where the orchestra's spectrum has already fallen away and where the ear's own threshold happens to be lowest. Nineteen decibels of advantage, in a place nobody is competing for.

timbre · The voice
The same intervals, played higher and higher. Sensory roughness for four fixed intervals as the pair is transposed up five octaves, computed from the same Plomp–Levelt model as the dissonance curve. Every one of them falls as it rises, and the small intervals fall furthest — so how consonant an interval is depends on where it is played, not only on what it is.

A chord is a register

The same three pitch classes are five times rougher at the bottom of a piano than in the middle, and the arrangement that minimises the roughness over a low bass turns out to be the bass's own fifth and sixth partials. The orchestration rule about low thirds is not a convention. It falls out of the width of a critical band, exactly.

intervals · The triad
How much of the writing the prohibitions forbid. I – vi – ii – V – I in C major, written in 3, 4, 5, 6 parts, with every ensemble covering the same total compass. A triad has three pitch classes, so n parts double n − 3 of them and a duet cannot state one at all. Of every ordered pair of complete voicings of V and I, the share with no parallel fifth or octave between any pair of voices: 91.2% at 3, 75.0% at 4, 44.3% at 5, 17.9% at 6.

Why the exercise is in four parts

Eight earlier essays move four voices, and nothing here ever chose four. Cut one choir's compass into three parts instead and the two most famous prohibitions in music cost exactly nothing — the cheapest realisation already obeys them. Cut it into six and they cost more than half a semitone per voice per chord change, because a parallel octave needs a doubled note and six voices sharing three notes can hardly avoid one. Four is the smallest number of parts at which the rules have a price at all, and it is the largest at which the price is small.

harmony · Voice-leading
3 tones, one power, and the interval between them. 3 tones of fixed total power, spread symmetrically about 440 hertz, drawn against the interval between neighbours. Piled on one pitch they are one sound of that power; separated by more than a critical band — 4.5 semitones here — they are 3 sounds whose loudnesses add, and the same power reaches 2.08 times the loudness at 5 semitones. The two lines are two models of the same rule and they disagree about how abrupt the change is, not about where it goes.

A chord is not as loud as its notes

The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.

perception · Loudness
How many intervals a duet is smooth at. Every ordered pairing of 5 radiators, with the number of wells in its dissonance curve over an octave from 262 hertz. Rows are the instrument underneath and columns the one above, so the grid is not symmetric about its diagonal and that asymmetry is the result. The count runs from 1 to 7 across the grid, and a cell and its mirror need not agree: a violin under a clarinet has 1 and a clarinet under a violin has 2. The fifth is a well in every one of the 25 pairings and no other interval is.

Which instrument is underneath

Every roughness curve so far compares two tones of the same timbre, which is a duet nobody plays. Give the two notes different instruments and the sum stops being symmetric: the same written interval, on the same two players, is up to five times rougher depending on which of them takes the lower note. From G3 upward the fifth is a well in all twenty-five pairings and no other interval is; below it, three pairings lose even that, and all three have a clarinet on top.

timbre · Spectrum
Crescendo, and what the impression does. A crescendo of 20 dB over 8 seconds, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.02 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.

Loud is relative, and it comes down slowly

The account of loudness had a model of a moment and the account of closure asked it for a model of a form. The published one exists and its content is a pair of numbers that are not the same: a listener's running impression of how loud the music is rises to meet a step in a fifth of a second and takes seven seconds to come back down. A twenty-decibel crescendo spread over eight seconds therefore buys almost no contrast at all, and the same twenty decibels taken as a step buys a factor of two.

form · Loudness
Six ways to put three players on three notes. The same chord — G3, B♭3, D4 — played by clarinet, oboe, voice in all 6 possible assignments, scored by the roughness each produces. Every bar is the same pitches and the same instruments; only who is on which note changes. The worst is 1.42 times the best, which is a factor a score can control and a chord symbol cannot express at all. Each row is labelled from the bottom note upward.

Which player on which note

An interval's roughness depends on which instrument is underneath, so the pair does not commute. Three players over three notes is the smallest thing that asymmetry has anywhere to go: six assignments, all of them the same chord, and across 450 of them the roughest averages half again the smoothest and reaches six times it. It is orchestration in the only form that can be computed here — not which chord, and not which voicing, but who is on which note.

timbre · Spectrum
Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 3 parts to 8 adds 4.3 decibels of power and 0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 6.0 to 3.9 as the parts crowd into the same register.

The dynamics are in the score already

Count the parts in each bar, realise them in their ranges, put every partial in its critical band, sum the loudnesses and run the result through the two smoothers built earlier. What comes out is a dynamic curve for a piece with no performance in it anywhere — and it says that doubling the number of parts inside a fixed register adds three decibels of power and about one of loudness, because the extra parts land in bands that were already occupied. Let the register widen with the parts and the same arithmetic gives eight phon, which is what a tutti actually is.

form · Loudness
Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness.

A final chord is not made loud by adding to it

An earlier essay on closure said the loudest cue an ending has needs a corpus rather than an arithmetic. The arithmetic was built one essay ago, so it does not. Run four ending textures through it and two things come out backwards: a final chord three parts thicker than the rest arrives *quieter* against the running impression than the passage it ends, and a texture that drops a part a bar does not get quieter at all until the bar where there is one part left.

form · Closure
Every assignment, at equal levels and at its own best balance. The 6 ways of putting 3 players on a chord, each drawn twice: hollow at equal levels, which is what an assignment ranking sees, and filled at the levels that balance the parts and then minimise roughness. Every scoring here is at the same total loudness, 23.3 sones, so two points are comparable. Solving the discrete problem first picks violin · clarinet · oboe; solving both at once picks violin · oboe · clarinet, and the two-stage answer costs 9.0 per cent more roughness. The orderings do not keep their places between the two columns, which is the whole of the argument: a ranking taken at equal levels is not a ranking.

Who plays what and how loud is one question

Two lines of argument, one about spectrum and one about loudness, each stopped at the same wall and each said so. One of them can choose who plays which note and has every player at the same level; the other can choose how loud each part is and has nobody assigned to anything. Put together they are a single problem with two kinds of variable, and solving it in stages picks a different answer from solving it at once — nine per cent rougher, at the same loudness, on an ordinary triad.

form · Orchestration
Roughness and loudness do not rise together. One four-note chord, played at levels from 35 to 95 decibels, with both quantities drawn as multiples of what they are at the quietest. Roughness is quadratic in pressure, so 60 decibels multiply it by 1.0e+6. Loudness is compressive — about ten phons to a doubling of sones — so the same range multiplies it by 96. The gap between the two lines is the quantity: roughness per sone rises by a factor of 1.0e+4 between a pianissimo and a fortissimo of the same chord.

The ranking survives the dynamic and the chord does not

Roughness is quadratic in pressure and loudness is compressive, so sixty decibels multiply a chord's roughness by a million and its loudness by ninety-six. Roughness per sone therefore rises ten thousandfold between a pianissimo and a fortissimo of the same four notes — and yet the ranking of which doubling is smoothest, over four hundred and eighty voicings, does not move by a single place.

harmony · Orchestration
Trumpet at three dynamics, as a spectrum rather than a level. The radiated partials of a trumpet at 45, 70, 95 decibels, each normalised to its own strongest partial so that only the SHAPE is compared. A linear source would give three identical pictures. This one does not: the spectral centroid moves from partial 2.19 to 6.41, a factor of 2.92, because the excitation is nonlinear and blowing harder steepens the pressure front rather than scaling it. The tilt used is 3 decibels per octave of partial number per ten decibels of level, referred to 70 dB — a stipulated, ordinal number, not a measurement of any instrument.

A dynamic mark changes what a note is

Every spectrum until now is a shape with a level in front of it, so that playing ten decibels louder raises every partial by ten. That is true of exactly one instrument in an orchestra. Everybody else steepens their own spectrum as they lean on it, and a trumpet's centre of gravity moves from the second partial to the sixth across a dynamic range while an organ flue pipe's does not move at all.

timbre · Orchestration
Which notes of a scored chord have to be played early. Four parts of one chord, each with its own instrument, its own pitch and its own dynamic, and the perceptual centre that comes out of all three. piano, sforzando on E1: an attack family of 8 milliseconds against a pitch floor of 97, so the pitch is what limits it, shortened by the dynamic to 65, heard 20.5 after it starts and needing to be played 12.0 early; flute, quiet on A5: an attack family of 60 milliseconds against a pitch floor of 5, so the instrument is, shortened by the dynamic to 69, heard 21.8 after it starts and needing to be played 13.2 early; violin, mezzo forte on E4: an attack family of 90 milliseconds against a pitch floor of 12, so the instrument is, shortened by the dynamic to 90, heard 28.5 after it starts and needing to be played 19.9 early; trumpet, forte on A3: an attack family of 30 milliseconds against a pitch floor of 18, so the instrument is, shortened by the dynamic to 27, heard 8.6 after it starts and needing to be played 0.0 early. The spread is 19.9 milliseconds, which is well above the two or three a listener resolves, so a conductor asking for these four to sound together is asking for four different physical onsets.

Which notes have to be played early

There are three separate contributions to one quantity — the instrument's attack family, the dynamic it is played at, and the note's own period — and every figure so far varies one and holds the others. Added together for a real scoring they do not add: a sforzando low piano note is pitch-limited to a hundred-millisecond attack and the sforzando shortens it back to sixty-five, so flattening the dynamics makes the ensemble's spread larger rather than smaller.

rhythm · Perceptual-centre
A cycle whose position is in the instrumentation. 3 isochronous layers over a cycle of 16 steps, at periods 16, 8, 4. Every layer on its own is perfectly symmetric and tells a listener nothing about where they are; the combination gives 4 distinct signatures over 16 steps, and hearing one of them leaves 2.81 bits unknown. The information is in which instruments sound rather than in where the onsets fall, which is a different answer from the one a single timeline gives — and it needs no asymmetry anywhere. The cost is 7 strokes a cycle, 0.44 to the step, spread over 3 players.

A cycle that says where it is

Euclidean timelines were asked how quickly they tell a listener where in the cycle they are, and answered it with rotational asymmetry: a symmetric pattern never locates at all. A colotomic cycle answers the same question with nothing asymmetric in it. Several isochronous layers at nested periods — a gong every sixteen, a kempul every eight, a kenong every four — put the position in which instruments sound, and the position is legible from a single stroke.

rhythm · Cyclic rhythm
A written dynamic is an instruction to the listener's impression. Every earlier scoring holds one chord still. A passage is a succession, and the running impression of loudness carries a chord into the one after it, so what a marking asks for and what playing the marking produces are different things. Here is a five-chord passage with a written shape. Playing each chord at its own written loudness gives the running impression 2.4, 3.0, 4.2, 5.6, 4.0 sones against the 2.4, 3.0, 4.2, 5.6, 2.0 that were asked for — right until the last chord, where it misses by 2.0. Solving for levels that make the impression arrive at the marking does not fix it: the last chord's target is I, two parts, and it is unreachable — the correction runs to silence and the impression still sits 1.1 sones above. A subito piano after a full chord is not a level a player can produce. It is a rate of change, and the smoother's two-second release is what refuses it.

A subito piano is a rate, not a level

All three earlier essays score one chord held still. An orchestration is a succession, and the running impression carries a chord into the one after it — so a written dynamic is an instruction to the listener's impression rather than to the instantaneous sound, and there are markings that cannot be produced at all. The correction runs to silence and the impression still sits above the target.

form · Orchestration
Intonation is a unison problem and nothing else. The roughness between two instruments on one note, against how far apart they are in cents, drawn for a unison and for the intervals beside it. A perfect unison is 0.0007 — the partials coincide and there is nothing to beat. Five cents apart it is 0.0465, 65 times as rough, and ten cents apart it is rougher than a major third played exactly. The mechanism is that partial n of a note mistuned by c cents is mistuned by c cents as well, which is n times as many hertz — so the top of the spectrum enters the critical band long before the fundamental does. The other curves are flat, because a third's roughness is set by which partials nearly coincide and a few cents does not change which.

Two players on one note

Six essays have put one instrument on each note of a chord, and the commonest thing an orchestrator actually does is put two on the same note. Two independent sources add in power, so the composite is neither of them — except that it nearly always is one of them, because the level at which ownership changes hands is rarely at zero. And a unison ten cents out is rougher than a major third dead in tune.

timbre · Spectrum
One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics.

The dissonance arrives and the dynamic does not

A scoring decides two things at once and both of them have to be integrated by a listener before they exist. The loudness smoother's release is two seconds and the roughness window is thirty-seven milliseconds, and that ratio of fifty decides which of the two survives at the pace music is actually played. Nothing anybody performs is fast enough to blur a dissonance, and a great deal of it is fast enough to average a dynamic.

form · Orchestration
A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance.

A page has two decibels

The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.

perception · Loudness
The passage that separates them, and a listener cannot hear it. A scoring changes at the halfway bar, and the two maps of required leads differ by 18.1 milliseconds at their widest. An ensemble that has internalised the map applies the new one on the first note of it and its spread never leaves zero. An ensemble that is listening to each other has to re-converge: its spread jumps to 11.8 milliseconds and takes 3 beats to get back under 5. The dashed line is twenty milliseconds, which is what a listener notices — and the disagreement never reaches it. So the two accounts are separable on a recording and very nearly not separable by ear, which is why nobody has noticed the distinction and why the measurement is worth making.

The passage that separates two players

An ensemble that has learnt where the asynchronies are applies them; one that is listening discovers them. In steady state the two are identical, which is why nobody has separated them. Change the scoring mid-phrase and they are not: one ensemble is wrong by twelve milliseconds for three beats and the other is not wrong at all — and twelve milliseconds is under what a listener notices and far above what a microphone resolves.

rhythm · Perceptual-centre
Where a cycle of 16 at 16, 8, 4 outruns the listener's memory. The residual uncertainty a listener is left with once the evidence has stopped accumulating, against how long one turn of a 16-step cycle takes. The listener's memory of a step halves after 3.5 seconds throughout; what changes is how many steps that is. At a cycle of 1.6 seconds it is 35 steps and every design reaches certainty, which is the regime a clave is played in. At 60 seconds it is 0.93 steps and none of them does: layers at 16, 8, 4 settles at 1.89 bits, son clave settles at 2.27 bits, the bossa-nova pattern settles at 2.38 bits, the best single line of 7 settles at 1.85 bits. That is the range a gong cycle occupies, and it is the design that wins there.

The cycle that outruns the memory

A timeline and a colotomy were compared at equal strokes and the comparison had no clock in it. A memory span is a number of seconds and a cycle is a number of steps, so the two only meet through a tempo — and at a clave's two seconds a listener's memory covers twenty-eight steps and forgets nothing, while at a gong cycle's forty it covers 1.4 and forgets almost everything. The single line is the better locator up to twenty-three seconds a cycle and the layered code is better after it, which is very close to where each is actually used.

rhythm · Cyclic rhythm
An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none.

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

perception · Loudness
The top voice arrives whole and the bottom one arrives as a sine. 4 parts sounding together, each of 8 partials, with every partial tested against the summed masked threshold of every component in the texture. A filled mark is a partial the listener receives and an open one is a partial the part would have had alone and does not have here. The bass at C3 keeps 1 of 8, the tenor at C4 keeps 2 of 8, the alto at E4 keeps 5 of 8, the soprano at G5 keeps 8 of 8, every one of them at 70 decibels. Every part is at the same level and the difference is entirely where each one sits: masking spreads upward, so the part at the top of the texture has nothing above it to be masked by and the part at the bottom has everything.

The listener is given the top voice, and the bass as a sine

Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.

perception · Masking
Who owns clarinet, oboe, voice at every balance. The composite of three players at 392 hertz belongs to whichever of them it is nearest in log-spectral distance, and here that is drawn over the whole plane of balances a conductor could set — the second and third players from 24 decibels below the first to 24 above. voice owns 79 per cent of the square. The three regions meet where all three distances are equal, which is the only balance at which the composite belongs to nobody: it is at -0.3 and -19.1 decibels, inside the square and therefore a balance an ensemble could actually be asked for. A trio has a colour of its own at one point, not over a region.

A section has a loudest member, not a colour

Two players on one note have a balance at which the composite belongs to neither, and that is what blending means. Three should have three such balances and no reason for them to agree — a trio with a rock-paper-scissors ownership would have no strongest member at all. Twenty trios, sixty pairwise comparisons, and not one disagreement: the possibility is real, arbitrary spectra do it once in twenty, and instruments never do.

timbre · Spectrum
What the chord before takes out of the chord after. Five chords at 1.2 seconds each, with the roughness each one has on its own — its simultaneous masking and the threshold of hearing already applied — and the roughness it actually has once the chord in front of it has raised the threshold. Four of the five are untouched. The fifth, i, two parts, follows the only step in this passage that falls more than fifteen decibels, and it arrives into a hole: it is entirely below threshold for its first 13 milliseconds and takes 240 to get all of itself back. Masking can only remove partials, so it can only lower a roughness — and the chords it can reach are the ones a written dynamic has just made quiet, which are already the smooth ones. Across the passage the dissonance contrast goes from 3021 to 3113: the mask widens it by 3.0 per cent rather than eating it.

A soft chord has to fade in

Forward masking sits between the two integration times already in play — two hundred milliseconds against a thirty-seven millisecond roughness window and a two-second loudness release — and it was owed as the term that might eat the dissonance contrast. It does not. It widens it, by three per cent at a chorale's pace and fifty-nine at four chords a second, because it can only ever remove partials and it can only reach the chord a dynamic has already made quiet. What it does instead is stranger: one chord in the passage is entirely inaudible for its first twelve milliseconds and takes a quarter of a second to arrive whole.

form · Orchestration
Four players on three notes, every arrangement. The 36 ways of putting 4 players on a 3-note chord so that every note is covered, ranked by roughness, all at one total loudness of 26.9 sones. Each row is shaded by which note carries the pair. The best is flue | clarinet+violin | oboe and the worst is clarinet | oboe+flue | violin, a factor of 2.21. Every earlier essay puts exactly one player on each note, which is a permutation; a doubling makes the arrangement a surjection instead, and the doubled note sounds neither of its two players but the composite they make. Which note gets the pair explains 7 per cent of the spread here and which players sit on the lowest note explains 89: the fourth player is a much smaller decision than the three that were already there.

The fourth player is a spectrum, not a decision

Six earlier essays put exactly one instrument on each note, which makes an arrangement a permutation — and the commonest operation in orchestration is a doubling, which does not. Four players on three notes give thirty-six arrangements instead of six, and the extra choice turns out to be the smallest thing on the page: which note carries the pair explains three per cent of the spread and which players sit on the bass explains eighty-nine. A doubled note can be priced as one player, and which one is not the one a spectral account would have named.

instruments · Orchestration
14 players summed, against the one at their average onset. Each thin line is one player's rising envelope, started at its own moment, with a spread of 30 milliseconds about the beat and a 90-millisecond attack. The heavy line is the section: nominally identical sources add incoherently, so their powers add and the sum is the root of the mean of their squares, drawn here as a fraction of the section's own peak. The dashed line is the single player who started at the section's average onset. The section reaches the 6 dB below peak criterion at 18.1 milliseconds and that player at 27.6, a difference of 9.5. The section is early because the players who started first are already sounding while the average one is still building, and nothing a late player does can make the sum quieter.

Twelve violins are more punctual than one

Every essay until now treats a part as one player, and an orchestral part is a dozen. Sectioning does two things at once and only one of them was expected: it pulls the part's heard moment forward, by four milliseconds against a map spanning twenty-six, and it makes the part's arrival more accurate by very nearly the root of the number of players. So the map of required leads applies to an orchestra better than it applies to a quartet, and the case where it fails is three trumpets rather than fourteen violins.

rhythm · Perceptual-centre
One doubling, held down a phrase. Where a single held arrangement of 4 players on 3 notes stands among the 36 at each chord of a 5-chord passage, best at the top, with what each chord would rather have named along the bottom. The held answer is flue pipe · trumpet · clarinet+violin, and it is the chord's own first choice at 4 of 5 of them. Holding it costs 16.2 per cent of the passage's roughness against re-scoring every chord — which is 2.4 per cent of the range the choice actually spans, since the arrangements at one chord differ by a factor of 7.8 on average. The cost is not spread over the passage: 1 chord carries nearly all of it.

An orchestrator doubles a line, not a chord

Three earlier essays made the objective a functional over a passage and a later one went back to holding one chord still. Put the doubling back into time and the retreat turns out to have been cheap: one arrangement held down a five-chord phrase is that phrase's own best answer at four of its five chords and costs 2.4 per cent of the range the choice spans — while the forward mask named earlier as the third temporal constant reaches for twenty milliseconds rather than two hundred, and cannot change the answer at any pace at all.

instruments · Orchestration
A bar of silence is worth 13.1 decibels written last and 6.0 written first. Four closing gestures, drawn against how many seconds of silence each contains. All four hold the same 12-part texture, write the same 6.0-decibel diminuendo over 4 seconds, and differ only in what order the diminuendo and the silence are written in. The quantity is how far the listener's running impression has fallen when the final chord arrives, in decibels of equivalent diminuendo. Written with the silence last, a general pause of a bar is worth 13.06 decibels and one of three and a half seconds is worth 28.3. Written with the silence first, both are worth 6.00 — exactly the diminuendo's own depth, because the music resuming after the silence puts the reference back at its own level. The silence written first is worth less than the silence written with no diminuendo at all, which reads 8.12 at a bar: a diminuendo placed after a general pause takes 2.12 decibels away and adds nothing.

A general pause is spent by the note after it

Whether a composer should write the pause before the diminuendo or after it looks like a question about how big the ensemble is. It is not. Forty decibels of ensemble are worth one decibel of silence, and the order is worth seven — because a running impression rises twenty times faster than it falls, so half a general pause is spent by ninety-seven milliseconds of sound.

form · Closure
An entrance stops being a loudness event and never stops being a colour one. The same oboe entering on the same note at the same level, against how many players were already sounding. Its contribution to the loudness falls from 15.5 phons to 0.32 — a factor of 48 — and crosses the one-phon difference limen at 5 players already playing. Its contribution to the roughness rises by a factor of 12.3 over the same range, because roughness is a sum over pairs and the entrant makes one new pair with everybody. Both curves are drawn as a share of their own largest value, since a phon and a squared pascal have no exchange rate. The claim is the two directions, not the crossing point of two units.

An entrance is a change of colour

Eight essays on orchestration move the assignment and hold the ensemble still, and a score does the opposite: it brings players in and takes them out. Loudness is a sum over parts and roughness is a sum over pairs, so the player who joins adds one term to the first and one to the second for everybody already there. What the entrance is worth in phons falls by a factor of forty-eight across the range an ensemble spans and crosses the difference limen at five players; what it is worth in roughness rises by twelve, and by a further factor of ten for every ten decibels the passage is played at.

form · Orchestration
Three clocks receive one entrance, and they do not agree about when. An oboe joining 5 players already sounding, at time zero, with each of the listener's three readings drawn as its own share of the change it eventually makes. The roughness window is 49 milliseconds wide and has half the change at 25; the short-term loudness smoother has half at 15; the long-term one, whose release is the two seconds an earlier essay is about, has half at 90. The two-second release is on the wrong side of the smoother to hide an entrance. Its attack is 99 milliseconds, so an entrance is received promptly and it is a departure that is not.

The release is on the wrong side

Whether the loudness model's two-second release makes an entrance inaudible has the answer no, for a reason the question did not anticipate. The smoother is asymmetric — ninety-nine milliseconds going up and two seconds coming down — so a rise is tracked twenty times faster than a fall, and an entrance is received promptly by every one of a listener's three readings. The colour of it arrives first, at twenty-five milliseconds against ninety, and the reading that moves with the ensemble is the one nobody would have picked.

form · Orchestration
The same player, arriving and leaving, read as a share of the change. One oboe joining 5 players and the same oboe leaving them again, with both loudness readings drawn as the share of their own change that has arrived. The entrance is half received in 90 milliseconds and the exit in 1.43 seconds, a factor of 15.9. The roughness readings, drawn faintly, are 25 and 25 milliseconds and lie on top of each other. A score that writes a diminuendo under a departing part is not softening the exit. It is doing the smoother's release for it, on a clock the smoother would otherwise take two seconds over.

A part that leaves is not a part that arrives

The same player, the same note, the same level, and the only difference is which way round it happens. A listener's loudness reading takes 1.43 seconds to receive half of a departure and 90 milliseconds to receive half of an arrival — a factor of sixteen with nothing asymmetric in the sound at all, since both readings integrate the same two states in the same order. The colour reading receives the two identically, because a window has no direction, so a departure is a change whose grain arrives at once and whose level takes most of two seconds.

form · Orchestration
Which chord of a passage has room for the part that is entering. An oboe entering on one note, tried at each chord of a five-chord passage, scored by how far its own partials sit above the threshold the ensemble already sounding puts over them. The best moment gives it 9.0 decibels of margin and the worst 0.3, a spread of 8.7 — and the best moment is not the quietest chord, which is vi, close below. Room for an entrance is spectral rather than dynamic. A chord with a hole in its written spacing need not have one in its spectrum, because the partials of its bass fill the middle whatever the notes above it do.

The chord that has room for an entrance

Three essays have made the ensemble something a score can change and none of them has asked when. The ensemble already sounding puts a masked threshold over whatever register an entering part takes, and that threshold is set by the voicing rather than by the dynamic — so the five chords of one passage differ by 8.7 decibels in how much of an entering oboe survives them, and the quietest chord of the five is the worst place in the passage to bring somebody in. Swept over the entrant's own pitch, the choice of moment is worth as much as the choice of register.

form · Orchestration
A doubled pizzicato gives its note away while it is still the louder. The power of a violin plucked, against a flue pipe holding the same note at 392 hertz, through the first 600 milliseconds of the pluck, with the pluck starting 12 decibels up and its fundamental decaying over 1 second. With each partial losing level in proportion to its number, the composite stops resembling the pluck at 70 ms, when the pluck is still 5.2 decibels the louder. With every partial fading together it would keep the note until 543 ms. The dashed line is the balance at which the steady-state doubling changes owner, minus 20.6 decibels: the release crosses the owner long before its balance gets there, because what hands the note over is the pluck's upper partials going, not its level.

A doubled pizzicato gives its note away early

The attack turns the balance between two players on one note by a few decibels and stops. A pluck does not stop — every partial of it decays, so a pizzicato doubled by a held instrument walks the balance for the whole note, and the expectation was a handover as slow as the decay. It is fast. A one-second pizzicato over a flute loses its note in 70 milliseconds, while it is still five decibels the louder, because what hands the note over is its upper partials going first. A uniform fade would have kept it eight times as long.

timbre · Spectrum
The schedule that hears every entrance best holds the high parts back. Six parts waiting to enter a five-chord passage over four sounding players, each entering once and staying: the schedule under which the least audible entrance is as audible as it can be made. brass on E3 enters at I, open with -1.4 decibels of mean margin over the mask; oboe on E4 enters at vi, close below with -2.6 decibels of mean margin over the mask; clarinet on G4 enters at I, open with -1.4 decibels of mean margin over the mask; voice on C5 enters at IV, close above with 2.1 decibels of mean margin over the mask; violin on G5 enters at V, bracketing with 6.9 decibels of mean margin over the mask; flue pipe on C6 enters at I, hollow with 4.0 decibels of mean margin over the mask. The least audible entrance is at -2.6 decibels and the margins sum to 7.6; of all 15625 schedules 0 have a better least audible entrance and 576 a larger sum.

Room is used up by whoever enters first

The chord with the most room for a part entering alone is a fact about that chord. It stops being a fact the moment two parts want it, because each part that comes in raises the mask over everybody after it. Given six parts waiting to enter a five-chord passage, choosing each part's moment the way one part's moment is chosen puts three of them into the same chord and lands in the bottom fifth of all 15,625 schedules. Placing them one at a time does no better. The schedule under which the least audible entrance is heard best is unique, and it brings the low and middle parts in while the texture is thin and holds the three highest back for the last three chords — because a high part keeps its room over a full texture and a middle part does not.

form · Orchestration
How long a doubled pizzicato keeps its note, seat by seat, in two rooms. How long a doubled violin pizzicato on 392 hertz keeps its note against the metres from the players, the pluck starting 12 dB up and decaying over 1 s with a loss exponent of 1. a concert hall, a flue pipe: 1 → 86 ms, 1.5 → 123 ms, 2 → 226 ms, 3 → 359 ms, 5 → 445 ms, 7 → 481 ms, 10 → 506 ms, 15 → 522 ms, 20 → 528 ms, 30 → 532 ms; a concert hall, an oboe: 1 → 52 ms, 1.5 → 55 ms, 2 → 59 ms, 3 → 77 ms, 5 → 149 ms, 7 → 195 ms, 10 → 224 ms, 15 → 242 ms, 20 → 248 ms, 30 → 254 ms; a concert hall, a clarinet: 1 → 44 ms, 1.5 → 45 ms, 2 → 45 ms, 3 → 47 ms, 5 → 53 ms, 7 → 65 ms, 10 → 86 ms, 15 → 105 ms, 20 → 112 ms, 30 → 118 ms; a large stone church, a flue pipe: 1 → 440 ms, 1.5 → 578 ms, 2 → 651 ms, 3 → 728 ms, 5 → 784 ms, 7 → 803 ms, 10 → 814 ms, 15 → 820 ms, 20 → 822 ms, 30 → 824 ms; a large stone church, an oboe: 1 → 56 ms, 1.5 → 65 ms, 2 → 89 ms, 3 → 207 ms, 5 → 281 ms, 7 → 303 ms, 10 → 315 ms, 15 → 322 ms, 20 → 325 ms, 30 → 326 ms; a large stone church, a clarinet: 1 → 44 ms, 1.5 → 43 ms, 2 → 43 ms, 3 → 44 ms, 5 → 53 ms, 7 → 67 ms, 10 → 80 ms, 15 → 89 ms, 20 → 92 ms, 30 → 94 ms. The mid-band critical distance is 5.3 m in a concert hall and 2.3 m in a large stone church. In none of the 60 cases does the note return to the pluck once it has left.

A room keeps a pizzicato from giving its note away

Doubled by a flute, a one-second pizzicato loses its note in 70 milliseconds dry, because its upper partials go first. The question left open was whether a room, whose reverberation keeps those partials alive, gives the note back afterwards. It does not give it back. It stops the note going: ten metres into a concert hall the pluck keeps it for 506 milliseconds, in a stone church for 814, and the room's own uneven decay takes back between a quarter and two fifths of that. In a room the loss law that decided everything dry matters a tenth as much, because the room's decay has become the clock.

timbre · Spectrum
With 4 of six required at the end, the best schedule hands parts over. Six parts entering, leaving and re-entering a five-chord passage over four sounding players, in the walk through all 64 sets of sounding parts that makes the least audible entrance as audible as possible, with no memory of the chord before and at least 4 of the six sounding at the last chord. I, open: clarinet on G4 enters at 0.6 dB; vi, close below: oboe on E4 enters at 0.3 dB, and clarinet leaves; IV, close above: brass on E3 enters at -0.7 dB, voice on C5 enters at 2.2 dB, and oboe leaves; V, bracketing: violin on G5 enters at 7.2 dB; I, hollow: flue pipe on C6 enters at 4.0 dB. The least audible entrance is -0.72 dB against -2.65 for the best schedule in which nobody leaves; the walk has 2 exits and 6 entrances.

An exit is worth nothing until the tutti is given up

Six parts entering a five-chord passage have a best schedule when each enters once and stays, and letting parts leave and come back was supposed to improve it. Searched over every set of sounding parts at every chord, it improves it by exactly nothing, with or without the chord before still masking — as long as all six must be playing at the end. Let one part be missing from the final chord and the weakest entrance gains 1.4 decibels; let two be missing and it gains 1.9, by a relay in which the parts with least room come in, are heard for one chord, and give way.

form · Orchestration

Named alongside it

The objects these essays reach for when they reach for this one.

Critical bandwidthRoughnessLoudnessMaskingDynamicsSpectrumVoicingRegisterTextureTimbreEnumerationIntegration window

All concepts