Concept

Dynamics — where it appears

How loud the music is, as a written instruction and as a physical fact about the sound. The two are not the same quantity: a marking is an instruction about effort rather than level, and effort changes the spectrum and the attack as well as the amplitude.

Named by 30 essays across 7 fields — each of them below, with the objects they name alongside it.

Two clocks on one contact, and which of them runs out first. Two times that could end a hammer's contact with a string, across the compass. The flat line is the felt's own recoil, which this site models as a fixed 1.6 milliseconds at a stated force and which does not know what note is being played. The falling line is the string's: a mass m driving a string of impedance Z loses its momentum in m/2Z, which is 11.9 milliseconds at the bottom of the keyboard and 0.76 at the top. They cross at G3. Below the crossing the felt lets go first and above it the string gets there first, and over six octaves the two are within a factor of two of each other — so the contact time is a coupled quantity and not a property of the felt.

The hammer that is heavier than its string

Three earlier essays varied where the hammer lands, how long it stays and how wide it is, and the third named the fourth variable: the exciter's own mass. It inverts across one instrument — the hammer is lighter than the wire it hits at the bottom of a piano and thirteen times heavier at the top — and it turns the contact time into a quantity the string decides rather than the felt.

instruments · Excitation point
How much earlier an accent is heard, by mechanism. An accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller.

Playing louder is playing earlier

An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.

perception · Perceptual-centre
Crescendo, and what the impression does. A crescendo of 20 dB over 8 seconds, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.02 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.

Loud is relative, and it comes down slowly

The account of loudness had a model of a moment and the account of closure asked it for a model of a form. The published one exists and its content is a pair of numbers that are not the same: a listener's running impression of how loud the music is rises to meet a step in a fifth of a second and takes seven seconds to come back down. A twenty-decibel crescendo spread over eight seconds therefore buys almost no contrast at all, and the same twenty decibels taken as a step buys a factor of two.

form · Loudness
Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 3 parts to 8 adds 4.3 decibels of power and 0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 6.0 to 3.9 as the parts crowd into the same register.

The dynamics are in the score already

Count the parts in each bar, realise them in their ranges, put every partial in its critical band, sum the loudnesses and run the result through the two smoothers built earlier. What comes out is a dynamic curve for a piece with no performance in it anywhere — and it says that doubling the number of parts inside a fixed register adds three decibels of power and about one of loudness, because the extra parts land in bands that were already occupied. Let the register widen with the parts and the same arithmetic gives eight phon, which is what a tutti actually is.

form · Loudness
Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness.

A final chord is not made loud by adding to it

An earlier essay on closure said the loudest cue an ending has needs a corpus rather than an arithmetic. The arithmetic was built one essay ago, so it does not. Run four ending textures through it and two things come out backwards: a final chord three parts thicker than the rest arrives *quieter* against the running impression than the passage it ends, and a texture that drops a part a bar does not get quieter at all until the bar where there is one part left.

form · Closure
Trumpet at three dynamics, as a spectrum rather than a level. The radiated partials of a trumpet at 45, 70, 95 decibels, each normalised to its own strongest partial so that only the SHAPE is compared. A linear source would give three identical pictures. This one does not: the spectral centroid moves from partial 2.19 to 6.41, a factor of 2.92, because the excitation is nonlinear and blowing harder steepens the pressure front rather than scaling it. The tilt used is 3 decibels per octave of partial number per ten decibels of level, referred to 70 dB — a stipulated, ordinal number, not a measurement of any instrument.

A dynamic mark changes what a note is

Every spectrum until now is a shape with a level in front of it, so that playing ten decibels louder raises every partial by ten. That is true of exactly one instrument in an orchestra. Everybody else steepens their own spectrum as they lean on it, and a trumpet's centre of gravity moves from the second partial to the sixth across a dynamic range while an organ flue pipe's does not move at all.

timbre · Orchestration
The tempo turns, and almost nothing moves. The earlier arrival reading — what is sounding at the final chord over what the listener has been hearing — swept over bar lengths from 0.5 to 5 seconds, which is 480 down to 48 beats a minute, at 3 closing lengths. Every curve is nearly flat. Across a tenfold change of tempo one gesture's reading moves by a factor of 1.201 and the other's by 1.098, while the gap between the two gestures — which is what that essay was measuring — is 1.228. The expectation was that the tempo would decide the answer, on the grounds that a two-second bar against a two-second release is a comparable pair. The premise is wrong in a way the sweep makes obvious: the thing being compared with the release is not a bar, it is the WHOLE ENDING, which is 2 to 8 bars long and is therefore far longer than the release at every tempo anybody plays. The running impression has caught up with the closing texture before the final chord arrives, at 0.5 seconds a bar and at 5, and what is left is the last bar's own jump.

The parameter that did not decide the answer

An earlier essay on closure ended by naming the tempo as the thing every number in it was resting on, and said it was the kind of parameter that had caused trouble before by turning out to decide the answer. Turned across a tenfold range at a closing gesture of fixed length it moves the reading by four per cent, against a twenty-three per cent gap between the gestures it is distinguishing. The parameter beside it in the same figure — how many bars the gesture occupies — moves it by twenty, and nobody had named that one at all.

form · Closure
Which notes of a scored chord have to be played early. Four parts of one chord, each with its own instrument, its own pitch and its own dynamic, and the perceptual centre that comes out of all three. piano, sforzando on E1: an attack family of 8 milliseconds against a pitch floor of 97, so the pitch is what limits it, shortened by the dynamic to 65, heard 20.5 after it starts and needing to be played 12.0 early; flute, quiet on A5: an attack family of 60 milliseconds against a pitch floor of 5, so the instrument is, shortened by the dynamic to 69, heard 21.8 after it starts and needing to be played 13.2 early; violin, mezzo forte on E4: an attack family of 90 milliseconds against a pitch floor of 12, so the instrument is, shortened by the dynamic to 90, heard 28.5 after it starts and needing to be played 19.9 early; trumpet, forte on A3: an attack family of 30 milliseconds against a pitch floor of 18, so the instrument is, shortened by the dynamic to 27, heard 8.6 after it starts and needing to be played 0.0 early. The spread is 19.9 milliseconds, which is well above the two or three a listener resolves, so a conductor asking for these four to sound together is asking for four different physical onsets.

Which notes have to be played early

There are three separate contributions to one quantity — the instrument's attack family, the dynamic it is played at, and the note's own period — and every figure so far varies one and holds the others. Added together for a real scoring they do not add: a sforzando low piano note is pitch-limited to a hundred-millisecond attack and the sforzando shortens it back to sixty-five, so flattening the dynamics makes the ensemble's spread larger rather than smaller.

rhythm · Perceptual-centre
A loud chord is a smaller chord. The share of a voicing's partials that stand above what the rest of it masks, and the share of its computed roughness that is between partials a listener actually has, from 30 decibels to 100. Both fall: 83 per cent of the partials survive at 30 decibels and 38 at 100, and the roughness share goes from 88 per cent to 67. The direction is the upward spread of masking, which grows faster than linearly with level: a loud partial masks a band above itself much wider than a quiet one does, so the chord's own top disappears into its own bottom. Two earlier essays are drawn at one level, and this is what they were holding.

A loud chord is a smaller chord

Two earlier essays hold the level fixed, and the level decides how much of a chord a listener is given. At thirty decibels twenty of a triad's twenty-four partials stand above what the rest of it masks; at a hundred, nine do. Every roughness figure until now counts partials that are in the score, and a partial the chord masks is not a partial the listener has.

harmony · Masking
A ritardando does not spend the diminuendo. The arrival reading under a deceleration into the ending, from no ritardando at all to a final tempo 30 per cent of the starting one — which stretches the closing bars from 12.0 seconds to 20.2. The expectation was that it would matter: a ritardando lengthens exactly the bars the gesture is happening in, so a diminuendo that would have been absorbed at a steady tempo gets more of the smoother's own time to be absorbed in. It moves the reading by 0.00 per cent. Every line here is flat to within the thickness of the line, which is the second time a tempo parameter has been swept here and found to do nothing.

The reading was a step response

Sweeping the tempo found it did not decide the answer. This one sweeps the deceleration across a factor of three and finds a null to five figures, and then sweeps the length of the closing gesture across a factor of forty-eight and finds it moves the reading by eight per cent — but not as a function of seconds. Sorted by seconds the twelve runs scatter; sorted by how many bars the instruction covers they fall into three tight groups. One sentence explains the null and the not-null together.

form · Closure
A written dynamic is an instruction to the listener's impression. Every earlier scoring holds one chord still. A passage is a succession, and the running impression of loudness carries a chord into the one after it, so what a marking asks for and what playing the marking produces are different things. Here is a five-chord passage with a written shape. Playing each chord at its own written loudness gives the running impression 2.4, 3.0, 4.2, 5.6, 4.0 sones against the 2.4, 3.0, 4.2, 5.6, 2.0 that were asked for — right until the last chord, where it misses by 2.0. Solving for levels that make the impression arrive at the marking does not fix it: the last chord's target is I, two parts, and it is unreachable — the correction runs to silence and the impression still sits 1.1 sones above. A subito piano after a full chord is not a level a player can produce. It is a rate of change, and the smoother's two-second release is what refuses it.

A subito piano is a rate, not a level

All three earlier essays score one chord held still. An orchestration is a succession, and the running impression carries a chord into the one after it — so a written dynamic is an instruction to the listener's impression rather than to the instantaneous sound, and there are markings that cannot be produced at all. The correction runs to silence and the impression still sits above the target.

form · Orchestration
One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics.

The dissonance arrives and the dynamic does not

A scoring decides two things at once and both of them have to be integrated by a listener before they exist. The loudness smoother's release is two seconds and the roughness window is thirty-seven milliseconds, and that ratio of fifty decides which of the two survives at the pace music is actually played. Nothing anybody performs is fast enough to blur a dissonance, and a great deal of it is fast enough to average a dynamic.

form · Orchestration
Nothing at all until fifteen decibels, and then it depends on the tempo. The fraction of a line's partials that stay above threshold, over how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement — forward masking leaves a threshold at most ten decibels below the masker, and a note's own partials mask each other from the same components at full level. The boundary is between twelve and eighteen decibels, and beyond it the loss grows with the tempo: at 280 to the crotchet and 36 decibels of contrast, 26 per cent of the line's partials are gone. Fifteen decibels is about the gap between a forte and a piano.

An equal note cannot be masked

Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.

perception · Masking
A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance.

A page has two decibels

The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.

perception · Loudness
A rest is a diminuendo, and a long one. How far a listener's running impression of loudness falls during a silence, converted into the diminuendo that would have taken it the same distance. Half a second of nothing is worth 3.2 decibels, a second and a bit is worth 8.1, and two and a half seconds is worth 17. The marked line is three and a half seconds, which is where a gap starts to be heard as an ending rather than as a pause: at that length the reference has fallen by 24 decibels, which is more than a fortissimo to a pianissimo. A tempo swept over a factor of ten, a deceleration over a factor of three and a gesture length over a factor of forty-eight all returned the same reading to four significant figures. This one moves it by twenty-four decibels.

A rest is a diminuendo

Three parameters swept over factors of ten, three and forty-eight returned the same reading to four significant figures. The one manipulation left unswept moves it by twenty-four decibels: a silence. A listener's running impression decays at the loudness smoother's two-second release, so a general pause is a diminuendo nobody wrote, and at the length that makes a gap an ending it is worth more than any marking a composer has.

form · Closure
How much of each instrument is narrow enough to steepen a wave. For each bore, the quantity that decides how nonlinear it is: the narrowest radius divided by the radius at each station, integrated along the tube. The shading is that integrand, so a bar that stays dark is a tube still doing damage to the wave and a bar that fades is a flare that has thinned it out. Divided by the instrument's own length the integral is a pure number: a plain cylinder 1.00, a tenor trombone 0.88, a trumpet 0.86, an F horn 0.59, a cone of a trumpet's length 0.24. A plain cylinder is 1 by construction, a cone of the same length and mouth is 0.24, and the ordering across the brass family is the one players give when asked which of them can be made to blare.

The partials the tube makes itself

Eleven earlier essays compute a passive linear resonator, and none of them ever says so. At a real fortissimo the air in a brass instrument is not linear: a compression outruns a rarefaction, the wave leans forward as it travels, and the fourth partial of a loud trumpet note is seventy decibels louder than a scaled-up quiet one — generated in the tube rather than at the lips. How much of it happens is an integral over the bore, and it is why a flugelhorn cannot be blown into being a trumpet.

instruments · Air column
An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none.

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

perception · Loudness
Articulation is worth 2.6 phons, and nobody counts it. A passage of notes at 80 decibels, 2 to the beat, at 7 tempi and 3 articulations, scored against the same level held continuously. The variable is the fraction of each inter-onset interval that is sounding — 0.95 is a legato, 0.4 a staccato — and the vertical axis is what that costs the passage's running loudness in phons. Nothing here is anybody playing harder or softer. At 40 to the beat the span from legato to staccato is 1.47 phons; at 200 it is 2.61, because a staccato note there lasts 60 milliseconds and no longer reaches its own loudness either.

A staccato is a dynamic mark

Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.

perception · Loudness
What the chord before takes out of the chord after. Five chords at 1.2 seconds each, with the roughness each one has on its own — its simultaneous masking and the threshold of hearing already applied — and the roughness it actually has once the chord in front of it has raised the threshold. Four of the five are untouched. The fifth, i, two parts, follows the only step in this passage that falls more than fifteen decibels, and it arrives into a hole: it is entirely below threshold for its first 13 milliseconds and takes 240 to get all of itself back. Masking can only remove partials, so it can only lower a roughness — and the chords it can reach are the ones a written dynamic has just made quiet, which are already the smooth ones. Across the passage the dissonance contrast goes from 3021 to 3113: the mask widens it by 3.0 per cent rather than eating it.

A soft chord has to fade in

Forward masking sits between the two integration times already in play — two hundred milliseconds against a thirty-seven millisecond roughness window and a two-second loudness release — and it was owed as the term that might eat the dissonance contrast. It does not. It widens it, by three per cent at a chorale's pace and fifty-nine at four chords a second, because it can only ever remove partials and it can only reach the chord a dynamic has already made quiet. What it does instead is stranger: one chord in the passage is entirely inaudible for its first twelve milliseconds and takes a quarter of a second to arrive whole.

form · Orchestration
Four players on three notes, every arrangement. The 36 ways of putting 4 players on a 3-note chord so that every note is covered, ranked by roughness, all at one total loudness of 26.9 sones. Each row is shaded by which note carries the pair. The best is flue | clarinet+violin | oboe and the worst is clarinet | oboe+flue | violin, a factor of 2.21. Every earlier essay puts exactly one player on each note, which is a permutation; a doubling makes the arrangement a surjection instead, and the doubled note sounds neither of its two players but the composite they make. Which note gets the pair explains 7 per cent of the spread here and which players sit on the lowest note explains 89: the fourth player is a much smaller decision than the three that were already there.

The fourth player is a spectrum, not a decision

Six earlier essays put exactly one instrument on each note, which makes an arrangement a permutation — and the commonest operation in orchestration is a doubling, which does not. Four players on three notes give thirty-six arrangements instead of six, and the extra choice turns out to be the smallest thing on the page: which note carries the pair explains three per cent of the spread and which players sit on the bass explains eighty-nine. A doubled note can be priced as one player, and which one is not the one a spectral account would have named.

instruments · Orchestration
A general pause of a bar is worth 4.9 decibels in a hall and 8.1 in silence. What a rest is worth as a diminuendo, in rooms with different reverberation times. The upper curve is the earlier figure, which assumed the sound stops when the players do; each curve below it lets the hall go on sounding, falling sixty decibels in its own reverberation time until it reaches the background. At the bar of silence a general pause usually is — about 1.2 seconds — the dry value is 8.1 decibels, shoebox concert hall keeps 60 per cent of it and gothic cathedral keeps 22. At 3.5 seconds, where a gap stops being a pause and becomes an ending, the same hall keeps 86 per cent. A long silence outlives any hall's tail and a short one does not, which is why the room costs the device most at exactly the length a composer writes it.

A rest needs a dry room

The twenty-four decibels a general pause is worth assume the sound stops when the players do. Put a hall under it and a bar of silence keeps 60 per cent of its value in a shoebox concert hall and 22 per cent in a cathedral, while a three-and-a-half-second one keeps 86 and 45 — because a hall's tail has a length and a rest either outlives it or does not. A written bar of silence is worth half its dry value at 2.76 seconds of reverberation, which falls between the concert hall and the stone church.

form · Closure
A 20-decibel crescendo is 27 phons on a bass note and 20 on a high one. The same change of level, from 60 to 80 decibels, converted to loudness at each register through ISO 226's equal-loudness contours rather than at one kilohertz. The heavy curve gives each note a string spectrum, so its partials are converted in their own bands and summed; the pale one is the fundamental alone. On the spectrum-aware curve the crescendo is worth 27.3 phons at C1 and 20.3 at C7. On the fundamental alone it is 77 at C1, which is not a finding but an artefact: a 60-decibel tone at 33 hertz sits 1.8 decibels above the threshold of hearing and is very nearly nothing. The honest correction is the smaller one, and it is still a difference of 7.1 phons across the compass for a mark written in the same ink.

A subito piano is four seconds longer in the bass

Every loudness figure with time in it converts level to loudness at one kilohertz, and the equal-loudness contours say that no other frequency works that way. Joining the two sorts the published numbers into those that were about the treble and those that were not. Three move a great deal — a twenty-decibel crescendo is worth 27 phons on a bass note and 20 on a high one, and the seven seconds a subito piano takes becomes eleven and a third. Three do not move at all, and the reason they do not is the same reason in every case.

perception · Loudness
A bar of silence is worth 13.1 decibels written last and 6.0 written first. Four closing gestures, drawn against how many seconds of silence each contains. All four hold the same 12-part texture, write the same 6.0-decibel diminuendo over 4 seconds, and differ only in what order the diminuendo and the silence are written in. The quantity is how far the listener's running impression has fallen when the final chord arrives, in decibels of equivalent diminuendo. Written with the silence last, a general pause of a bar is worth 13.06 decibels and one of three and a half seconds is worth 28.3. Written with the silence first, both are worth 6.00 — exactly the diminuendo's own depth, because the music resuming after the silence puts the reference back at its own level. The silence written first is worth less than the silence written with no diminuendo at all, which reads 8.12 at a bar: a diminuendo placed after a general pause takes 2.12 decibels away and adds nothing.

A general pause is spent by the note after it

Whether a composer should write the pause before the diminuendo or after it looks like a question about how big the ensemble is. It is not. Forty decibels of ensemble are worth one decibel of silence, and the order is worth seven — because a running impression rises twenty times faster than it falls, so half a general pause is spent by ninety-seven milliseconds of sound.

form · Closure
A final chord stands above the impression for a fraction of a second. How far a final chord at the tutti's own level stands above the listener's running impression at the instant it is released, against how long it lasts, for four ways of arriving at it. Straight out of the tutti it stands above nothing at any length; after 1.2 s of silence the impression is 8.1 dB down, and the chord stands highest, 5.14 dB, when it lasts 54 ms; after 3.5 s of silence the impression is 23.7 dB down, and the chord stands highest, 14.42 dB, when it lasts 28 ms; after a 6 dB diminuendo the impression is 5.0 dB down, and the chord stands highest, 3.18 dB, when it lasts 42 ms. Every curve is level again by half a second, because the impression's attack of 99 ms catches the note's attack of 22 ms, so a held chord is released at the impression's level whatever preceded it.

A final chord stands out for a twentieth of a second

A general pause drives a listener's running impression down, and the final chord that follows is supposed to cash the fall in. It cashes in at most two thirds of it. The note's own loudness rises with a 22-millisecond constant and the impression with a 99-millisecond one, so after a bar of silence the chord stands furthest above the impression 54 milliseconds in, by 5.1 of the 8.1 decibels the silence bought, and after 206 milliseconds the two are within a phon of each other. A short stamp spends most of its life standing out; a chord held a second and a half spends a seventh of it.

form · Closure
Level does not dilute the register's roughness, it multiplies it. The mean roughness of the I – vi – IV – V – I arrivals at 4 registers, each relative to the register as written, read three ways. Level-free, the bass is 8.6 times rougher than the treble. With every note at 70 dB it is 8.6 times, the same factor, because one level rescales every pair alike. With each chord played at the level that makes it as loud as the written register's chords — 82.8 dB −2 octaves, 75.5 dB −1 octave, 70.0 dB as written, 66.8 dB +1 octave — the bass is 343 times rougher than the treble, because roughness grows with the square of the pressure and the bass needs more of it to be heard at the same loudness.

A rough arrival is rough because of its spacing

The pair the expectation essays report for every chord — how surprising it was, how rough its voicing is — has no level in it. Putting level back in answers the question it left open, and not the way it was framed. At one written dynamic the arrivals keep their order from 40 to 90 dB at three registers of four, and the bass stays 8.6 times rougher than the treble. Made equally loud, the bass has to be played 12.8 dB harder, and it is 343 times rougher: level does not explain the register's roughness away, it multiplies it.

harmony · Tonal-expectation
The thirds' difference-tone line, note by note, against the dynamic. Each note of the line the difference tone f₂ − f₁ draws under a scale in just thirds, as its level above the higher of the threshold of hearing and the primaries' masking, against the level of the primaries. The product sits 50 dB below the primaries at 60 dB and grows twice as fast as they do. C2 above the limit from 73.5 dB; A1 above the limit from 76 dB; C2 above the limit from 73.5 dB; F2 above the limit from 70 dB; G2 above the limit from 68.5 dB; F2 above the limit from 70 dB; G2 above the limit from 68.5 dB; C3 above the limit from 66 dB.

A combination-tone bass needs a forte

A scale in just thirds draws a diatonic bass line through its difference tones, and in sixths the cubic product draws one. Given the two published level laws, with their constants swept, the thirds' bass is not heard at all below primaries of about 66 dB and is heard whole only from 71 to 81. The cubic products are a different kind of object: the primaries mask them decibel for decibel as they rise, so no dynamic changes whether they are heard. Most of the thirds' inner line never is, and the sixths' bass needs a forte and a gentle law.

intervals · Combination tone
The ghost bass drops when the passage gets louder. The note the whole crowd of products names, as a multiple of the fundamental the interval implies, against how loudly the interval is played. a major third: 2.9999999999999996 times the fundamental below 70 decibels and 1 times above it, a drop of 19 semitones; a minor third: 4 times the fundamental below 62 decibels and 2 times above it, a drop of 12 semitones; a fourth: 2 times the fundamental below 72 decibels and 1 times above it, a drop of 12 semitones. Softly, only the cubic products clear their thresholds, and they are an exact series on (2p − q) times the fundamental with no gaps in it. Loudly, the difference tones fill in the low harmonics, no template on the higher note can explain them, and the fit falls. Nothing about the interval has changed; the listener is simply being given a different subset of the same harmonic series.

The ghost bass drops a twelfth at a forte

Both crowds arrive at once and every member of both is a multiple of the same absent fundamental, so a listener is never given a choice between them — only a different subset of one harmonic series at every dynamic. Softly, the subset is an exact gapless series on three times the fundamental. Loudly, the difference tones fill in the low harmonics and no template on the higher note survives them. Between 62 and 72 decibels, depending on the interval, the note the crowd names falls by an octave or a twelfth, and the two qualities of third cross at different levels.

intervals · Combination tone
The breath is the looser ceiling nearly everywhere. How long a trained singer can hold a phrase on one breath, across a compass and at four dynamics, against the 8-second ceiling the psychological present puts on the same phrase. The flow through the folds rises with pitch and with loudness, so the breath ceiling falls both ways: at 60 decibels it runs 32.6 seconds at the bottom of the compass to 21.2 at the top; at 70 decibels it runs 23.1 seconds at the bottom of the compass to 15.0 at the top; at 80 decibels it runs 16.4 seconds at the bottom of the compass to 10.6 at the top; at 90 decibels it runs 11.6 seconds at the bottom of the compass to 7.5 at the top. The shaded line is the listener's ceiling and it does not move. The breath binds only where the two lines cross — 1 of the 40 cells drawn, all of them loud and high. So the constraint everybody names when asked why a phrase is the length it is, is almost never the constraint that decides it.

The ceiling everybody names is the loose one

Ask why phrases are the length they are and the answer given is the breath. It is arithmetic — usable lung volume over the air a note costs per second — and it comes out between fifteen and twenty-three seconds at a comfortable dynamic and between seven and twelve at a loud one. The ceiling the present moment imposes, the two-to-eight seconds inside which a stretch is heard as one thing rather than as a series, is two to three times tighter at almost every note and dynamic. A singer in an adagio is not running out of breath at the phrase end. They are running out of present.

form · Phrase
A louder final chord stands higher and still stands for a fraction of a second. How far a final chord stands above the listener's running impression at the instant it is released, against how long it lasts. The chord is a struck six-note tonic; "a step louder" is the hammer velocity doubled, which raises its loudness 7.12 phons above the tutti's. At the tutti's level, after 1.2 s of silence, the impression is 8.09 dB below the chord and the chord stands highest, 5.14 dB, at 54 ms; a step louder, straight out of the tutti, the impression is 7.12 dB below the chord and the chord stands highest, 4.51 dB, at 36 ms; a step louder, after 1.2 s of silence, the impression is 15.21 dB below the chord and the chord stands highest, 9.46 dB, at 38 ms. Every curve returns to zero by half a second: the impression climbs to whatever level the chord is played at, so a louder mark is not a stand that lasts but a deeper fall to climb out of, and a silence and a mark add as depths.

A louder final chord is a deeper silence and a brighter sound

A final chord marked a step louder than the passage was supposed to stand above a listener's running impression for as long as it sounded, since the impression can climb no higher than the chord. It climbs exactly that high, and the stand closes in half a second as it always did. What a louder mark actually buys is depth — about seven phons, the same depth a second of silence buys — and a spectrum whose balance point sits most of a whole tone higher, which, unlike the stand, lasts for the whole chord.

form · Closure
Put back beside its notes, the crowd names the bass at every dynamic. A just major third on complex tones, drawn on one axis of harmonic numbers of the fundamental its ratio implies: the partials of the two played notes, and the products of those partials that clear threshold, at 55 and 80 dB. At 55 dB the products alone name 3 times the fundamental with 0 empty slots; the products and the notes together name 1 times it with 4 empty slots, and the notes alone name it with 9. At 80 dB the products alone name 1 times the fundamental with 0 empty slots; the products and the notes together name 1 times it with 0 empty slots, and the notes alone name it with 9. The soft reading on a higher note exists only when the loud notes are set aside. Taken together, the products do not decide which fundamental is named; they decide how many holes its template has.

The played notes already name the ghost bass

The products of a just third's partials, fitted on their own, name a note a twelfth above the bass when the interval is soft and drop to the bass when it is loud. Put the two played notes back beside them and the drop disappears: the notes and their products name the bass at every dynamic, because the notes' own partials are harmonics of it already. What the dynamic changes is not which note is implied but how complete its harmonic series is — nine holes from the notes alone, four when soft, none when loud.

intervals · Combination tone

Named alongside it

The objects these essays reach for when they reach for this one.

LoudnessOrchestrationClosureIntegration windowCritical bandwidthTextureRoughnessTemporal integrationCadenceMaskingNotationSilence

All concepts