A part entering is not a change of level
Assumes: A page has two decibels · Loud is relative, and it comes down slowly
A page has two decibels built up a final chord one part at a time and found the whole operation worth 1.8 decibels. It ended by naming what every rung of this ladder had been assuming:
Every rung above adds parts that are already sounding by the time the loudness is measured, and the moment an orchestrator actually controls is the entry — a part beginning, into a texture that was already there.
That is a fair description of the omission. A chord with eight parts in it and a chord with seven are two sonorities, and comparing them is not the same question as asking what happens when the eighth part comes in. The second question has a moment in it.
This essay is that question, and the answer has two halves that point the same way.
Sixty-one entries, and the middle one is inaudible as a level change
Take a five-part texture spread the way an ensemble spreads one — a bass, a tenor, and a close upper triad — and hold every part at sixty-two decibels. Then let one more part enter, at the same level as the others, and try it at every semitone from two octaves below middle C to three octaves above.
The best available entry is worth 4.2 phons and the median is worth 0.89.
A phon is a decibel at these levels, and the difference limen for loudness is about one of them. So the entry in the middle of the range is a change the listener cannot reliably detect, and only twenty-nine of the sixty-one entries clear the limen at all.
They are all in the treble. The lowest entry that is worth a whole phon is at A♭4, four hundred and fifteen hertz — above the top of the bass staff, near the top of a tenor’s range — and everything below it is inside the strip.
Put the other way round: the same increase bought by everybody already playing harder costs 0.98 decibels of extra effort, and the best entry in the whole range is worth 4.7. A player has sixty decibels to spend and an entry buys one of them.
Why the bass is worth nothing and the treble is worth something
Neither half of that is new to this ladder; what is new is that the question has been asked of a moment rather than of a comparison.
Loudness adds across critical bands and compresses inside them. A chord is not as loud as its notes is the rung that established it, and it is the reason a new part in an occupied band is worth so little: within a band the ear takes something much nearer the largest component than the sum, so the entering part’s fundamental has to find somewhere the texture is not.
And there is much less room down there than a keyboard suggests. The ear’s analysis bands are evenly spaced on the Bark scale, which is nearly linear in hertz below about five hundred and roughly logarithmic above it. A band at C2 is thirty-six semitones wide and a band at C6 is under three. So the bottom two octaves of a texture are two or three bands and the top two are twenty, and a part entering in the bass is entering somewhere the texture already is, whatever note it plays.
Those two together say that where a part enters decides almost everything and that it enters decides almost nothing, which is a strange thing to be true of an orchestra and is what the arithmetic keeps returning.
The practical form of it is a rule orchestrators already use without this reason. A piccolo entering is an event. A third bassoon is not. It has usually been explained by saying the piccolo is piercing, and on this account the piccolo is simply the only part with anywhere to go.
There is a third term that pushes the same way and is not in this arithmetic. One sound hides another, and it hides upward is the masking ladder’s first rung, and it says a part entering below a texture arrives into a threshold the texture has already raised, while a part entering above it does not. That does not change the loudness sum — a masked partial is still in the excitation — but it means the low entry is losing spectrum at the same time as it is buying no level. Both mechanisms leave the bass entry with nothing and they are independent of each other.
The two models disagree, and the disagreement is worth having
The figure draws two lines, because this ladder carries two ways of adding a sonority’s loudness and they do not agree in the bass.
The band model groups components greedily along the Bark axis and converts each group’s summed intensity to loudness once. The excitation model integrates a specific loudness along the axis, with every component casting a skirt, so nothing depends on where a boundary was drawn.
Above the tenor register they agree to within a few tenths of a phon. Below it they part company, and the way they part company is diagnostic: the band model reports twenty-four entries of the sixty-one that make the texture quieter, and the excitation model reports none.
A part cannot be added to a sonority and reduce its excitation — the pattern is a sum of skirts and every component only adds to it. So a model that says otherwise is reporting its own grouping rule. Adding a part between two existing groups can chain them into one, and one group at the combined intensity is worth less than two groups at their own, so the total falls.
This matters for reading the rung below. A page has two decibels found its build-up curve non-monotone — four parts louder than five, and than eight — and read it as the equal-loudness contours making a bass part cheap. That reading is right about small, and cannot be right about negative: a part that contributes almost nothing contributes almost nothing, not less than nothing. The non-monotonicity was the grouping rule and not the contours, and the 1.8-decibel headline survives it because a span between two ends is barely affected by what happens between them.
That is the kind of correction a second model is for, and it is why the ladder has carried both since the second rung rather than choosing.
The other half is that an entry has a rise time, and it is not the slow one
An entry is not an instant. The part builds — over three milliseconds on a marimba and a hundred and ten on a sung vowel, which is the range the onset ladder measures across the families.
The running impression of loudness has a rise time of its own. It is two smoothers in series, and the published long-term attack constant is ninety-nine milliseconds — which sits inside the instruments’ range rather than under it. So there are two clocks and it is not obvious which one a listener is reading.
The answer is that the listener is the slower clock, and it flattens the difference between instruments rather than reporting it.
A marimba’s entry is built in three milliseconds and the running impression covers half the distance to its new value at ninety-four. A sung vowel’s is built in a hundred and ten and the impression covers half at a hundred and fifty-eight. So a hundred and seven milliseconds of difference between the instruments arrives as sixty-four — and an entry with no rise time at all, a part that simply appears, would still take ninety-three.
Eight of the nine sources are faster than the ear’s own attack constant. Only the sung vowel is slower, and it is the one instrument whose entries a listener does hear as gradual.
There is a supporting fact from the auditory-scene ladder that makes this modelling legitimate. A blown note does not start late computed how far apart a wind instrument’s own partials are as it speaks, and at a tenth of the steady amplitude the spread is between 2.2 and 6.9 milliseconds across five families. That is far inside anything the ear treats as an asynchrony, so an entering part builds as one object with one rise, and drawing it as a single ramp is not an idealisation that hides anything.
Which makes an entry a grouping event and not a dynamic one
Both halves say the same thing and it is worth stating plainly.
An entry is not detectable as a change of level in the middle of its range, and it is not fast enough to be heard as an attack on the loudness axis, because the ear’s own integrator is slower than the instrument’s.
What it is instead is an object arriving. The ear builds objects out of a signal, and the strongest cue any account of that gives is onset synchrony: partials that start together belong together, and partials that start at a different time do not. An entering part’s components all start at the same moment, and that moment is not shared with anything else in the texture.
So the answer to the question the rung below posed — whether an entering part is heard as an increase in loudness or as a new object — is the second, and by a wide margin at every register. The loudness change is under the limen for half the range and never above five phons; the onset cue is unanimous everywhere.
That is a reallocation rather than a discovery. The effect an entry has is real and everybody agrees it is large. What the arithmetic says is that it is not being delivered by the axis this ladder measures, and that the ladder’s own units are the reason it took six rungs to say so.
What that leaves an orchestrator with
Three things, and none of them is the one the word entry suggests.
An entry high enough is a level change. Above about four hundred hertz the arithmetic starts to pay: a part entering at B♭6 into this texture is worth 4.2 phons, which is four or five just-noticeable steps and a real gesture. That is the piccolo, and it is the whole of what a high entry is doing that a low one is not.
An entry at any pitch is a new object. The onset cue does not weaken in the bass — it is not about the spectrum at all — so a bassoon entering under a texture is as clearly an arrival as a piccolo over it, and it is simply not an arrival on the loudness axis.
And a level change is the thing an entry is not competing with. Loud is relative, and it comes down slowly priced the dynamic devices in one currency: the ratio of the loudest moment to the running impression at that moment, which a subito forte of twenty decibels takes to 2.07. Put the median entry through the same machinery — a step of 0.89 phons, with a flute’s rise on it — and it comes to 1.03. The best entry in the range comes to 1.17. They are different orders of magnitude and they are also different kinds of event, which is why a passage can have both at once and lose neither.
The one place the three come together is the joint scoring problem. Who plays what, and how loud, is one question takes the levels as a continuous variable and searches over the assignment; this essay says why the discrete part of that problem — which parts are playing at all — is worth so little on the objective’s loudness term, and therefore why a scoring is decided almost entirely by the continuous half. An entry is a change to the texture that the loudness constraint barely notices, and that is a fact about the constraint rather than about the entry.
Which computation produced the numbers
The base texture is five parts at MIDI 48, 55, 60, 64 and 67 — a bass, a fifth above it, and a close triad on middle C — which is the spacing four-part orchestral writing has used since the classical period and is not a neutral choice. Every part is a string-like spectrum of eight partials falling as one over n, and every part including the entering one is at sixty-two decibels.
The entering part is tried at every semitone from C2 to C7. Its worth is the ratio of the texture’s loudness with it to the texture’s loudness without it, converted to phons by the standard doubling: a doubling of sones is ten phons, and near this level a phon is a decibel.
Both loudness models are this ladder’s own, unchanged. The band model is the second rung’s; the excitation model is the same rung’s smoother alternative, with each component casting the site’s own spreading function as its skirt.
The two clocks are computed by adding a power ramp to a steady base and running the published two-stage integrator over the result — twenty-two milliseconds and ninety-nine going up, fifty and two seconds coming down. The rise times are the onset ladder’s table, which gives a range for each family; the central value is used.
Where the model stops
Every part is one timbre and one level. A real entry is a trombone into strings, at whatever level the passage is at rather than at the level of the parts already there. Both of those would raise the numbers, and neither would raise them by an order of magnitude, because the constraint is the band structure rather than the source.
Sixty-two decibels is a quiet ensemble. The band widths do not move with level but the spreading function does, and a texture at ninety would mask more of an entering part’s spectrum than one at sixty-two. That pushes the entry’s worth down rather than up.
The base texture is one texture. An entry into a thin texture is worth much more than an entry into a full one, and the figure is drawn against a full one. That is the case an orchestrator is usually in and it is not the only case.
The rise is a ramp in power, not an instrument. A real onset is a spectrum changing shape as well as a level rising, and the running-loudness model reads only the level.
And ninety-nine milliseconds is a fitted constant. It comes from listeners judging the overall loudness of time-varying sounds measured in seconds, and the entries here are events of tens of milliseconds. The constant is being used near the short end of what it was fitted on.
What the picture cannot show
It cannot show attention. A listener following a line hears an entry because they were waiting for it, and no threshold decides that. The figure is about what the signal offers, not about what is taken from it.
Nor can it show what an entry is for. Most entries are not there to make the music louder; they are there to state a subject, to fill out a harmony, to change a colour. This essay prices the one thing an entry might be doing that the ladder has units for, and finds it small.
It cannot show the room. A hall answers an entry with a reverberant field that builds over a second or two, which is longer than either clock here and is the same order as the long-term release.
It cannot show a change of texture that is not an addition. A part entering while another stops is two events at once, and the arithmetic here is nested by construction.
And it cannot show the notation. A score marks an entry with nothing at all — the part simply has notes where it had rests — so whatever an entry is worth, no marking is claiming it.
Whose music, and when
The texture is generic and the levels are generic. Nothing here is a measurement of a repertoire.
The observation with a period in it is about where entries are put. A fugal exposition brings its voices in from the middle outward and its last entry is very often the bass, which on this arithmetic is the entry worth the least in loudness and exactly as much as any other in arrival. That is consistent with the device being about identity rather than about weight, which is what everybody says it is about.
The nineteenth-century orchestral tutti does the opposite and adds at the top — piccolo, triangle, upper woodwind — and that is the arrangement the numbers say is worth something on the loudness axis. Whether the composers were choosing the audible option or the bright one is not a question this figure can answer, and the two happen to coincide.
Where this ladder goes next
Six rungs. What one tone’s loudness is; what more than one tone’s is, which the critical band decides; what a passage’s is, which two constants decide; what a page’s is, which needs no performance; what a page’s is worth, which is two decibels; and now what an entry is worth, which is one, and which turns out to be a question about grouping wearing a loudness costume.
What is owed after this is the note’s own length. Every figure on this ladder, including this one, is of a sound that has been going on long enough — and the running impression has an integration time in it, so a short note and a long note at the same level are not the same loudness at all. That is not a small correction: the collection already knows that a note has to last a certain time before its pitch exists, and the loudness version of the same question has never been asked here. It needs no corpus and no listener, only the two constants this ladder already has and a note with a duration in it, and what would come out is whether the marks a score uses for how long each note sounds are dynamic marks that nobody counts.
Part 6 of 8
One essay in the series on loudness. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Critical bandwidthDifference limenDynamicsIntegration windowLoudnessOnsetOrchestrationTexture
- A general pause is spent by the note after it dynamics, integration window, loudness, orchestration, texture
- The dissonance arrives and the dynamic does not critical bandwidth, dynamics, integration window, loudness, orchestration
- The dynamics are in the score already critical bandwidth, dynamics, loudness, orchestration, texture
- A final chord is not made loud by adding to it dynamics, loudness, orchestration, texture
- A soft chord has to fade in dynamics, integration window, loudness, orchestration
- A subito piano is a rate, not a level dynamics, loudness, orchestration, texture