Perception and the listener

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

Assumes: A page has two decibels · Loud is relative, and it comes down slowly

A page has two decibels built up a final chord one part at a time and found the whole operation worth 1.8 decibels. It ended by naming what every rung of this ladder had been assuming:

Every rung above adds parts that are already sounding by the time the loudness is measured, and the moment an orchestrator actually controls is the entry — a part beginning, into a texture that was already there.

That is a fair description of the omission. A chord with eight parts in it and a chord with seven are two sonorities, and comparing them is not the same question as asking what happens when the eighth part comes in. The second question has a moment in it.

This essay is that question, and the answer has two halves that point the same way.

An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none.
Fig. 1 A five-part texture at 62 decibels a part, with one more part entering at the same level, tried at every semitone over five octaves. The vertical axis is what the entry is worth in phons; the shaded strip is the difference limen for loudness. The median entry is inside it.

Sixty-one entries, and the middle one is inaudible as a level change

Take a five-part texture spread the way an ensemble spreads one — a bass, a tenor, and a close upper triad — and hold every part at sixty-two decibels. Then let one more part enter, at the same level as the others, and try it at every semitone from two octaves below middle C to three octaves above.

The best available entry is worth 4.2 phons and the median is worth 0.89.

A phon is a decibel at these levels, and the difference limen for loudness is about one of them. So the entry in the middle of the range is a change the listener cannot reliably detect, and only twenty-nine of the sixty-one entries clear the limen at all.

They are all in the treble. The lowest entry that is worth a whole phon is at A♭4, four hundred and fifteen hertz — above the top of the bass staff, near the top of a tenor’s range — and everything below it is inside the strip.

Put the other way round: the same increase bought by everybody already playing harder costs 0.98 decibels of extra effort, and the best entry in the whole range is worth 4.7. A player has sixty decibels to spend and an entry buys one of them.

Why the bass is worth nothing and the treble is worth something

Neither half of that is new to this ladder; what is new is that the question has been asked of a moment rather than of a comparison.

Loudness adds across critical bands and compresses inside them. A chord is not as loud as its notes is the rung that established it, and it is the reason a new part in an occupied band is worth so little: within a band the ear takes something much nearer the largest component than the sum, so the entering part’s fundamental has to find somewhere the texture is not.

And there is much less room down there than a keyboard suggests. The ear’s analysis bands are evenly spaced on the Bark scale, which is nearly linear in hertz below about five hundred and roughly logarithmic above it. A band at C2 is thirty-six semitones wide and a band at C6 is under three. So the bottom two octaves of a texture are two or three bands and the top two are twenty, and a part entering in the bass is entering somewhere the texture already is, whatever note it plays.

The same chord is louder in the treble, at the same power. A 3-note chord spanning 7 semitones, drawn at 5 registers with the total power held constant, and scored against one note carrying all of it. Low down the whole chord fits inside one critical band — 35 semitones wide at C2 — so it is analysed as one thing and buys 0.99 times the loudness. At C6 the band is 2.8 semitones, the notes are in separate bands, and the same power is 2.24 times as loud.
Fig. 2 The same fact from the other side: one chord at five registers with the total power held constant. Low down the whole chord is inside one band and is analysed as one thing; high up its notes are in separate bands and the same power is worth twice the loudness.

Those two together say that where a part enters decides almost everything and that it enters decides almost nothing, which is a strange thing to be true of an orchestra and is what the arithmetic keeps returning.

The practical form of it is a rule orchestrators already use without this reason. A piccolo entering is an event. A third bassoon is not. It has usually been explained by saying the piccolo is piercing, and on this account the piccolo is simply the only part with anywhere to go.

There is a third term that pushes the same way and is not in this arithmetic. One sound hides another, and it hides upward is the masking ladder’s first rung, and it says a part entering below a texture arrives into a threshold the texture has already raised, while a part entering above it does not. That does not change the loudness sum — a masked partial is still in the excitation — but it means the low entry is losing spectrum at the same time as it is buying no level. Both mechanisms leave the bass entry with nothing and they are independent of each other.

The two models disagree, and the disagreement is worth having

The figure draws two lines, because this ladder carries two ways of adding a sonority’s loudness and they do not agree in the bass.

The band model groups components greedily along the Bark axis and converts each group’s summed intensity to loudness once. The excitation model integrates a specific loudness along the axis, with every component casting a skirt, so nothing depends on where a boundary was drawn.

Above the tenor register they agree to within a few tenths of a phon. Below it they part company, and the way they part company is diagnostic: the band model reports twenty-four entries of the sixty-one that make the texture quieter, and the excitation model reports none.

A part cannot be added to a sonority and reduce its excitation — the pattern is a sum of skirts and every component only adds to it. So a model that says otherwise is reporting its own grouping rule. Adding a part between two existing groups can chain them into one, and one group at the combined intensity is worth less than two groups at their own, so the total falls.

This matters for reading the rung below. A page has two decibels found its build-up curve non-monotone — four parts louder than five, and than eight — and read it as the equal-loudness contours making a bass part cheap. That reading is right about small, and cannot be right about negative: a part that contributes almost nothing contributes almost nothing, not less than nothing. The non-monotonicity was the grouping rule and not the contours, and the 1.8-decibel headline survives it because a span between two ends is barely affected by what happens between them.

That is the kind of correction a second model is for, and it is why the ladder has carried both since the second rung rather than choosing.

A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance.
Fig. 3 The curve the paragraph above is about, drawn earlier: a texture built up one part at a time, and the same chord played by people from forty to a hundred decibels. The span from end to end is what that essay measured and it stands; the dips inside it are the grouping rule.

The other half is that an entry has a rise time, and it is not the slow one

An entry is not an instant. The part builds — over three milliseconds on a marimba and a hundred and ten on a sung vowel, which is the range the onset ladder measures across the families.

The running impression of loudness has a rise time of its own. It is two smoothers in series, and the published long-term attack constant is ninety-nine milliseconds — which sits inside the instruments’ range rather than under it. So there are two clocks and it is not obvious which one a listener is reading.

A part entering, and the impression arriving after it. A texture at 70 decibels with one part entering at a second and a half, building over the 60 milliseconds a flute takes and adding 6 decibels of power. The pale line is what is sounding, the dashed line is the short-term loudness of the moment, and the heavy line is the long-term loudness a listener would report as the loudness of the passage. The entry is complete in under a tenth of a second and the impression is still 63 per cent short of its new value at that point; the whole change is worth 6.00 phons when it finally arrives.
Fig. 4 One entry as three loudnesses on one time axis: what is sounding, the loudness of the moment, and the long-term loudness a listener would report as the loudness of the passage. The part is fully in before the impression has moved a third of the way.

The answer is that the listener is the slower clock, and it flattens the difference between instruments rather than reporting it.

A marimba’s entry is built in three milliseconds and the running impression covers half the distance to its new value at ninety-four. A sung vowel’s is built in a hundred and ten and the impression covers half at a hundred and fifty-eight. So a hundred and seven milliseconds of difference between the instruments arrives as sixty-four — and an entry with no rise time at all, a part that simply appears, would still take ninety-three.

Eight of the nine sources are faster than the ear’s own attack constant. Only the sung vowel is slower, and it is the one instrument whose entries a listener does hear as gradual.

Every entry takes a tenth of a second, whatever is entering. Nine sources entering a texture that was already sounding, each adding 6 decibels of power and building over its own measured rise time. The horizontal axis is that rise; the vertical is when the running impression has covered half the distance to the new loudness, and the upper marks are nine tenths. The thin diagonal is what a listener with no integration time of their own would report. A marimba rises in 3 milliseconds and is registered at 94; a sung vowel rises in 110 and is registered at 158. So 107 milliseconds of difference between instruments arrives as 64, and an instantaneous entry would still take 93 — the long-term attack constant is 99 milliseconds and 8 of the nine sources are faster than it.
Fig. 5 Nine sources entering a texture, each building over its own measured rise. The thin diagonal is what a listener with no integration time would report; the gap between it and the curves is the listener.

There is a supporting fact from the auditory-scene ladder that makes this modelling legitimate. A blown note does not start late computed how far apart a wind instrument’s own partials are as it speaks, and at a tenth of the steady amplitude the spread is between 2.2 and 6.9 milliseconds across five families. That is far inside anything the ear treats as an asynchrony, so an entering part builds as one object with one rise, and drawing it as a single ramp is not an idealisation that hides anything.

Which makes an entry a grouping event and not a dynamic one

Both halves say the same thing and it is worth stating plainly.

An entry is not detectable as a change of level in the middle of its range, and it is not fast enough to be heard as an attack on the loudness axis, because the ear’s own integrator is slower than the instrument’s.

What it is instead is an object arriving. The ear builds objects out of a signal, and the strongest cue any account of that gives is onset synchrony: partials that start together belong together, and partials that start at a different time do not. An entering part’s components all start at the same moment, and that moment is not shared with anything else in the texture.

The partials start together by a factor of four hundred. How far each partial of a struck piano string is from the first in its onset, computed from the string's own dispersion: a stiff string's high partials travel faster, so component n reaches the bridge ahead of component 1 by t₁(1 − 1/√(1 + Bn²)). The largest offset here is 119 microseconds at the sixteenth partial. The line across the top is 20 milliseconds, which is the asynchrony at which a partial stops being heard as part of the note. The margin is a factor of 168. So the cue every account of grouping calls the strongest is, for this source, unanimous: every partial votes to fuse, and nothing in the physics of the string comes near to changing that.
Fig. 6 The cue the entry is actually delivering, from the auditory-scene essays: the partials of one source start together to well inside any threshold, which is what makes them one thing and makes them a different thing from what was already sounding.

So the answer to the question the rung below posed — whether an entering part is heard as an increase in loudness or as a new object — is the second, and by a wide margin at every register. The loudness change is under the limen for half the range and never above five phons; the onset cue is unanimous everywhere.

That is a reallocation rather than a discovery. The effect an entry has is real and everybody agrees it is large. What the arithmetic says is that it is not being delivered by the axis this ladder measures, and that the ladder’s own units are the reason it took six rungs to say so.

What that leaves an orchestrator with

Three things, and none of them is the one the word entry suggests.

An entry high enough is a level change. Above about four hundred hertz the arithmetic starts to pay: a part entering at B♭6 into this texture is worth 4.2 phons, which is four or five just-noticeable steps and a real gesture. That is the piccolo, and it is the whole of what a high entry is doing that a low one is not.

An entry at any pitch is a new object. The onset cue does not weaken in the bass — it is not about the spectrum at all — so a bassoon entering under a texture is as clearly an arrival as a piccolo over it, and it is simply not an arrival on the loudness axis.

And a level change is the thing an entry is not competing with. Loud is relative, and it comes down slowly priced the dynamic devices in one currency: the ratio of the loudest moment to the running impression at that moment, which a subito forte of twenty decibels takes to 2.07. Put the median entry through the same machinery — a step of 0.89 phons, with a flute’s rise on it — and it comes to 1.03. The best entry in the range comes to 1.17. They are different orders of magnitude and they are also different kinds of event, which is why a passage can have both at once and lose neither.

The one place the three come together is the joint scoring problem. Who plays what, and how loud, is one question takes the levels as a continuous variable and searches over the assignment; this essay says why the discrete part of that problem — which parts are playing at all — is worth so little on the objective’s loudness term, and therefore why a scoring is decided almost entirely by the continuous half. An entry is a change to the texture that the loudness constraint barely notices, and that is a fact about the constraint rather than about the entry.

Sforzando, and what the impression doesOne loud note in a quiet passage, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 2.22 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.2.22×02468101205101520secondsloudness, sonessounding nowshort-term — theloudness of a notelong-term — theloudness of a passage
Fig. 7 For scale: what a sforzando does to the same three loudnesses. Everything on this page is a smaller event than this one and a different sort of event.

Which computation produced the numbers

The base texture is five parts at MIDI 48, 55, 60, 64 and 67 — a bass, a fifth above it, and a close triad on middle C — which is the spacing four-part orchestral writing has used since the classical period and is not a neutral choice. Every part is a string-like spectrum of eight partials falling as one over n, and every part including the entering one is at sixty-two decibels.

The entering part is tried at every semitone from C2 to C7. Its worth is the ratio of the texture’s loudness with it to the texture’s loudness without it, converted to phons by the standard doubling: a doubling of sones is ten phons, and near this level a phon is a decibel.

Both loudness models are this ladder’s own, unchanged. The band model is the second rung’s; the excitation model is the same rung’s smoother alternative, with each component casting the site’s own spreading function as its skirt.

The two clocks are computed by adding a power ramp to a steady base and running the published two-stage integrator over the result — twenty-two milliseconds and ninety-nine going up, fifty and two seconds coming down. The rise times are the onset ladder’s table, which gives a range for each family; the central value is used.

Where the model stops

Every part is one timbre and one level. A real entry is a trombone into strings, at whatever level the passage is at rather than at the level of the parts already there. Both of those would raise the numbers, and neither would raise them by an order of magnitude, because the constraint is the band structure rather than the source.

Sixty-two decibels is a quiet ensemble. The band widths do not move with level but the spreading function does, and a texture at ninety would mask more of an entering part’s spectrum than one at sixty-two. That pushes the entry’s worth down rather than up.

The base texture is one texture. An entry into a thin texture is worth much more than an entry into a full one, and the figure is drawn against a full one. That is the case an orchestrator is usually in and it is not the only case.

The rise is a ramp in power, not an instrument. A real onset is a spectrum changing shape as well as a level rising, and the running-loudness model reads only the level.

And ninety-nine milliseconds is a fitted constant. It comes from listeners judging the overall loudness of time-varying sounds measured in seconds, and the entries here are events of tens of milliseconds. The constant is being used near the short end of what it was fitted on.

What the picture cannot show

It cannot show attention. A listener following a line hears an entry because they were waiting for it, and no threshold decides that. The figure is about what the signal offers, not about what is taken from it.

Nor can it show what an entry is for. Most entries are not there to make the music louder; they are there to state a subject, to fill out a harmony, to change a colour. This essay prices the one thing an entry might be doing that the ladder has units for, and finds it small.

It cannot show the room. A hall answers an entry with a reverberant field that builds over a second or two, which is longer than either clock here and is the same order as the long-term release.

It cannot show a change of texture that is not an addition. A part entering while another stops is two events at once, and the arithmetic here is nested by construction.

And it cannot show the notation. A score marks an entry with nothing at all — the part simply has notes where it had rests — so whatever an entry is worth, no marking is claiming it.

Whose music, and when

The texture is generic and the levels are generic. Nothing here is a measurement of a repertoire.

The observation with a period in it is about where entries are put. A fugal exposition brings its voices in from the middle outward and its last entry is very often the bass, which on this arithmetic is the entry worth the least in loudness and exactly as much as any other in arrival. That is consistent with the device being about identity rather than about weight, which is what everybody says it is about.

The nineteenth-century orchestral tutti does the opposite and adds at the top — piccolo, triangle, upper woodwind — and that is the arrangement the numbers say is worth something on the loudness axis. Whether the composers were choosing the audible option or the bright one is not a question this figure can answer, and the two happen to coincide.

Where this ladder goes next

Six rungs. What one tone’s loudness is; what more than one tone’s is, which the critical band decides; what a passage’s is, which two constants decide; what a page’s is, which needs no performance; what a page’s is worth, which is two decibels; and now what an entry is worth, which is one, and which turns out to be a question about grouping wearing a loudness costume.

What is owed after this is the note’s own length. Every figure on this ladder, including this one, is of a sound that has been going on long enough — and the running impression has an integration time in it, so a short note and a long note at the same level are not the same loudness at all. That is not a small correction: the collection already knows that a note has to last a certain time before its pitch exists, and the loudness version of the same question has never been asked here. It needs no corpus and no listener, only the two constants this ladder already has and a note with a duration in it, and what would come out is whether the marks a score uses for how long each note sounds are dynamic marks that nobody counts.

Part 6 of 8

One essay in the series on loudness. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Critical bandwidthDifference limenDynamicsIntegration windowLoudnessOnsetOrchestrationTexture