Perception and the listener

A staccato is a dynamic mark

Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.

Assumes: A part entering is not a change of level · Loud is relative, and it comes down slowly

This collection has a way of finding the next rung of a ladder, and it is to look for the parameter every figure on it sets to the same value. How long a note has to be is the worked example: every difference limen on the pitch ladder had been measured on a tone lasting as long as the listener needed, nobody had said so, and varying the duration moved the quoted five-cent limen to nineteen.

Run the same search over this ladder and the answer is the same word. Every loudness figure here is of a sound that has been going on long enough. A chord, a section, a tutti, a dynamic curve — the levels change and the durations never do, because there are no durations in any of them.

Music has nothing else. A note has a length, the length is written in the score, and the running-loudness model has an integration time in it.

Articulation is worth 2.6 phons, and nobody counts it. A passage of notes at 80 decibels, 2 to the beat, at 7 tempi and 3 articulations, scored against the same level held continuously. The variable is the fraction of each inter-onset interval that is sounding — 0.95 is a legato, 0.4 a staccato — and the vertical axis is what that costs the passage's running loudness in phons. Nothing here is anybody playing harder or softer. At 40 to the beat the span from legato to staccato is 1.47 phons; at 200 it is 2.61, because a staccato note there lasts 60 milliseconds and no longer reaches its own loudness either.
Fig. 1 A passage of quavers at one level, at seven tempi and three articulations, against the same level held continuously. The variable is the fraction of each beat that is sounding. Nobody is playing harder or softer anywhere on this figure.

What a note of a stated length reaches

Start with one note, because the passage is made of them.

The published model is two one-pole smoothers in series applied to the instantaneous loudness, and their attack constants differ by a factor of four and a half: twenty-two milliseconds for the short-term loudness, which is the loudness of the note, and ninety-nine for the long-term loudness, which is what a listener reports as the loudness of the passage.

Those two constants give two different answers to how long a note has to be.

The note’s own loudness is nine tenths of the way there by fifty milliseconds. Below that a note is audibly quieter than its own level, which is a real effect and a narrow one — fifty milliseconds is a semiquaver at three hundred to the crotchet.

Its contribution to the passage’s loudness needs two hundred and forty-five. A note of a hundred and fifty milliseconds delivers three quarters of it and a note of fifty delivers two fifths.

A note reaches its own loudness at 50 ms and the passage's at 245. One note at 80 decibels against a 35-decibel floor, at durations from ten milliseconds to two seconds, with each of the two published loudness integrators drawn as a fraction of what the same level reaches when it is held. The short-term loudness — the loudness of the note itself — is nine tenths of the way there by 50 milliseconds. The long-term loudness, which is what the note contributes to the loudness of the passage it is in, needs 245. A note of a hundred and fifty milliseconds — a semiquaver at a hundred to the crotchet — reaches 0.75 of it, and one of fifty reaches 0.40.
Fig. 2 One note at eighty decibels, at durations from ten milliseconds to two seconds, with each integrator drawn as a fraction of what the same level reaches when it is held. The gap between the two curves is the whole of this essay.

Two hundred and forty-five milliseconds is a quaver at a hundred and twenty to the crotchet. So the ordinary note of ordinary music is right on the boundary, and everything shorter than it is delivering less to the passage’s loudness than its level says it should.

Which makes an articulation a dynamic

A passage is not one note. It is notes with gaps between them, and the gaps are what an articulation mark is an instruction about.

Take a stream of quavers at eighty decibels, hold the level absolutely fixed, and vary only the fraction of each inter-onset interval that is sounding: 0.95 for a legato, 0.7 for an ordinary detached playing, 0.4 for a staccato.

beats a minute legato ordinary staccato span
40 −0.04 −0.45 −1.51 1.47
120 −0.04 −0.42 −1.86 1.82
200 −0.04 −0.57 −2.65 2.61

The numbers are phons against the same level held continuously, and near this level a phon is a decibel.

An articulation mark commands between one and a half and two and two thirds decibels of the passage’s loudness, at a fixed dynamic.

That is a larger number than it looks, because of what it is next to. A page has two decibels measured what a score’s texture can do — every part it can add or take away, one at a time, at a fixed level — and got 1.8 decibels. So on the same page, in the same ink, the articulation marks are worth as much as the whole scoring, and more of it at speed.

The scoring is what an orchestration textbook is about. The articulation is a dot.

And the second mechanism arrives at speed

The table’s rows are not parallel, and the way they fail to be parallel says there are two things happening rather than one.

At forty to the crotchet a staccato quaver lasts three hundred milliseconds, which is past the boundary above, so the note reaches its own loudness in full and the entire loss is the gaps — the long-term integrator averaging in four hundred and fifty milliseconds of near-silence between each note and the next, with a release constant of two seconds that will not let it fall very far.

At two hundred a staccato quaver lasts sixty milliseconds. Now the note is inside the boundary too, and it never reaches its own loudness before it stops. The two losses add, and the span widens from 1.47 to 2.61.

Push the subdivision and the second mechanism takes over completely.

Semiquavers, where the note runs out before the gap does. A passage of notes at 80 decibels, 4 to the beat, at 7 tempi and 4 articulations, scored against the same level held continuously. The variable is the fraction of each inter-onset interval that is sounding — 0.95 is a legato, 0.4 a staccato — and the vertical axis is what that costs the passage's running loudness in phons. Nothing here is anybody playing harder or softer. At 50 to the beat the span from legato to staccato is 3.58 phons; at 230 it is 8.12, because a staccato note there lasts 16 milliseconds and no longer reaches its own loudness either.
Fig. 3 The same computation on semiquavers rather than quavers, with the articulation taken from a full legato down to a quarter of each note’s own space. At two hundred and thirty to the crotchet the span is 8.1 phons.

At two hundred and thirty to the crotchet a semiquaver’s own space is sixty-five milliseconds and a quarter of it is sixteen. A note of sixteen milliseconds on its own reaches just over half of its own loudness and a fifth of what it would give the passage, and the span from a full legato to a quarter articulation at that tempo is 8.1 phons.

Eight phons is a forte against a piano. It is written with a dot.

What a phon is worth, because the unit is doing work here

Everything above is in phons and the argument depends on how big one is, so it is worth pinning down.

A phon is a decibel by definition at one kilohertz, which is where these numbers are computed. The loudness in sones doubles for every ten phons above about forty, which means the 8.1 phons an articulation commands at speed is a factor of 1.75 in sones — the passage sounds not quite twice as loud legato as it does at a quarter articulation, with nobody having played a note differently.

The other end matters more. The difference limen for loudness is about one decibel, so the 1.47 phons an articulation commands at a slow tempo is one and a half just-noticeable steps: audible, certainly, and small. And the 0.04 of the legato line is a twenty-fifth of a limen, which is a clean zero.

Twice as loud is ten decibels, not twice the pressure. Loudness in sones against loudness level in phons, from Stevens's power law: above 40 phons the sone value doubles for every ten phons. Ten equal sources are ten times the power and about ten decibels, so they sound roughly twice as loud as one — which is why a section of ten violins is not ten violins loud.
Fig. 4 The scale everything on this page is stated in, established earlier: loudness in sones against loudness level in phons, with the number of equal sources it takes to reach each. Ten players are ten times the power and about twice the loudness.

Read against the same scale, the dynamic range a player has is about sixty decibels and the range a whole ensemble has is a little more. Articulation is therefore between a fortieth and an eighth of what a performance commands — which is small, and is the same order as what the whole scoring commands, and is the point.

The nulls, and both of them are informative

Two of the sweeps return nothing and each says why.

The legato line is flat. At a duty of 0.95 the loss is 0.04 phons and it is 0.04 at every tempo from forty to two hundred. That is the right answer and it is worth stating: the ear’s release constant is two seconds, so a gap of a few tens of milliseconds is not a gap as far as the running impression is concerned. The effect is entirely in how much silence there is, not in how it is arranged.

And the tempo on its own does almost nothing. Hold the articulation at an ordinary 0.7 and take the tempo from forty to two hundred — a factor of five — and the loss moves from 0.45 to 0.57 phons. An eighth of a decibel. Fast music at a fixed dynamic and a fixed articulation is not quieter than slow music.

So the variable is the articulation and not the speed. The speed only matters because it decides how long a stated fraction of a beat lasts, which is exactly the sense in which the first figure’s rows fan out rather than shift.

That is a useful narrowing, because “fast music is quieter” is the kind of claim that would be easy to make from the first figure alone and it is false.

Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 3 parts to 8 adds 4.3 decibels of power and 0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 6.0 to 3.9 as the parts crowd into the same register.
Fig. 5 The other half of what a page can do, drawn earlier: the dynamic curve of a written form, from its texture alone. This is the 1.8 decibels the articulation is being measured against.

What this does to the ladder’s own claim

The fourth rung computed a passage’s loudness from a score with no performance in it, and the fifth priced what that was worth and got two decibels. Both of them read the score for one thing: how many parts are sounding, and where.

They were reading half the page. A score also says how long each note lasts, and that is worth as much again — so a written dynamic reading that ignores articulation is understating the range a page commands by about a factor of two, and by a great deal more in a fast movement.

The correction is not a scaling and it does not commute with the first. Texture and articulation are independent variables: a passage can thin and shorten, in which case they add, or thin and lengthen, in which case they cancel. A composer writing a diminuendo without a marking has two devices and has always had them.

They are also independent of the third thing a page can do, which a part entering is not a change of level priced at one decibel and found to be almost entirely a grouping event rather than a dynamic one. Putting the three together, a score commands about two decibels through its texture, two to eight through its articulation, and one through each entry — and the middle term, which is the largest of the three, is the one no account of written dynamics mentions. That ordering is the finding, and it survives any reasonable argument about the exact numbers, because the terms are so far apart.

It is worth being clear about which direction the correction runs. Nothing here makes a passage louder; every number is a loss against the same level held. So what a page commands is a range downward from its own legato, and the forte end of a written dynamic is always the sustained end. That is consistent with what the critical band does to a chord, where the loudest arrangement of a fixed power is the widest one and every other arrangement is a reduction from it: in both cases the score’s dynamic control is a set of ways to spend less than it has.

It also changes what the fifth rung’s refusal was about. That essay concluded that the dynamic component of an ending is a performance quantity because a score contains markings and textures and neither is a loudness. A score contains a third thing — the note lengths — and that one is a loudness, computed here without any performance data at all. The refusal narrows: what a score cannot supply is the level, and it can supply rather more of the shape than the fifth rung credited it with.

A step up is absorbed in a fifth of a second and a step down takes seven. How long the running impression of loudness takes to come within a tenth of a step of a stated size, drawn separately for a step up and a step down. The two curves are the same model with the same input and differ only in which of two published time constants applies — 99 milliseconds while the loudness is rising and 2 seconds while it is falling. A 25-decibel step is absorbed in 0.23 s going up and 7.7 s coming down, a ratio of 34 to one.
Fig. 6 The two constants this whole essay is made of, established earlier: how long the running impression takes to absorb a step of a stated size, up and down. The two-second release is why a legato passage’s gaps are invisible and a staccato passage’s are not.

What a performer is trading

There is a second collection of numbers about short notes here and it points the other way, which makes the pair worth putting together.

An equal note cannot be masked found that a detached note buys back partials that a legato note loses. A legato predecessor ends exactly when its successor begins, so the successor arrives into the highest threshold the decay ever offers; a gap of thirty-seven milliseconds takes that residual threshold from sixty decibels to twenty-seven.

So detaching a note costs loudness and buys spectrum. Both are computed on this site, from different literatures, and they are a genuine trade rather than two accounts of one thing: the loudness loss is the ear’s integrator and the spectral gain is the ear’s forward masking, and the two mechanisms have nothing to do with each other.

That is a reasonable description of what staccato sounds like — thinner and clearer at once — and neither half of it is the usual explanation, which is that the notes are shorter.

There is a third term in the same trade that this ladder cannot price. The shape of a note is the collection’s account of what an envelope is, and a short note is not merely a long note stopped early: the attack occupies a larger fraction of it, so its average spectrum is brighter as well as its total being quieter. Whether a listener trades brightness against loudness is a question about neither integrator, and this ladder has no term for it at all.

What each dynamic device buys, against the running impression. Six dynamic devices, each scored twice: how much louder the loudest moment is than the running impression at that moment, and how much quieter the quietest is. A crescendo of twenty decibels spread over eight seconds buys a contrast of 1.02 — the impression tracks it almost exactly — while the same twenty decibels taken as a step buys 2.22. The largest number in the figure is a step down rather than any step up, which is the asymmetry in the two time constants read as a piece of orchestration advice.
Fig. 7 For scale: what each written dynamic device buys against the running impression. A crescendo of twenty decibels spread over eight seconds buys a contrast of 1.02, which is less than a change of articulation does.

Which computation produced the numbers

The passage is a stream of identical notes at eighty decibels against a thirty-decibel floor, rectangular in level, with the sounding fraction of each inter-onset interval as the only variable. Ten seconds of it are run and the mean long-term loudness over the last four is taken, so nothing depends on the attack at the beginning.

The single-note figure is the same model with one note in it against a thirty-five-decibel floor, and the quantity plotted is the peak each integrator reaches as a fraction of what the same level reaches when it is held indefinitely.

The integrator is the published one, unchanged: two one-pole smoothers in series with different constants each way, twenty-two and ninety-nine milliseconds rising and fifty milliseconds and two seconds falling. The conversion to phons is the standard doubling — ten phons per doubling of sones.

The level is converted to loudness at one kilohertz throughout, which is the model’s own convention and is discussed under what the picture cannot show.

Where the model stops

A note that stops is not what most instruments do. A piano’s note decays for seconds and a bowed note stops when the bow does. The rectangular note here is a wind instrument or a voice; on a piano the sounding fraction of a beat is not something the player controls at all above a certain tempo, and the whole effect would be smaller.

And a room does not stop either. A hall keeps every note going for a second or two, which fills the gaps this rung is entirely about. In a reverberant space the articulation should be worth much less, and in a dry one much more — which is a prediction about halls this collection can make and has not tested.

The floor is a choice. Thirty decibels below the notes is a quiet room. A noisier floor would raise the staccato passage’s loudness and shrink the span, because the running impression would have less far to fall in each gap.

A real staccato is not the same note shortened. A player asked for staccato usually plays with a sharper attack and often a little louder, so the level and the duration move together and this figure separates something a performance does not.

And the model reads only level. A short note has a different spectrum from a long one — its attack transient is a larger fraction of it — and none of that reaches the arithmetic here.

What the picture cannot show

It cannot show what an articulation is for. Staccato marks phrasing, character and attack as much as length, and the loudness consequence computed here is a side effect of one of those rather than the point of any of them.

Nor can it show a listener who knows the piece. The running impression is a model of a reported number, fitted on listeners judging unfamiliar sounds measured in seconds, and a listener following a phrase they can predict is not obviously doing the same thing.

It cannot show notation’s own hierarchy. A dot over a note is not read as a dynamic instruction by anyone, and the fact that it functions as one does not make it one. The mark that is not a level argues that even the marks that are dynamic instructions are ordinal rather than metric, and a dot is not on that scale at all.

And it cannot show the register. Everything here is computed at one kilohertz, and the equal-loudness contours are the one object in this subject whose whole content is that level and frequency do not separate. A staccato in the bass and a staccato in the treble are the same figure here and are almost certainly not the same event.

Whose music, and when

The passage is generic. Nothing here is a measurement of a repertoire.

The observation with a period in it is about when articulation is notated. Detailed articulation marking is a nineteenth-century habit; before it, the length of a note was a matter of convention, instrument and style, and a great deal of what a modern edition prints as a slur or a dot is an editor’s. On this arithmetic the dynamic content of a piece with unmarked articulation is genuinely underdetermined by up to eight decibels, which is more than the notated dynamic range of most eighteenth-century scores.

That is not an argument for any particular performance practice. It is an argument that the question of how long the notes are is a dynamic question, and it is one that the repertoire’s own notation answered later than it answered the level.

Where this ladder goes next

Seven rungs. What one tone’s loudness is; what more than one tone’s is; what a passage’s is, which two constants decide; what a page’s is, from the texture alone; what a page’s is worth, which is two decibels; what an entry is worth, which is one; and now what the note lengths are worth, which is two and sometimes eight.

What the ladder owes now is the register, and the debt is easy to state and slightly embarrassing. Every figure on this ladder that has time in it — every running loudness, every settling time, every number in this essay — converts level to loudness at one kilohertz, because that is the convention the published model is stated in. The first rung of this same ladder is the equal-loudness contours, whose entire content is that the map from level to loudness is a different map at every frequency, and whose curves are further apart at the bottom of the spectrum than in the middle. So a twenty-decibel crescendo on a double bass is not the same loudness event as a twenty-decibel crescendo on a flute, a subito piano should take a different length of time to fade in each register, and the two halves of this ladder have never been introduced. It needs no corpus and no listener — the contours are tabulated and the integrator is four lines — and what would come out is whether every dynamic number this ladder has published is a number about the treble.

Part 7 of 8

One essay in the series on loudness. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DurationDynamicsEnvelopeIntegration windowLoudnessNotationTempoTemporal integration