Form and structure

The reading was a step response

Sweeping the tempo found it did not decide the answer. This one sweeps the deceleration across a factor of three and finds a null to five figures, and then sweeps the length of the closing gesture across a factor of forty-eight and finds it moves the reading by eight per cent — but not as a function of seconds. Sorted by seconds the twelve runs scatter; sorted by how many bars the instruction covers they fall into three tight groups. One sentence explains the null and the not-null together.

Assumes: The parameter that did not decide the answer · An ending is a deceleration

The parameter that did not decide the answer swept the tempo across a factor of ten and found the closing dynamic reading moved by four per cent — while the number of closing bars, swept beside it without being named, moved it by twenty. It ended by naming what it had not tried:

Holding the tempo fixed across the closing bars is now known to be nearly harmless; letting it move is a different question, because a ritardando lengthens exactly the bars in which the gesture is happening.

It is a different question and it has the same answer, to five figures. The sweep beside it — the length of the closing gesture — has the opposite answer, and the useful thing about this rung is that one sentence accounts for both.

A ritardando does not spend the diminuendo. The arrival reading under a deceleration into the ending, from no ritardando at all to a final tempo 30 per cent of the starting one — which stretches the closing bars from 12.0 seconds to 20.2. The expectation was that it would matter: a ritardando lengthens exactly the bars the gesture is happening in, so a diminuendo that would have been absorbed at a steady tempo gets more of the smoother's own time to be absorbed in. It moves the reading by 0.00 per cent. Every line here is flat to within the thickness of the line, which is the second time a tempo parameter has been swept here and found to do nothing.
Fig. 1 The arrival reading under a deceleration into the ending, from no ritardando at all to a final tempo a third of the starting one. Every line is flat to within its own thickness.

The question, which was a good one

The reasoning behind the ritardando question was sound and it is worth restating, because a sound argument that fails is more useful than a vague one that succeeds.

The reading being taken is the ratio of what is sounding at the arrival of the final chord to what the listener has been hearing — the running impression carried by the loudness ladder’s long-term smoother, whose release is two seconds. Loud is relative, and it comes down slowly is the essay that measures that release, and the whole of the sixth and seventh rungs of this ladder is about racing it.

A ritardando lengthens the closing bars. If a diminuendo takes six bars and those bars are stretched from twelve seconds to twenty, the smoother gets eight more seconds to follow the diminuendo down — so the impression at the arrival should be lower and the ratio higher.

The prediction is right in direction and wrong in size, and “wrong in size” is an understatement. A ritardando taking the closing bars from twelve seconds to twenty leaves the reading at 0.76415 in every one of the eight runs — not approximately, identically, to every figure the arithmetic carries. The same holds for the other three textures: 0.92932, 0.95218, 1.00000, eight times each.

And the length of the gesture does, but not in seconds

The obvious next move is to shorten the thing. If twelve seconds is far too long to race a two-second release, then a shorter gesture should race it.

Sweeping the closing gesture from half a second to twenty-four — a factor of forty-eight, spanning from a quarter of a release to twelve of them — moves the reading from 0.934 to 1.008, which is eight per cent and is not nothing. But plotted against seconds the twelve runs do not lie on a curve. They lie in a scatter: 0.934 at half a second, 0.934 at two and a half, 0.970 at two, 1.001 at four, 0.934 at six, 1.008 at twenty-four.

A quantity that is 0.934 at six seconds and 1.001 at four is not a function of seconds. The sweep turned two dials at once — four bar counts against three bar lengths — and sorting the same twelve runs by the other dial resolves the scatter completely:

closing bars at 0.5 s a bar at 1.2 s at 3.0 s
1 0.934 0.934 0.934
2 0.934 0.934 0.934
4 0.970 1.000 1.002
8 1.001 1.006 1.008

Three tight rows and a sixfold change of tempo inside each of them. The gesture’s length in bars moves the reading by eight per cent; its length in seconds moves it, within a bar count, by three per cent at worst and three hundredths of one at best.

That is the point at which the sweep stops being about music and starts being about the instrument. Two parameters were turned; one of them is not in the quantity at all and the other is — and the one that is not in it is the one made of seconds, which is the one a smoother with a two-second release was supposed to care about.

The length of the gesture does not decide it either. The arrival reading against how long the closing gesture lasts, from 0.5 seconds to 24 — a factor of 48, spanning from well inside the running impression's two-second release to twelve times outside it. The reading moves by 6.6 per cent at most. That is the second null in a row and it is what makes the third figure necessary: two sweeps of two different parameters returning the same number to four digits is not a fact about endings, it is a fact about the quantity being measured.
Fig. 2 The arrival reading against how long the closing gesture lasts, from half a second to twenty-four. The vertical line is the smoother’s own two-second release, and the points do not know it is there — but they are not on a curve either. Each of the three groups is one bar count read at three tempi, and the grouping is the finding.

What the quantity actually is

One line accounts for the null, for the grouping, and for the size of both.

The reading can see one thing: how far the final bar’s texture stands from the bar before it. A ritardando changes how long the earlier bars lasted and leaves that difference alone, so it is a null. A bar count changes which bar is second-to-last, and therefore changes the difference itself: a drop-a-part-a-bar instruction from six parts leaves the final chord standing against six parts over one or two closing bars, against four parts over four, and against a single line over eight. Three second-to-last textures, three rows in the table, and the tempo cannot reach any of them because a part count is not a duration.

Sweeping the one parameter nobody had varied confirms it, and does so in the form the account predicts.

The reading is taken a fixed 0.3 seconds after the final chord arrives. Sweeping that delay from 0.05 seconds to four gives an exponential that flattens exactly where the release says it should: the thinning ending reads 0.796 at fifty milliseconds, 0.758 at two hundred, 0.764 at three hundred, 0.882 at two seconds, and 0.882 for ever after.

That is a step response sampled at a delay, and its time constant is the release. A shape read over six bars would have no such constant in it.

So the quantity is: the size of the last step, read at a stated point on the smoother’s own exponential. The size of the step is set by what the final chord does relative to the bars before it — a fortissimo arrival after a diminuendo is a big step, an unchanged texture is no step. The point on the exponential is set by the delay.

Which makes the delay a resolution knob rather than a convention, and it is worth asking where it is sharpest. Reading the four endings at each delay and taking the spread between the widest pair as how much the measurement can distinguish:

delay spread between the four endings
0.02 s 0.1425
0.05 0.2045
0.10 0.2304
0.20, the maximum 0.2421
0.30, the value used 0.2359
0.50 0.2187
0.80 0.1942
2.00 and beyond 0.1178

The curve has an interior maximum and the reason is that two exponentials are running against each other. Too early and the short-term smoother, whose attack is about a tenth of a second, has not yet arrived at the final chord’s own level; too late and the long-term smoother has begun following it. The best sample is after the first has finished and before the second has started, which is a window of two or three tenths of a second wide.

Three hundred milliseconds turns out to sit inside it, at 97 per cent of the maximum available at two hundred. That is worth saying plainly because the convention was not chosen for this reason, or for any stated reason, and it could easily have been wrong: twenty milliseconds would have cost 41 per cent of the resolution and two seconds would have cost half. The reading was lucky, and a sweep is how a lucky choice is told from a good one.

It also has the honest half. Two hundred milliseconds after a chord arrives is still inside the attack of a bowed or blown note, so a sample taken at the optimum is a sample of a moment the listener has not finished receiving — which is a reason to prefer three hundred rather than an argument for moving to two.

Nothing else enters. A ritardando changes what happens in the bars before the step and those bars are finished by the time the sample is taken. The tempo changes the same thing. Only the bar count changes the step itself.

The reading was a step response read at a delay. The same arrival reading against the one parameter nobody swept: how long after the final chord it is taken. It is an exponential with the running impression's two-second release as its time constant, and it flattens once the delay passes one release — which is what a step response does and what nothing else here does. That is the explanation of both nulls. The quantity is the sounding loudness over a running average, sampled at a fixed delay after a step, so it depends on the size of the last step and on where on the exponential the sample falls, and on nothing the music does over six bars. A ritardando and a longer gesture both change things that are already finished by the time the sample is taken. The reading is real and it measures one thing: the final chord. A piece that wanted the running impression to matter would have to put its gesture inside two seconds of the arrival, which is a bar or less at any tempo — the opposite of a ritardando.
Fig. 3 The reading against the one parameter nobody swept: how long after the final chord it is taken. It is an exponential with the running impression’s release as its time constant, and it flattens once the delay passes one release.

What that does to the sixth rung

This is an audit of the ladder’s own instrument and it is worth being clear about what survives.

A final chord is not made loud by adding to it compared four closing textures and found three of them producing a reading and one not. The three that do — the thinning to a single line, the sudden full final chord, and the diminuendo into a loud final chord — are exactly the three with a step at the arrival. The one that does not is the unchanged texture, which has none and reads 1.000 by construction.

That result stands and it is now explained rather than reported. The reading measures a step, so a texture with a step produces one and a texture without one does not — and the largest reading on that page belongs to the largest step, which is two parts becoming one.

What does not survive is a reading of the sixth rung as being about gestures. The three textures that score are not scoring because they are shaped differently over six bars; they are scoring because their last bar differs from the bar before it. The six bars of thinning that precede the loud final chord contribute nothing measurable to the number beyond setting what the final chord is measured against — which is the whole of what the bar-count table above is saying.

So the ladder’s dynamic reading has one degree of freedom, and it is the last chord.

Three outcomes a sweep can have, and this rung got two

This is the third parameter sweep this ladder has run and the fourth in the collection, and the pattern across them is worth stating because it is becoming a method.

A sweep takes a number a model asserts and asks how much of the conclusion depends on it. There are three outcomes and all three are useful.

The parameter decides the answer. Then the model is really a model of that parameter, and the honest thing is to say so. The number every ladder here has been quoting is the worked example: sweeping one assumed width of a perceptual category inverted the collection’s most-quoted result about hearing.

The parameter does not decide the answer. Then the conclusion is stronger than it looked, because it survives not knowing the number. That is the seventh rung’s tempo and it is the profile exponent on the harmony ladder.

The parameter does not decide the answer and a parameter nobody was asking about does. Then the model is not the model anybody thought it was, and the right response is to find the one sentence that predicts both — which is what the step account is.

The third case is the new one and it is the reason this rung exists. A null on its own is a robustness claim and a positive on its own is a finding; a null and a positive that the same sentence predicts is a mechanism, because the sentence had to get both right and could have failed twice. The ritardando had to come out flat and the bar count had to come out in steps, and a shape-based reading would have produced the opposite pair.

Three ways to arrive at the same final tempoTempo against position in the closing passage, ending at 35 per cent of the opening tempo, for curvature exponents 1, 2, 3. All three begin and end at the same tempo, so what separates them is the middle: at the halfway point they read 68 per cent for linear in score position, 75 per cent for constant deceleration, 80 per cent for q = 3. The straight line is the one nobody plays. Measured ritardandos fit the decelerating curves, which is the whole of Kronman and Sundberg's argument: a closing gesture has the shape of a body stopping rather than of a dial being turned, and the parameter that varies between performances is the final tempo rather than the shape.the final tempo — 35%linear in score positionhalfway: 68%constant decelerationhalfway: 75%q = 3halfway: 80%all three endat the same tempo0.000.200.400.600.801.0000.20.40.60.81position in the closing passagetempo, as a fraction of the opening
Fig. 4 The deceleration curve this essay swept, established earlier. Its one free parameter is the final tempo, and moving it across a factor of three moves the closing bars from twelve seconds to twenty and the reading by nothing at all — the same five figures at every setting.

What would race the release

The construction says exactly what a piece would have to do for the running impression to matter, and it is a strange instruction.

The smoother’s release is two seconds. For the impression to be caught out, the thing being asked about has to happen inside two seconds of the arrival — which at any tempo anybody plays is a bar or less.

A ritardando does the opposite. It makes the closing bars longer, which gives the smoother more time to follow, which is the direction of no effect rather than of a large one. The intuition that stretching the gesture would help was backwards.

What would help is a gesture compressed into the last bar: a subito change, an unprepared arrival, an accent after a silence. Those are the closing devices that a running impression cannot follow, and they are the ones this reading is sensitive to.

That is a small and real result and it points at something outside this ladder. A silence long enough to be an ending is about a gap of three and a half seconds; a gap of that length lets the impression fall a long way, so the chord after a long silence arrives against a much lower reference. The interaction between a silence and the running impression is a term this ladder has never computed and it is the one place the two-second release should be visible.

Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness.
Fig. 5 The earlier figure: four closing textures with the running impression under each. The three with a step at the arrival produce a reading and the one without a step does not, which this essay explains rather than repeats. The biggest of the three steps is downward.

The reading is not wrong, it is narrow

It would be easy to read this as demolishing the sixth rung and that is not what the arithmetic says.

A quantity with one degree of freedom is a quantity. What the sixth rung measures is how large the step at the arrival is, in the units of a listener’s running impression rather than in decibels — and that is a genuinely useful conversion, because a step of six decibels arriving after a long crescendo and the same step arriving after a diminuendo are different perceptual events and the raw decibels do not distinguish them.

What the reading cannot do is compare gestures. Two endings that arrive at the same final chord from different shapes over six bars get the same number, and the ladder’s sixth rung presented four textures as though the reading were about the shape rather than about the last bar.

So the repair is a restatement rather than a retraction. The dynamic component of an ending is the size of its final step against the listener’s running level, and the six bars before it enter only through where they leave that level. That is a smaller claim and it is one the figure fully supports.

It also makes the component commensurable with the others in a way it was not. What makes an ending an ending lists five components and every one of the others is an event at the arrival — a cadence, a long final note, a return to the tonic. The dynamic component now joins them as an event rather than sitting apart as a shape.

Five signals, computed separately, and no total. The five components of closure for 6 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.
Fig. 6 The five components of an ending, of which the dynamic one is what this essay has been auditing. After the audit it is the same kind of thing as the other four: something that happens at the arrival.

Which computation produced the numbers

The level track is the sixth rung’s scoreDynamics: a thirty-two-bar scheme with a texture instruction over the closing bars, each bar’s loudness computed as a mean over the voicings its part count admits, converted to a level and run through the loudness ladder’s two smoothers.

The ritardando warps the bar durations over the closing bars using the deceleration curve the third rung fitted, with its one free parameter — the final tempo as a fraction of the starting one — swept from one to a third.

The arrival reading is the short-term loudness over the long-term loudness at a stated delay past the start of the final bar. The reading at a delay of 0.3 seconds and a ritardando of one reproduces the sixth rung’s own number exactly, which is the check that this is the same quantity.

The gesture sweep runs four bar counts against three bar lengths and reports twelve runs, so a plot of it against seconds has two of its points at nearly the same abscissa from different bar counts. That is why the scatter reads as scatter: the figure’s horizontal axis is a product of the two dials and only one of the two is in the answer.

Where the model stops

The scheme is one form. A thirty-two-bar song with a stated harmonic plan is a specific and unusual object, and a symphonic ending is nothing like it. What is being audited is the reading rather than the repertoire, so this matters less than it would for a claim about endings.

The texture is a part count. The dynamic of a bar is the loudness of a chord voiced for that many parts, which is the fourth rung of the loudness ladder’s model and which has no written dynamic marks in it at all. A real ending has both.

The smoother’s constants are from one standard. The attack and release times are those of a published loudness meter, and the release is the number everything here turns on. A different standard’s release would move the exponential’s time constant and change nothing else.

And a ritardando is not only a lengthening. Real players slow down and get quieter together, and the two are correlated in performance data; modelling the tempo alone holds a correlation at zero that is not zero.

What the picture cannot show

It cannot show a listener. The running impression is a model of a meter, not of a person, and whether the ratio it computes corresponds to anything a listener has is the assumption the whole of the sixth rung rests on.

Nor can it show the other components. What makes an ending an ending lists five and this is one of them; the harmonic, melodic, rhythmic and durational components are all elsewhere in the ladder and none of them is in this number.

And it cannot show whether a null means the model is wrong. An instrument that returns the same number under a manipulation is either measuring one robust thing or measuring nothing, and no amount of staring at the null decides which. What decides it here is that the same account which predicts the null also predicts a specific non-null of a specific shape, and the bar-count table is that prediction coming back right. A reading that had been empty would have been flat under both.

Whose endings, and when

The thirty-two-bar song form is a twentieth-century popular one and it is used throughout this ladder because it is short, standard and has a stated harmonic plan. The deceleration curve is fitted to measurements of Western classical performance from the twentieth century, where a ritardando into a final cadence is close to obligatory.

The devices this rung says would register are also datable. A subito piano before a final chord is a classical and romantic gesture; the unprepared loud ending after a general pause is a nineteenth-century one; the fade-out is a recording-studio device of the second half of the twentieth century and is the one case where the gesture is deliberately made much longer than any release, so that the impression follows it perfectly and the ending is heard as a departure rather than as an event.

That last case is the cleanest confirmation available. A fade is engineered to be unracing, and it is heard as no ending at all.

Where this ladder goes next

Eight rungs. An ending is five components with no total; half the cadences withhold some of them; the performance slows on a curve; nothing in the statistics announces a stop; a silence is an ending after three and a half seconds; the dynamic component is asymmetric between adding parts and taking them away; the tempo did not decide it; and now neither did the deceleration, while the number of closing bars did — because the reading is a step response, its only variable is the last step, and a bar count is the only thing swept that changes which step that is.

What is owed after this is the silence. The one manipulation that should move a running impression is a gap, because two seconds of nothing is one full release and the chord after it arrives against a reference that has fallen. The fifth rung has a silence with a measured threshold in it and this rung has the smoother; putting them together would say what a general pause is worth in the same units as a diminuendo, and it is the one place in this ladder where the two-second constant should finally do some work.

Part 8 of 13

One essay in the series on closure. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CadenceClosureDynamicsFinal lengtheningLoudnessRitardandoTempoTemporal integration