The bars a key is made of
Assumes: The listener who forgets · The margin the dynamic program already had
The listener who forgets put a memory on this ladder’s key-finder and found that forgetting costs a great deal of confidence and very little accuracy. It ended by saying what was wrong with the memory it had used:
A discount that treats every bar alike is the simplest memory a model can have and it is the wrong shape: a listener remembers the first chord of a phrase and the cadence at the end of it far better than the middle.
That is a serial-position effect, and it has a large literature. The ingredients for a memory that has one are already here — the collection knows which bars carry a cadence, which open a section and which close one — so the objection can be answered rather than repeated.
It turns out to have two answers and they point opposite ways.
Take a bar away and see what it was worth
Before asking what shape a memory should have, there is a prior question nobody on this ladder has asked: which bars is the reading actually built from?
It has a direct answer. The key-finder assigns every bar an emission score — how well its notes fit each degree of each key — and that score can be switched off for one bar without touching anything else. The transitions still run through it; the bar simply stops being evidence. Do that thirty-two times and the passage says which of its bars were load-bearing.
Two things can be measured when a bar goes, and the whole of what follows is that they are different.
The certainty is the mean margin between the best two readings, which is the quantity the margin the dynamic program already had introduced. Removing a bar lowers it, because a reading with less evidence behind it is less far ahead of its rivals.
The answer is the sequence of key names. Removing a bar may leave every name where it was, or it may move some of them.
A bar can supply a great deal of the first and none of the second, and most of them do.
On the thirty-two-bar form the eighteen bars an analysis would point at — a section opening, a section closing, a dominant, the chord a dominant resolves to — average 0.079 bits of certainty and half a bar of answer each. The fourteen ordinary bars average 0.047 bits and 0.07 bars.
So the structural bars carry 1.7 times as much certainty and seven times as much answer. On the first quantity the difference is real and modest. On the second it is nearly categorical: take away an ordinary bar and the reading almost never changes at all.
Four bars out of forty
The per-bar picture is more extreme than the averages make it sound, and the extreme is the readable part.
On the thirty-two-bar form, five of the thirty-two bars move the reading at all and four of the five are structural. One of them — bar nineteen, the applied dominant in the middle of the bridge — moves six bars on its own, which is more than the other four together. The remaining twenty-seven bars can each be deleted with no consequence whatever for what the model says the passage is in.
That is the shape of an over-determined inference and it is what the listener who forgets found from the other direction. A key is decided early and cheaply; almost everything after the deciding is confirmation. The bars that move the answer are the bars where the deciding is still going on, and in a form with one modulation there are four or five of them.
And three bars have a negative influence. Removing bar twenty raises the mean margin by 0.18 bits; removing bar twenty-one raises it by 0.16, and bar eighteen by 0.13. All three are applied dominants inside the bridge, and all three are bars whose notes fit the winning reading badly.
There is nothing wrong with a negative number here. A margin is the gap between the best reading and its nearest rival, and a bar that fits neither well drags the winner down toward the field. Deleting it is a small improvement to the model’s confidence and a small loss to its honesty.
Which makes the bridge the interesting object. It is the only stretch of the form that is genuinely ambiguous, it holds every bar that moves the answer, and it holds every bar whose removal makes the model surer. A memory built to keep the structural bars keeps the bridge; a memory built to keep the recent ones does not, unless the listener happens to be inside it.
The rondo says the same thing more sharply
The thirty-two-bar form is an awkward test because it barely modulates. A form with a stated key plan gives the measurement more to move.
On the rondo the split is cleaner and the first half of it reverses. The twenty-seven structural bars average 0.085 bits of certainty against the thirteen ordinary bars’ 0.097 — the ordinary bars are worth more. And the ordinary bars move the reading in not one case out of thirteen, against 0.19 bars for the structural ones.
Four bars of the forty move the reading, and they are bars eight, nine, ten and seventeen: the last bar of the opening refrain, the first bar of the episode, the dominant seventh inside it, and the first bar of the refrain’s return. They are the two seams of the one modulation the model finds, and nothing else in the form is doing any work on the question of what key it is in.
Nothing else in the form is even close. The bar with the largest effect on the certainty is bar thirty, an ordinary subdominant in the middle of the minor episode, and removing it moves no reading anywhere.
That is the result, and it is worth stating as a sentence rather than as a table: the structural bars decide which key it is and the ordinary bars decide how sure the model is.
Once said it is not mysterious. A margin is a sum over bars, and a bar of plain diatonic harmony fits its key very well and its rivals badly, so it adds a lot to the sum. A dominant seventh fits fewer keys but the ones it fits it pins down, because it contains a tritone and a tritone belongs to two keys rather than seven. The ordinary bars are the mass of agreement; the structural bars are where the disagreements are settled.
Which means an analysis and a confidence are about different bars, and the ladder has been quoting one number for both since the key plan.
What that predicts about a structured memory
The obvious inference is the one the debt made: if the structural bars carry the answer, a listener who remembers them and forgets the rest should read the passage nearly as well as a listener who remembers everything.
That is testable and it is wrong.
The test has to hold something fixed or it is not a test. What is held fixed is the budget: the total attention a listener has to spend across the passage. Five memories spend the same total and spend it in different shapes — evenly; with a decay toward the recent bars, which is the shape the twelfth rung used; concentrated on the structural bars; concentrated on the bars that are both structural and recent, which is the object the debt actually asked for; and concentrated on everything an analysis ignores, as a control.
At half the full attention, spreading it evenly names the same key as an unforgetting reading in ninety-seven per cent of bars. Concentrating it on the structural bars gives ninety-one. Concentrating it on the structural and recent bars gives nothing at all: the reading agrees in no bar of the passage.
The control is worse still — the memory built of the bars an analysis ignores agrees in three per cent — so it is not that structure is irrelevant. It is that a memory shaped like an analysis, at the same cost, is a worse instrument than a memory shaped like nothing.
Why concentration is a bad idea and not a bad shape
The failure is not that the structural bars are the wrong bars. It is that attention with a budget behaves multiplicatively, and a memory concentrated on a sixth of a passage does not read that sixth well — it reads it too well.
An emission score scaled up counts a bar’s evidence more than once. A memory that puts six bars’ worth of attention onto three bars is not a listener who noticed those bars; it is a listener who heard them three times. The reading it produces is the reading of a passage that is not the passage.
That is a modelling artefact in one sense and a real statement in another. A budget on attention is the right constraint — a listener cannot attend to everything and the whole point of a memory is that something is lost — and given that constraint, spreading it is what preserves the reading.
The two halves of the finding therefore do not contradict each other and are easy to read as though they did. The structural bars are where the reading is decided. A memory made only of them is not the reading, because a key is established by the mass of ordinary bars and only changed at the structural ones, and a listener who has kept none of the mass has nothing for the changes to be changes to.
The comparison is not stable, and that is a result
Halve the budget again and the ordering changes. At a quarter of full attention the recency memory reads the thirty-two-bar form best, the structural memory beats the even one, and the even one has fallen to fifty-nine per cent.
Run it on the rondo at a third of full attention and the even memory is back on top at seventy-eight per cent, the structural one at twenty and the structure-and-recency one at forty-five.
Nothing here is monotone in the budget and nothing is consistent across the two forms except one thing: the even memory is never the worst, and every shaped memory is the worst somewhere.
That is a weaker claim than a ranking and it is the claim the numbers support. A ranking would have been more satisfying and would not have survived the second form.
The instability itself is worth reading rather than apologising for. What makes it jump is that a shaped memory changes which bars are ambiguous, and the reading at an ambiguous bar can flip on very little. A memory that has kept bars sixteen to nineteen of the thirty-two-bar form is reading a bridge; a memory that has kept bars one to eight is reading a tonic. The two produce different sequences of names, and how many names they share is not a smooth function of how much either one remembers.
So the honest report is a floor rather than a ranking: whatever shape a memory has, it agrees with an unforgetting reading somewhere between nothing and ninety-seven per cent of the time, and the shapes that never fall to nothing are the ones that keep something everywhere. That is a real constraint on any account of how a listener holds a key, and it is much weaker than the account the debt proposed.
What the ladder should take from this
Three of this anchor’s rungs now report a quantity that depends on an assumption about memory, and the assumptions are not interchangeable.
How much of the reading arrives late prices hindsight, which is about the next few bars. The listener who forgets prices recency, which is about the last few dozen. This essay prices shape, which is about neither: it is about whether a listener’s attention is distributed like an analysis, and the answer is that if it were, the readings would be worse.
There is a fourth assumption underneath all three and it has never been examined either. Every bar of every scheme here is one chord of equal length, so the model’s attention and the music’s own weighting are the same thing by construction. A real passage gives a cadence four beats and a passing chord one, and the metre ladder’s account of weight says that difference is not decoration.
What the measurement cannot say
It cannot say what a listener attends to. The budget is a modelling device with no measured value, and the whole comparison is between shapes at an arbitrary size. Nothing in this collection gives a number for how much of a thirty-two-bar form a listener retains, and the repetition ladder’s account of recognising a return is the nearest thing to one.
It cannot say that the roles are the right roles. A section opening, a section closing, a dominant and its resolution are four categories read off a scheme rather than off a hearing. A theorist asked to mark the structural bars of a rondo would mark others, and might not mark all of these.
It cannot separate a bar’s role from its content. The dominant bars are also the bars containing a tritone, and the measurement cannot say whether they carry the answer because they are cadential or because a diminished fifth belongs to fewer keys than a perfect one. The chord that is not played at once makes the same distinction about arpeggiation and does not resolve it either.
And it cannot say anything about a first hearing. Every reading here is of a scheme the model has seen once. The form a first hearing cannot have is the collection’s account of what changes on the second, and a memory that keeps cadences is exactly the memory that would change most.
Which computation produced the numbers
The key-finder is unchanged from a key-finder that keeps the order: states are a key and a scale degree, the emission is how well a bar’s pitch classes match that degree’s triad, the transitions are root-motion frequencies, and changing key costs a fixed amount. It is held at 0.8 here rather than the 2.2 used elsewhere on the ladder, because at 2.2 neither form modulates at all and a measurement of what moves the reading has nothing to move.
The margin is the gap in bits between the best path through the best key at a bar and the best path through any other, which is the tenth rung’s quantity.
Attention scales a bar’s emission before the dynamic program runs. At one the bar counts fully; at zero its emission is flat across every state and it contributes only its transition. Every policy is normalised so the attention over the passage sums to the same total, and each figure asserts that it does.
The general dynamic program is checked against the ladder’s own at every placement: at one collection and full attention it reproduces the published margins to the last bit, which is what makes this a measurement of the same model rather than of a new one.
Where this ladder goes next
Thirteen rungs in, the ladder has a key-finder, a map of key distance under three metrics, a count of how much evidence a modulation needs, an account of what happens when two keys sound at once and when they alternate, an ordered model beside a histogram one, a tuned key cost, a margin at every bar, a split between what a listener could have and what an analyst has, a memory, and now a measurement of which bars the answer is in.
What is owed after this is the weighting inside the bar. Every figure on this ladder gives each bar one chord and one unit of evidence, and a bar is not one thing: the harmony changes within it, the chord tones sit on some beats and not others, and which notes are the chord is a decision the metre makes before the key-finder ever sees a pitch class. A key reading built on a segmentation that is itself uncertain is carrying an error bar nobody has drawn, and the two decisions are already computed side by side in three decisions that constrain each other. Joining them needs no corpus and no listener — it needs the segmentation’s own posterior fed into the key-finder’s emission instead of its best guess, which is arithmetic this collection has both halves of.
Part 13 of 21
One essay in the series on Key-relations. The essays either side of this one:
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
CadenceHidden markov modelInferenceKey-findingMemory decayMetrical weightModulationMusical form
- The cadence as evidence cadence, inference, key-finding
- The passage built to make them disagree cadence, key-finding, modulation
- A bass line is not a list of roots inference, key-finding
- A chord, given a key and a predecessor cadence, key-finding
- A count and a correlation cadence, key-finding
- A count is not an estimate memory decay, musical form