Harmony and voice leading

The listener who forgets

Setting an analyst's reading of a key against a listener's measures what arrives late. Both passes assume perfect recall of their own half — which is as wrong going forward as knowing the future is going back. Put a decay on the forward pass and a listener with a memory of a few bars keeps a quarter of the confidence and nine tenths of the answers.

Assumes: How much of the reading arrives late · The margin the dynamic program already had

How much of the reading arrives late split the ladder’s key-finder in two. The two-sided pass uses every bar of a passage, including the ones after the bar being read; the causal pass drops the second half and is what a listener could have at the time. The difference is the retrospective element in tonal hearing, and that rung measured it.

Its last paragraph names the other half of the same objection:

The causal pass has perfect recall of everything before the bar, which is as wrong as the two-sided pass’s knowledge of everything after — so the honest listener’s reading is a forward pass with a decay.

A decay is one parameter and one line: discount the accumulated score once per bar, so a chord eight bars ago counts for less than the one just heard.

A listener has a quarter of an analyst's confidence and the same answer. The mean margin between a passage's best two key readings, against how many bars a listener's memory of the evidence takes to halve. The two-sided reading — the one that uses bars that have not happened yet — sits at 3.90; the forward pass with perfect recall at 1.79; a forward pass whose evidence halves every 6.6 bars at 1.00. The dots' size is how often that reading names the same key as the two-sided one: 100 per cent at perfect recall and 81 at a one-bar half-life. So forgetting costs a great deal of confidence and very little accuracy — the key is robust and the certainty is not.
Fig. 1 The mean margin between a passage’s best two key readings, against how many bars a listener’s memory of the evidence takes to halve. The dots’ size is how often that reading agrees with the two-sided one about the key.

Three readings of one passage

Run the thirty-two-bar form the ladder has used since its second rung and the mean margin between the best two readings comes out:

mean margin agrees with the analyst
two-sided, with hindsight 3.90
forward, perfect recall 1.79 100%
forward, halving every 6.6 bars 1.00 91%
forward, halving every bar 0.23 81%

The eleventh rung’s finding was the first gap: fifty-four per cent of a two-sided reading’s confidence comes from bars that have not happened. This rung adds the second: a listener whose memory of the evidence halves over a few bars keeps about a quarter of the analyst’s confidence, and about half of the perfect-recall listener’s.

And the answers barely move. At perfect recall the forward pass names the same key as the two-sided one in every bar. At a half-life of six and a half bars it does in ninety-one per cent of them. At a one-bar half-life — a listener who has essentially forgotten the bar before last — it still agrees eighty-one per cent of the time.

That is the shape of the result: forgetting costs a great deal of confidence and very little accuracy.

Which is a reasonable thing for a key to be

It is worth saying why that is not a surprise once it is stated, because a result that only surprises is usually a bug.

A key is over-determined by its evidence. Every bar of a tonal passage carries a triad drawn from seven degrees, and the joint probability of a run of them under the wrong key falls fast — so by two or three bars in, the correct key is already far ahead, and the remaining twenty-nine bars are adding to a lead rather than deciding a contest.

A decay removes most of the lead and none of the decision. The margin is a sum over bars and shrinks with the discount; the ranking is decided early, and a decay that keeps the last two or three bars keeps the part that did the deciding.

So the model says a key is easy and a listener’s certainty about it is expensive, which is a distinction the ladder has not had and which its earlier rungs need. The margin the dynamic program already had turned every figure in this anchor from an analysis into a measurement of ambiguity, and this rung says that measurement is in units nobody has: the number depends entirely on how long the listener is assumed to remember.

And where a decay ought to change an answer

There is one quantity on this ladder a decay should move rather than merely shrink, and it is worth naming because this rung does not compute it.

How much evidence a modulation needs counts the chords it takes for the model to switch keys, and it counts them against a forward pass that remembers the old key perfectly. A listener who is forgetting the old key should switch sooner, because the evidence for the previous reading is decaying while the evidence for the new one is fresh.

That is a shift in a number the ladder reports rather than a shrinkage of a confidence, and it is a testable one: the modulation length should fall with the decay rate, and it should fall fastest for modulations that arrive after a long stretch in one key.

The same argument runs the other way for two keys in alternation, where the model’s difficulty is that it keeps being dragged back to a key it has just left. A memory that fades should make alternation easier to follow, not harder — which would be the first time on this ladder that forgetting improved a reading.

Where the confidence goes

Bar by bar the loss is not spread evenly, and where it falls is the interesting half.

Where the confidence goes when the memory fades. The margin at every bar of a thirty-two-bar form: the two-sided reading, and three forward readings whose evidence halves after infinity, 6.6 and 1.9 bars. The two-sided curve is highest everywhere by construction. What the decay takes is not spread evenly: the bars that lose most are the ones deep inside a section, where the perfect-recall pass has accumulated a long unbroken run of evidence, and the bars that lose least are the ones just after a cadence, where the evidence is fresh anyway. A listener with a short memory therefore hears a form whose certainty resets at every phrase, which is closer to what listening is like than a curve that climbs for thirty-two bars.
Fig. 2 The margin at every bar of the form, for the two-sided reading and three forward ones. The decay flattens the long climbs and leaves the fresh bars alone.

The bars that lose most are deep inside a section, where the perfect-recall pass has accumulated an unbroken run of consistent evidence and its margin has been climbing for eight or twelve bars. The bars that lose least are the ones just after a cadence or a change of key, where the evidence in the last two bars is most of what there is anyway.

So a listener with a short memory hears a form whose certainty resets at every phrase, rather than one whose certainty climbs monotonically from the first bar to the last. That is a much better description of what listening to a thirty-two-bar song is like, and it is what the decay buys.

It also predicts where the two readings differ. A bar where the perfect-recall pass is confident and the decaying one is not is a bar whose key is established by material the listener has stopped counting — which is exactly the kind of bar an analyst points at and a listener does not notice.

What a plausible half-life is, and the ladder has one

The decay rate is a free parameter and this rung does not measure it. What it can do is say where the collection’s own numbers put it.

The repetition ladder’s account of how a return is recognised has an account of how a musical memory fades, and the melody ladder has the horizon over which a listener holds a sequence of pitches. Both are of the order of a few seconds to a few tens of seconds, which at two seconds a bar is a half-life somewhere between two and ten bars.

Over that whole range the margin runs from 1.32 down to 0.61 and the agreement from ninety-four per cent to eighty-four, which is the band the key plan should be read inside. The conclusion does not depend on which end of the range is right, which is the useful thing about a parameter whose effect is monotone and shallow.

What would change the conclusion is a half-life under a bar, and nothing in this collection suggests one.

How much of the reading comes from what has not happened yet. Every margin reported earlier is two-sided: the best path through a key at a bar is the best score into it plus the best score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line, because the dynamic program already had both halves separately. The mean margin falls from 3.90 bits with hindsight to 1.79 without it, so 54 per cent of this passage's certainty is retrospective. The two passes never disagree about which key is best here, so the hindsight buys confidence rather than a different answer. This is the quantity every earlier essay has assumed and none has measured.
Fig. 3 The earlier comparison, which is the top two lines of the figure above. This essay is what happens when the second of them is also made honest.
How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is.
Fig. 4 How many bars a key change takes to be heard, computed earlier. Every number on it was computed with perfect recall, and a decay would lengthen it — which is the one place on this page where forgetting should change an answer and not only a confidence.

What the two corrections say together

The eleventh rung and this one are corrections in opposite directions to the same object, and putting them side by side gives the anchor something it has not had: a bracket.

An analyst’s reading is the two-sided one, at 3.90. It uses the future and has perfect recall, and is the upper bound on what anybody could extract from the notes.

A listener’s reading is the forward one with a memory, at about 1.00 for a plausible half-life. It uses neither. The middle line — the forward pass with perfect recall, which is what the eleventh rung called the listener’s reading — turns out to be a third object that nobody is: it knows no future and forgets nothing.

The ratio is a factor of four, and everything a published analysis says about how strongly a passage is in a key is being said at the top of that bracket. That is not a criticism — an analysis is meant to be an account of the notes rather than of a hearing — but the anchor has been treating the two as the same reading with different amounts of information, and they are not: they differ by more than either differs from naming the wrong key.

The ambiguous passage, where it should matter most

The eleventh rung built a passage designed to make the two passes disagree — a phrase that sits in one key and ends with a cadence in another — and used it to isolate the retrospective element.

That passage is short, which is exactly the case a decay cannot touch. With eight bars and a half-life of six, the discount on the oldest evidence is under a factor of two, and the decaying reading is nearly the perfect-recall one.

A decay is a long-form effect. It does nothing to a phrase and a great deal to a movement, which is the opposite of the retrospective effect the eleventh rung measured: hindsight matters most over a few bars, because that is how far ahead a cadence is.

So the two corrections to the key-finder act on different time scales and do not compete. One is about the next few bars and one is about the last few dozen.

How much of the reading comes from what has not happened yet. Every margin reported earlier is two-sided: the best path through a key at a bar is the best score into it plus the best score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line, because the dynamic program already had both halves separately. The mean margin falls from 0.69 bits with hindsight to 0.18 without it, so 74 per cent of this passage's certainty is retrospective. And at 9 of the 11 bars the two passes name different keys — those are the bars a listener hears one way and an analyst reads another, and bar 1 is the first, C at the time and G afterwards. This is the quantity every earlier essay has assumed and none has measured.
Fig. 5 The constructed passage from that earlier essay, where hindsight decides the reading. A memory decay has almost no effect on it, because there is not enough of it to forget.
How sure the reading is, bar by bar. Every earlier essay reports one best reading. A dynamic program that finds a best path has, by construction, the best score into every state at every bar — so the gap between the best reading and the best reading in any other key is already computed and has never been printed. Here it is, in bits, for the thirty-two-bar AABA. The mean margin is 1.14 bits and 13 of 32 bars are inside one bit of a rival reading, which is where a listener would be genuinely undecided. The reading itself names C, E, B, D, G; the margin says what that naming is worth, and at the weakest bar — bar 31, C over F — it is worth 0.07.
Fig. 6 The margin itself, computed earlier, which is the quantity every number on this page is a version of. It is a difference of model scores and its size has no units a listener would recognise — which is why the ratios above are the reportable part.

What this does to the ladder’s own instrument

The tenth rung’s contribution was that a dynamic program which finds a best path has, for free, the score of every other path — so the margin between the top two readings is available at every bar and turns an analysis into a measurement.

Two rungs later that measurement has acquired two knobs. Whether the reading is two-sided or causal changes the mean margin by a factor of 2.2; how long the listener remembers changes it by another factor of up to 8. The margin is not a property of a passage; it is a property of a passage and a listener.

That is worth stating plainly because the ladder has been quoting margins for three rungs as though they described the music. They describe the music under an assumption about recall, and the assumption every rung before the eleventh made — perfect memory of everything, in both directions — is the one furthest from any listener.

The useful form of the measurement is therefore a ratio rather than a number: how much more ambiguous one bar is than another, under the same assumptions. That survives every setting of both knobs, and the figure above shows why — the curves move up and down together and keep their shape.

The one number the ordered key-finder was tuned on. For each rate of alternation between two keys, the cost of changing key at which the model stops hearing two keys and starts hearing borrowed chords in one. The threshold rises with the period — 1.70 at 8 bars — so the parameter and the rate trade off against each other exactly. The value tuned earlier, 2.2, sits above every threshold on this axis, which means its verdict about fast alternation was a consequence of the tuning rather than a finding about the music. Filled means the model names two keys; hollow means it names one and calls the rest borrowings.
Fig. 7 The other free parameter carried throughout: how much the model charges for changing key. Everything on this page holds it at 2.2, and the decay and the key cost pull on the same quantity from opposite ends.

Two free parameters, and they are not the same one

It would be reasonable to suspect that a memory decay is the key cost in disguise. Both make the model less willing to hold a reading over a long stretch, and both have a single number in them.

They are not the same. The key cost is charged when the reading changes, and it makes modulations expensive without touching how confident the model is in between. The decay is charged every bar, and it makes long stretches of agreement worth less without making a change any cheaper.

The difference shows in the shape rather than in the mean. Raising the key cost lifts the margin everywhere and flattens the dips at the modulations; raising the decay leaves the dips where they are and pulls down the peaks. On the bar-by-bar figure they are visibly different operations, and the ladder now has both.

That leaves the anchor with two parameters and one measurement, which is honest and is not comfortable. The key cost was swept rather than chosen for the same reason three rungs ago, and the decay will have to be too until somebody has a number for how long a listener holds a bar.

Which computation produced the numbers

The key-finder is this ladder’s own hidden Markov model, unchanged: twelve keys times seven degrees as states, an emission score from how well each bar’s pitch classes match each degree’s triad, a transition weight from the root-motion frequencies, and a fixed cost for changing key.

The margin is the tenth rung’s: for each bar, the best path through each key at that bar, and the gap between the top two.

The causal pass drops the backward half. The decay multiplies the accumulated forward score by a constant once per bar before the new bar’s evidence is added, which is an exponential discount on everything older.

The half-life is that constant read as a number of bars: the log of a half over the log of the decay.

The passage is the thirty-two-bar song form with its stated key plan, at one chord a bar.

Where the model stops

A discount on a log score is not a model of memory. It is the simplest thing that has the right shape, and a real memory does not decay exponentially — it is better for the beginning of a passage than for the middle, which is a serial-position effect no exponential has.

The decay applies to everything equally. A real listener remembers a cadence better than a passing chord, and this discount does not know the difference. Weighting the discount by metrical or structural importance is the obvious next thing and it is not one parameter any more.

The margin is in log-probability. It is a difference of model scores, not a psychological quantity, and calling it confidence is a convenience the tenth rung already flagged.

And the transition weights are a corpus somebody else’s. The root-motion frequencies come from published counts and this collection has recorded three ladders wanting the same corpus to do better.

And the key plan is a compromise. The root-motion probabilities come from common-practice European music and the thirty-two-bar form is a twentieth-century popular one, which is this ladder’s standing mismatch and is stated at every rung that uses both.

What the picture cannot show

It cannot show two keys at once. A passage that is in two keys is the case where the margin is the whole subject, and a decay changes it by an amount this essay has not computed on that passage.

It cannot show a listener who is not looking for a key. The whole apparatus assumes the question is which key, and most listening is not that.

It cannot show the two knobs interacting. The key cost and the decay are swept one at a time here, and they act on the same accumulated score; a joint sweep would say whether a high decay and a low key cost are indistinguishable, which is the question a fitted model would face first.

Nor can it show the melody. Every bar here is a triad, and a listener’s evidence includes the tune, the bass line and the instrumentation.

It cannot show what confidence is for. A margin of one against a margin of four is a difference in the model’s certainty, and whether that corresponds to anything a listener could report is untested.

It cannot show a modulation being noticed. A key change is the event a listener does report, and how much evidence one needs is this ladder’s account of it — computed, like everything before this rung, with perfect recall.

And it cannot show a second hearing. A return has to be remembered, and a listener who knows the piece has a memory of the whole thing — which is the two-sided reading, arrived at by a completely different route.

Whose music, and when

The form is a twentieth-century popular one and the transition model is common-practice European. Neither is a claim about a repertoire.

The observation with a period in it is about length. A decay that halves in a few bars makes very little difference to a thirty-two-bar song and a great deal to a sonata movement, where the perfect-recall pass accumulates evidence for hundreds of bars and a listener does not. So the gap between an analyst’s reading and a listener’s should grow with the length of the form, and it should have grown historically as forms did.

That is a prediction the model makes and cannot test here, because it needs the analysed corpus three ladders on this site have now recorded wanting.

Where this ladder goes next

Twelve rungs. Keys are neighbours and the map is computed; the key plan is the form; three distance measures disagree; a modulation takes a measurable number of chords; the circle is a circle and the map is not; two keys at once; two keys in alternation; a model that keeps the order; the number that decides which reading it gives; what the reading is worth; how much of it a listener could have had at the time; and now how much of that survives forgetting.

What is owed after this is the weighting. A discount that treats every bar alike is the simplest memory a model can have and it is the wrong shape: a listener remembers the first chord of a phrase and the cadence at the end of it far better than the middle, which is a serial-position effect with a large literature. This ladder already computes which bars are metrically strong and which carry a cadence, so the ingredients for a structured memory are here — and the question it would answer is whether a listener’s key reading is built from the bars a theorist would call structural, which is an assumption every analysis makes and none of them tests.

Part 12 of 21

One essay in the series on Key-relations. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Hidden markov modelInferenceKey-findingKey-relationsMemory decayModulationMusical form