Scales and modes

A modulation and a borrowing are one number apart

The key-finder that keeps the order has one tuned parameter, and said so. Sweep it and the parameter turns out to be the whole verdict: below a threshold the model hears two keys alternating, above it one key with borrowed chords. The threshold rises with how slowly the keys alternate — and the value it chose sits above every threshold in range, so its finding about fast alternation was a consequence of the tuning.

Assumes: A key-finder that keeps the order · The alternation a key-finder cannot follow

The eighth rung of this ladder built a key-finder with a state — eighty-four states, one per key and degree, and 7,056 transitions between them — so that it could tell an alternation between two keys from a simultaneity of both, which a bag-of-notes model cannot. It has one free parameter: the cost of changing key, which stops the model from re-analysing every chord into whichever key fits it best.

That rung set the cost to 2.2, said out loud that the value was tuned to produce sensible answers, and named the consequence in its last paragraph:

Both models are wrong about fast alternation and one of them is wrong in an interesting way: the ordered model reads the second key as borrowed chords in the first, which is what a tonal analyst would write. So the two readings — a modulation and a set of borrowings — are the same evidence under two different priors on how often keys change, and the parameter that separates them is the one number this rung tuned.

The one number the ordered key-finder was tuned on. For each rate of alternation between two keys, the cost of changing key at which the model stops hearing two keys and starts hearing borrowed chords in one. The threshold rises with the period — 0.95 at 1 bar, 0.95 at 2 bars, 0.95 at 4 bars, 2.00 at 8 bars, 3.50 at 16 bars — so the parameter and the rate trade off against each other exactly. The value tuned earlier, 2.2, sits above every threshold on this axis, which means its verdict about fast alternation was a consequence of the tuning rather than a finding about the music. Filled means the model names two keys; hollow means it names one and calls the rest borrowings.
Fig. 1 The tuned number, swept. For each rate of alternation between two keys, the cost at which the model stops hearing two keys and starts hearing one with borrowings. The threshold is 0.95 for alternations of one, two and four bars, 2.0 at eight and 3.5 at sixteen — so the value chosen earlier, 2.2, sits above every threshold up to eight bars and below the one at sixteen.

What the parameter actually is

The cost is a prior. It is the log-odds penalty the model pays for saying that the key has changed between one bar and the next, and it does the same job a Bayesian prior does in any inference: it decides how much evidence is needed before a hypothesis is preferred to the default.

Set it to zero and the model has no preference for staying, so it re-analyses every bar into whichever key fits that bar best — which produces a reading in which the key changes constantly and means nothing — and which is exactly what the circle-of-fifths map would predict, since every chord is near some key. Set it very high and the model never changes key at all, and reads a genuine modulation as an increasingly implausible sequence of chromatic chords in the original.

Between those there is a threshold, and the threshold is what the sweep finds. What makes it interesting is that there is not one threshold. It depends on how fast the two keys are alternating, because a fast alternation asks the model to pay the cost many times and a slow one asks it to pay once.

Two keys 4 bars at a time, read twiceAn alternating passage — 4 bars in C, 4 in G, 8 blocks — with the key actually sounding on the top row, a 8-bar histogram model's reading on the second, and an ordered model's on the third. The histogram is right where a block is long enough to fill its window and wrong at every change; the ordered model changes key when the grammar says a cadence has happened, which is not the same place.what is soundinghistogram, 8 barsordered modelCCCCCCCCCCCCGGCGCCGGCGGCCGCCGCCGCCGCGGCGGCGGCGGCCGCCGCCGCCGCGGCGGCGGCGGCCGCCGCCGCCGCGGCGGCGGCGGCbarshaded where the reading is right
Fig. 2 The model itself, from an earlier essay: a passage that alternates between two keys, with the ordered model’s reading through it. At the tuned cost the model names one key and calls the rest borrowings. Every number in this essay is that reading recomputed at other costs.

The trade, and what it means

The thresholds come out at 0.95 for alternations of one, two and four bars, 2.0 for eight and 3.5 for sixteen. The pattern is clear: the slower the alternation, the higher the cost has to be before the model refuses to hear two keys — and it is flat at the fast end, because below about four bars the passage is changing key faster than any evidence can accumulate and one threshold serves them all.

That is exactly right as a piece of inference and it is worth saying why. A passage that spends sixteen bars in one key and sixteen in another presents a great deal of evidence for each; the model has to be very reluctant indeed to refuse it. A passage that changes every bar presents one bar’s evidence at a time, and a small reluctance is enough to overrule it.

So the parameter and the rate trade off, and the model’s verdict about any particular passage is a point in a two-dimensional space rather than a fact about the music.

The tuned value of 2.2 sits above the thresholds for one, two, four and eight bars and below the one for sixteen. Which means the eighth rung’s headline — that the ordered model reads a fast alternation as borrowings — was not a finding about fast alternation. It was a finding about 2.2, and the model would have read any alternation shorter than about twelve bars the same way. What the tuned value does distinguish is a genuine long-range modulation from a short-range oscillation, which is a reasonable thing for it to do and is not what the eighth rung claimed.

That is not a refutation of the eighth rung, which said its number was tuned. It is what the sweep the rung asked for was for.

The one number the ordered key-finder was tuned on. For each rate of alternation between two keys, the cost of changing key at which the model stops hearing two keys and starts hearing borrowed chords in one. The threshold rises with the period — 0.53 at 4 bars — so the parameter and the rate trade off against each other exactly. The value tuned earlier, 2.2, sits above every threshold on this axis, which means its verdict about fast alternation was a consequence of the tuning rather than a finding about the music. Filled means the model names two keys; hollow means it names one and calls the rest borrowings.
Fig. 3 The same sweep for an alternation to a different key — the supertonic rather than the dominant. The threshold moves, because a different key shares a different number of chords and the evidence for a change is correspondingly stronger or weaker. So the parameter interacts with key distance as well as with rate, and the two-dimensional space above is really a three-dimensional one.

What would settle it, and why it is a corpus statistic

The parameter is a prior on how often keys change. That is not an opinion; it is a rate, and a rate is something a corpus measures.

Count key changes per bar across a body of music and the prior is determined. A repertoire in which keys change every four bars implies a low cost; one in which they change every thirty-two implies a high one. The number is different for a Bach chorale, a Schubert song and a jazz standard, and it should be — a key-finder analysing one of those should carry a different prior from one analysing another.

That is the honest form of what this rung produces: the parameter is not a nuisance to be tuned away, it is a property of a repertoire that a corpus would supply, and until somebody supplies it every reading this ladder produces is conditional on a guess.

Which puts this anchor in the same position as four others on this site: an argument that is complete except for a count nobody has made. The melody ladder needs post-skip reversal counted; the microtiming ladder needs a distribution of short-note durations; the clave rung needs a few hundred timelines. This one needs key changes per bar, which is the smallest of the four.

How far every key is from C major. The 11 other major keys under three measurements. Notes in common and steps round the circle are the same measurement — the first is seven minus the second until it bottoms out at two — and voice-leading distance between the tonic triads is not. E major shares 3 notes and is 2 semitones away; A♭ major shares 3 notes and is 2 semitones away, against the dominant's 6 and 3.
Fig. 4 Why the threshold depends on which key: three ways to measure how far a key is, which disagree with each other. A change to a near key needs less evidence than a change to a distant one under every one of the three, so the prior that decides whether a change is heard is doing different work at different distances — and this essay has swept it at only two.

The reading a tonal analyst would write

There is a thing worth noticing about the borrowing reading, which is that it is not the wrong answer.

A trained analyst handed a passage that alternates between C and G every four bars will usually write one key with a series of applied dominants and secondary chords, not eight modulations. That is standard practice, it is what the pedagogy teaches, and it is what the eighth rung noticed the ordered model doing.

So the model at a cost of 2.2 agrees with a human analyst, and the sweep says the agreement is a property of the number rather than of the model’s structure. The model was tuned, by hand, until it did what an analyst does — which is a perfectly reasonable way to set a parameter and is also the definition of fitting to the answer.

The honest reading is therefore not that the eighth rung’s finding is wrong. It is that the finding “the ordered model reads fast alternation as borrowings, which is what an analyst would write” contains no information about the model, because the parameter was chosen to make it true. What the model does contribute is the part the analyst does not: it says the threshold moves with the rate and with the key distance, and it says by how much.

How fast two keys can alternate before the finder stops following. The share of bars a moving key-finder names correctly, once its reading is shifted back by its own lag, against how many bars each key holds for. One line per window. Below a block of three bars the second key is never named at all — 2 of the sweep's readings report a single key for the whole passage — and above about twice the window the tracking is over ninety per cent. The lag itself is about half the window: 30 bars at a window of 2, 0 bars at a window of 4, 3 bars at a window of 8.
Fig. 5 The earlier comparison: the histogram model against the ordered one on the same alternating passage, at three window lengths. The histogram model cannot represent an alternation at all — a bag of notes has no order — and this essay’s sweep is about the one number that decides what the ordered model does instead.

Which computation produced the numbers

The passage is a repeating alternation between two keys, generated by taking a four-chord progression, transposing it, and alternating the two at a stated period in bars. The distance between the keys is a parameter and is a fifth in the hero figure and a fourth in the figure above.

For each cost, the ordered model is run over the whole passage and two things are read off: how often its per-bar key matches the key the passage was generated in, and how many distinct keys it names at all. The second is the verdict — one key is a borrowing reading and two is a modulation reading — and the threshold is the cost between the last value that names two and the first that names one.

The sweep runs from 0.05 to 10 in sixteen steps, roughly logarithmic, which resolves the thresholds to about twenty per cent. Below 0.05 the model changes key constantly and the reading is meaningless; above 10 it never changes and the reading is equally meaningless.

The thresholds at one, two and four bars are identical and that is a feature of the model rather than a resolution limit: below about four bars the passage supplies so little evidence per key that the model’s decision is made by the prior alone, and the prior does not know how fast the alternation is. The threshold only starts moving once each stretch is long enough to argue for itself.

The count the verdict should have been

The verdict above is a count of distinct keys, and the section on where the model stops calls that a coarse instrument without saying how coarse. There is a sharper one available for one extra pass of the same recursion: hold the state space to a single key, re-run, and compare the best one-key path with the best path of any kind. The difference is what the borrowing reading gives up, in the units the model already works in. (The next rung reports a different margin — per bar, against the runner-up key — and the two answer different questions; this one is per passage and is about the verdict this rung gives.)

The margin is exactly piecewise-linear in the cost, and each slope is exactly the number of key changes the winning reading uses. That is not an approximation to three decimal places; it is what a log-odds penalty means, and it makes the verdict legible in a way the count of distinct keys never could. The model is choosing how many times to change key so as to maximise the evidence gained less the number of changes times the cost, and every kink in the margin curve is one reading giving way to a cheaper one.

Which turns the coarse verdict into a sharp one. The passage alternates in eight blocks, so it contains seven true key changes at every rate, and the honest question is not how many distinct keys the model names but whether it recovers that seven. Sweeping the cost finely and recording every change count the model will produce:

alternation change counts the model will produce, and where
1 bar 5, 3, 1, 0 — seven never appears
2 bars 7 below 0.34, then 3, then 0
4 bars 15, 13, 3, 1, 0 — seven never appears
8 bars 33, 13, 11, 9, 7 between 0.80 and 0.92, 1, 0
16 bars 71 down to 10, 7 between 1.24 and 1.52, 1, 0

Two things follow and both are worse than a mis-tuned parameter. At one bar and at four the true reading is not reachable at any cost whatever — the model jumps from thirteen changes to three with nothing in between, so there is no prior on key change that makes it hear the passage correctly, and four bars is the rate every figure in this rung and the last is drawn at. And the three windows in which the true reading is reachable — below 0.34, between 0.80 and 0.92, between 1.24 and 1.52 — are disjoint. No single value of the parameter reads all five rates correctly, which is a stronger statement than the one this rung has been making.

The corpus argument survives it, but narrowed. A repertoire’s rate of key change would fix the parameter for that repertoire; what the disjoint windows say is that the right value also depends on the rate of the individual passage, and the rate of the individual passage is the thing the model exists to infer. A prior that has to be set from the answer is not doing the work of a prior.

There is also a smaller correction to make. Near every threshold in the hero figure the winning two-key reading uses exactly one key change — the margin’s slope is one there, at every rate — so the reading the count of distinct keys is calling a modulation is a single move to the second key and a stay, not an alternation. The two-key verdict and the alternation reading were never the same object, and the count could not tell them apart.

There is one more thing the same machinery will answer, and it is the question of whose confidence the margin is.

A listener has a quarter of an analyst's confidence and the same answer. The mean margin between a passage's best two key readings, against how many bars a listener's memory of the evidence takes to halve. The two-sided reading — the one that uses bars that have not happened yet — sits at 1.14; the forward pass with perfect recall at 0.68; a forward pass whose evidence halves every 6.6 bars at 0.56. The dots' size is how often that reading names the same key as the two-sided one: 94 per cent at perfect recall and 84 at a one-bar half-life. So forgetting costs a great deal of confidence and very little accuracy — the key is robust and the certainty is not.
Fig. 6 The mean margin between a passage’s best two readings, against how many bars a listener’s memory of the evidence takes to halve. The two-sided reading — the one that uses bars which have not happened yet — is the analyst’s, and it sits highest; a forward pass with a short memory is the listener’s.

A listener has a fraction of an analyst’s confidence and the same answer, which is the shape this ladder keeps producing: the reading is robust and the certainty is not. The parameter this essay is about sits inside the analyst’s reading, so a corpus that settled it would settle what an analyst should write down, and would leave the listener’s margin roughly where it is.

Where the model stops

The passages are constructed and not found. They are a progression alternating with its transposition, at an exact period, forever. Real music alternates irregularly, changes the progression as it goes, and modulates to more than two places — and a key-finder’s behaviour on constructed alternations is a diagnostic rather than a measurement.

The verdict is a count of distinct keys, which is a coarse instrument. A reading that names two keys and a reading that names two keys with one of them appearing for a single bar are both “two”, and they are very different analyses. The section before this one replaces the count with the number of switches and with the margin, and both guesses in the sentence that used to sit here turn out to be right — the threshold moves, and it is a gradient.

One parameter is swept and there are nine. The model also carries eight root-motion weights, which are ordinal and asserted, and the eighth rung said so. Sweeping the key cost with those held is an audit of one dimension of a nine-dimensional space, and there is no reason to think the thresholds found here are stable under changes to the others.

A cost is not a probability. The parameter is used as a log-odds penalty and interpreted as a prior on key changes, and the interpretation is only exact if the rest of the model is a properly normalised probability. It is not — the emission scores are similarities rather than likelihoods, and the root-motion weights do not sum to anything. So “count key changes in a corpus and the parameter follows” is a direction of travel rather than a recipe.

And borrowing is not modelled as such. The model has no representation of a borrowed chord: what it does is fail to change key and then score the foreign chord badly. A model that actually had borrowed chords as a category would have a different structure and would probably need a different parameter, and the equivalence this essay is about — a modulation and a borrowing as two readings of one evidence — is a claim about how the model behaves rather than about what it contains.

Whose music, and when

The distinction between a modulation and a borrowing is a piece of analytical vocabulary with a history, and it is not neutral.

Nineteenth-century German theory tends to hear tonicisation — brief excursions read as elaborations of one key — where earlier theory would hear a series of cadences in different keys. That is a shift in the prior, made by people, over a century, for reasons that had to do with a view of a piece as an organic whole rather than as a succession of events. The parameter this essay sweeps is that argument, written as a number.

Which repertoire it is set for matters. A Bach fugue subject answered in the dominant and a Wagner passage that touches six keys in eight bars are not the same object, and a single value applied to both will read one of them wrongly. The eighth rung’s 2.2 is, on the evidence here, a value that produces late-nineteenth-century readings of everything.

And there is a case for saying the argument is not resolvable by a corpus at all. A corpus count of key changes presupposes an analysis that says where the key changes, which is what the model is for — so the measurement is circular unless the analyses come from somewhere independent, and the somewhere independent is human analysts who disagree with each other along exactly this axis. That is the same circularity the key-plan rung works inside, and it is why every corpus of harmonic analysis carries an editor’s name.

What the picture cannot show

Which reading a listener has. The model produces a reading and there is no reason a listener produces one at all; how much evidence a modulation needs measured how many chords it takes before a listener is confident, and confidence is not the same as a labelled analysis.

One step out, in each direction, is not one thing. For each neighbouring key, the twelve pitch classes with the home key's seven marked, the new key's notes outlined, and the note the modulation adds picked out. The two closest keys both share six of seven notes with home; what separates them is which note the seventh is.
Fig. 7 What the model is weighing, chord by chord: for each neighbouring key, which pitch classes it shares with home and which one note gives it away. A change to a near key is announced by a single tone and a change to a far one by several, so the evidence a modulation supplies varies as much as the prior this essay sweeps. How much evidence a modulation needs measured the first and this essay sweeps the second, and nothing here has put the two numbers in the same expression.

And it cannot show ambiguity as ambiguity. The model always commits to one key per bar. A passage that is genuinely poised between two readings is one the model resolves arbitrarily, and the thing worth knowing about such a passage is precisely that it is poised — which is a property of the distribution over readings and not of the best one.

Where this ladder goes next

Nine rungs. Keys are neighbours and the map is computed; the key plan is the form; three distance measures disagree; a modulation takes a measurable number of chords to detect; the circle is a circle and the map is not; two keys at once are not in the vocabulary; two keys in alternation are followed only above a rate; a model that keeps the order can tell those apart; and now the number that decides which of the two readings it gives.

What is owed is the distribution the last section names. Every rung of this ladder produces a single best reading, and the interesting passages in the repertoire are the ones where two readings are nearly equally good. The machinery for that already exists inside the model — a dynamic program that finds a best path has, by construction, the scores of every other path — and reporting the margin between the best two readings instead of the best one would turn every figure in this ladder from an analysis into a measurement of how ambiguous a passage is.

And after that, a model whose state carries the rate. The disjoint windows above say the key cost is standing in for a quantity it cannot represent — how often this passage changes key, as against how often passages in this repertoire do — and a parameter asked to hold both will always be wrong about one of them. A second state variable, or a prior over rates with the rate inferred alongside the keys, is the shape of that repair, and it is a larger piece of work than any rung on this ladder has been.

Part 9 of 21

One essay in the series on Key-relations. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 10.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Borrowed chordHarmonic analysisKey distanceKey-findingModulationPriorSegmentationTonal hierarchy