Harmony and voice leading

A key-finder that keeps the order

Every key-finding model until now begins by throwing order away — a histogram over a window, correlated against twenty-four profiles — and the thing that most obviously declares a key is an ordered pair of chords. A model that keeps the order costs 84 states and 7,056 transitions against 24 hypotheses and none, and what it buys is the one test a histogram was said not to pass: two keys in turn and two keys at once have histograms 98 per cent alike, and are two different sequences.

Assumes: The alternation a key-finder cannot follow · Two keys at once

The alternation a key-finder cannot follow ended by naming a rung and by saying it would be the first in this ladder to need a different kind of model. Every method here begins by throwing order away. A key-finder that kept it — one that scored a sequence of chords against a grammar rather than a bag of pitch classes against a profile — would be able in principle to tell an alternation from a simultaneity, and the interesting question is not whether it could but how much it would cost.

The test a histogram cannot pass. Two passages built from the same two keys: one takes them in turn and the other sounds them together. Over a window long enough to hold both, their pitch-class histograms are 98.6 per cent alike, so they are very nearly the same object to a profile model — and it names two keys in both. The ordered model names up to four for the alternation and one for the simultaneity, because a simultaneity has no sequence in it that any single key's grammar will not fit.
Fig. 1 Two passages built from the same two keys: one takes them in turn, the other sounds them together. Their pitch-class histograms are 98 per cent alike over any window long enough to hold both, so a profile model sees very nearly one object and names two keys in each. The ordered model names several keys for the alternation and exactly one for the simultaneity, which is the distinction said earlier to be unavailable.

It is worth saying which of the two readings is the more surprising. That an ordered model can tell the two apart is expected — it is what the extra structure is for. That the histogram cannot is the stronger half, and it is not a matter of degree: 98 per cent alike is the figure over a window long enough to hold both keys, and the residual two per cent is the difference in how often each chord happens to appear rather than anything about order at all. Shorten the window and the two passages become identical rather than nearly so.

The two passages are the same object to a histogram and different objects to a grammar. That is the whole case for the more expensive model, and it is worth having in one figure because it is not an improvement in accuracy — both models are about equally right about which keys are present — but a difference in what can be asked.

Why the histogram cannot, and it is not a defect of the profiles

It is worth being precise about the impossibility, because it is stronger than “the histogram happens to be bad at this”.

A pitch-class histogram over a window is a count. It records how much of each of the twelve pitch classes was present and nothing about when. Two passages containing the same notes for the same total durations have identical histograms whatever order those notes came in — that is what a count is — and an alternation and a simultaneity built from the same two progressions contain very nearly the same notes for very nearly the same durations.

So no choice of profile can separate them, and no choice of window either: a window short enough to see one key at a time cannot see the alternation, and one long enough to see the alternation contains both keys and is a simultaneity as far as the count is concerned. The seventh rung found this as a failure at a rate; it is a failure in principle.

How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is.
Fig. 2 The profile model doing the job it was built for, from an earlier essay: a single modulation, three window widths, and the lag before the new key is named. It is good at this and the window trade is the whole of its behaviour — narrow windows are quick and noisy, wide ones are slow and sure. Nothing about that trade helps with the problem above.

What the ordered model is

A hidden Markov chain over pairs. The state is (key, degree): which of the twelve major keys the music is in, and which degree of it the current chord is. Twelve times seven is eighty-four states.

The emission is how well the chord actually sounding matches that degree’s triad in that key — shared tones over the union, which is the crudest possible measure and is deliberately crude, because the whole argument is about the transitions.

The transitions carry a grammar. Within a key the weight depends on the root motion: down a fifth is the strongest, up a step next, a repeat weakest — which is the same ordering the progression ladder’s map is drawn from, read as probabilities rather than as distances. Between keys there is a fixed cost, and it is the one free parameter.

Decoding is Viterbi: the single most probable path through the whole passage, which is a global reading rather than a windowed one. That is the second difference from the profile model and it may be the more important — a histogram model has a window and must choose its width, and this one does not.

Two keys 4 bars at a time, read twiceAn alternating passage — 4 bars in C, 4 in G, 8 blocks — with the key actually sounding on the top row, a 8-bar histogram model's reading on the second, and an ordered model's on the third. The histogram is right where a block is long enough to fill its window and wrong at every change; the ordered model changes key when the grammar says a cadence has happened, which is not the same place.what is soundinghistogram, 8 barsordered modelCCCCCCCCCCCCGGCGCCGGCGGCCGCCGCCGCCGCGGCGGCGGCGGCCGCCGCCGCCGCGGCGGCGGCGGCCGCCGCCGCCGCGGDGGDGGDGGDbarshaded where the reading is right
Fig. 3 Both models bar by bar on one alternating passage: what is sounding on the top row, an eight-bar histogram’s reading on the second, the ordered model’s on the third. The histogram is right in the middle of a block and wrong at every change, because a window straddling a change contains both keys. The ordered model changes where its grammar says a cadence has closed, which is not the same place and is not always the right one either.

The two failure modes are worth naming because they are the two things an analyst does. A model that smears is hedging; a model that assimilates is committing. Neither is wrong as an analysis and they answer different questions — what is sounding now and what is this passage in — and the histogram is built to answer the first while the grammar is built to answer the second. That the two answers differ is a property of the music rather than a defect of either.

What it costs

The seventh rung stated the trade precisely: what a model can distinguish against how many hypotheses it must carry. Both halves are now numbers.

What keeping the order costs, on a logarithmic axis. The two models' sizes. A histogram model carries 24 hypotheses — twelve major keys and twelve minor — and no transitions at all, because it has nothing to transition between. An ordered model over (key, degree) carries 84 states and 7,056 transitions between them. That is a factor of 3.5 in hypotheses and an infinite factor in transitions, and it is the trade named earlier: what a model can distinguish against how many hypotheses it must carry.
Fig. 4 The two models’ sizes on a logarithmic axis. Twenty-four hypotheses and no transitions against eighty-four states and 7,056 moves between them. The states are a factor of three and a half; the transitions are a factor of infinity, because a model with nothing to transition between has none.

Eighty-four states and 7,056 transitions. That is small — a laptop decodes a hundred-bar passage in twenty milliseconds — and it is the wrong way to read the number. What matters is that the transitions are specified: every one of the 7,056 is a weight somebody has to choose, and the profile model’s twenty-four hypotheses need only twelve numbers each, taken once from a published experiment.

Though in this implementation they are not 7,056 independent choices. Every transition weight is built from two things — a root-motion table with a handful of entries and one number for changing key — so the model has fewer than ten free parameters and reconstructs the rest. That makes it much cheaper to justify than the raw count suggests and much more exposed to any one of them, which the last section of this rung is about: a model with 7,056 transitions and eight parameters is a model in which one parameter moves thousands of weights at once.

The pair model paid the same trade when it went from twenty-four candidates to three hundred, and it paid it in the same currency: more to distinguish, more to specify.

That is the honest reading of the cost. A model with more parameters is cheaper to run and more expensive to justify, and the justification is where the corpus debt lands. The profile model’s twenty-four hypotheses rest on one published experiment that has been replicated; this one’s 7,056 transitions rest on eight numbers chosen by somebody who knows the repertoire.

Where the extra structure is not free

The ordered model is worse at the thing the histogram is good at, and the reason is instructive.

A histogram over a window is a local reading and gives an answer for every bar independently. Viterbi is a global one: it finds the single best path, and a path is penalised for changing key, so a passage that changes often is read as a passage that changes less often than it does. Below about three bars a block the ordered model gives up and names one key throughout.

Two keys 2 bars at a time, read twiceAn alternating passage — 2 bars in C, 2 in G, 8 blocks — with the key actually sounding on the top row, a 8-bar histogram model's reading on the second, and an ordered model's on the third. The histogram is right where a block is long enough to fill its window and wrong at every change; the ordered model changes key when the grammar says a cadence has happened, which is not the same place.what is soundinghistogram, 8 barsordered modelCCCCCCGCCGCCCCGCCGGCGGCGCCFCCFGCFGCFCCCCCCGCCGCCbarshaded where the reading is right
Fig. 5 The same comparison at two bars a block. The histogram is wrong nearly everywhere, because its window never contains one key alone. The ordered model is wrong everywhere in a different way: it has decided the passage is in one key and read every bar of the other key as a borrowed chord, which is a perfectly reasonable analysis and is not what is happening.

Three bars is a suggestive number, and the one free parameter decides it. Sweeping the key cost and asking for the shortest block the model still tracks:

key cost shortest block it follows
0.5 4 bars
1.0 8
1.5 12
2.2 never

At the cost these figures use it is about three or four bars, and at the model’s own default of 2.2 there is no block length at which the alternation is tracked at all — the path takes the whole passage as one key and reads every bar of the other as borrowed. So the number is not a property of the construction; it is the parameter, read out.

That removes the coincidence the paragraph above was reaching for. The amount of evidence a modulation needs is a measurement made against a window filling with pitch classes; this is a penalty somebody chose. The two agree at one setting out of the four tried and disagree at the other three, so the agreement is a fact about the setting rather than about two constructions meeting.

What the sweep does establish is the shape of the trade, and it is steeper than more penalty means slower to change. The threshold roughly doubles for each half-point of cost — four bars, eight, twelve — so the parameter is not a fine adjustment: two neighbouring plausible values give models that disagree by a factor of two about how long a passage has to stay put before the model will believe it.

One row of the sweep is worth flagging as a caution rather than a result. At the essay’s own cost the model tracks a four-bar block and an eight-bar one and fails a six-bar one, which is not monotone in the block length. A Viterbi path is a single global commitment and it can lock onto a phase that fits most of the passage and mis-reads the rest; that is a property of taking one best path rather than a distribution over them, and it means the threshold above should be read as about four bars and not as a boundary a passage crosses cleanly.

Both models fail fast alternation and they fail it differently. The histogram smears; the grammar assimilates. A listener does something else again, and how much evidence a modulation needs put that at a handful of chords, which is between the two.

There is a third thing the ordered model buys that is worth recording even though this rung does not use it. Its state is (key, degree), so its output is not a key but an analysis: at every bar it names the key and the roman numeral together. A profile model returns a key and leaves the numeral to be worked out afterwards from the key, which is the order every textbook uses and is the opposite of what a listener plausibly does — a cadence is recognised as a cadence before its key is named.

Which computation produced the numbers

The passage is the seventh rung’s own alternatingPassage, unchanged: a four-chord progression in one key, then the same progression transposed, in blocks of a stated length.

The histogram reading is keyReading, also unchanged — a weighted pitch-class histogram over a window, correlated against the twenty-four rotated Krumhansl–Kessler profiles, exactly as the fourth rung of this ladder built it.

The ordered reading is a Viterbi decode over the eighty-four states. The emission is the shared-tone match above; the transition weights are eight ordinal numbers for the eight root motions, chosen to put descending fifths first and stated in the source as ordinal rather than measured; the key-change cost is one number and is the only thing tuned, at 0.5 in log units.

The key-change cost is where the model can be attacked. At zero it changes key freely and names six; at two it never changes and names one. The value used here is chosen so that it names exactly the two keys that are present at block lengths of three bars and more, which is fitting a parameter to the answer — and it is stated here rather than buried because the control result does not depend on it. At every value between 0.2 and 0.8 the model separates alternation from simultaneity, and the histogram does not at any window width.

The histogram similarity between the two passages is the cosine of their pitch-class distributions, which is 0.979 at four bars a block and above 0.92 everywhere.

Where the model stops

The grammar is asserted. Eight numbers for eight root motions is a statement about the common practice, not a measurement of it, and a corpus would give better ones — which is the same corpus debt three other ladders on this site have recorded.

Twelve keys, not twenty-four. The model carries major keys only, which halves the states and removes the hardest confusion in key-finding, between a major key and its relative minor. Adding the minors takes it to 168 states and 28,224 transitions and needs a minor grammar, which is a different set of eight numbers because the minor mode’s degrees are not the major’s rotated.

And a chord is given, not found. The passage arrives as a list of chords. A real key-finder is handed notes, and which notes are the chord is a decision that depends on the metre, which depends on the segmentation — the three-way problem the previous rung of the progression ladder solved jointly. Putting an ordered key model inside that search multiplies its cost by eighty-four.

Two keys 12 bars at a time, read twiceAn alternating passage — 12 bars in C, 12 in G, 4 blocks — with the key actually sounding on the top row, a 8-bar histogram model's reading on the second, and an ordered model's on the third. The histogram is right where a block is long enough to fill its window and wrong at every change; the ordered model changes key when the grammar says a cadence has happened, which is not the same place.what is soundinghistogram, 8 barsordered modelCCFCCFCCFCCFCCCCCCCCCCCCCCGCCGCCGCCGGCGGGGGCGGGGGGDGGDGGDGGDGGCGGCGGCGGCCGCCGCCGCCGCCCGCCGCCGCCGCCFCCFCCFCCFGCGGGGGCGGGGGGGGGGGGGGGGGGDGGDGGDGGDbarshaded where the reading is right
Fig. 6 And at twelve bars a block, where both models are comfortable. The histogram is right on nine bars in ten and the ordered model on more, and neither is doing anything interesting — a block long enough for a window is a block long enough for anything. The two models are only worth distinguishing in the region where at least one of them is in trouble.

Whose music, and what the model is a model of

The claim is about a class of algorithm rather than about a repertoire, and the algorithms in question were written for one: the profiles are Krumhansl and Kessler’s, measured on listeners raised on European tonal music, and the grammar’s descending fifths are that tradition’s own.

Where it becomes a claim about music is in the construction the whole rung turns on. Alternation and simultaneity are both real practices and neither is common. Rapid alternation between two keys a fifth apart is a device of the eighteenth-century galant style — the phrase that answers in the dominant and returns — and simultaneity, in the sense of two keys genuinely sounding at once, is a twentieth-century one, Stravinsky and Milhaud and Ives — which the bitonal rung found is mostly not heard as two keys at all.

That the two produce nearly identical histograms is not a coincidence about these two passages; it follows from the profile model being a statistic over a window, and a window long enough to see an alternation cannot tell when things happened inside it. So a profile key-finder run over the Petrushka chord and over a galant antecedent–consequent pair will report something similar about both, and the difference between a joke about tonality and a normal use of it is exactly the thing it discards.

What the picture cannot show

Whether a listener does either of these things. The ordered model is not proposed as a psychological account; it is proposed as the cheapest thing that passes a test the other model fails. Counting produced the hierarchy is the essay about where the profiles came from, and nothing equivalent exists for a grammar.

Nor whether the simultaneity is heard as one key. The ordered model names one key for the simultaneity, which is a claim about the model rather than about the ear. What two keys at once established is that a listener hearing genuine bitonality mostly does not hear two keys either — so the ordered model’s answer may be right for the wrong reason, and the test it passes is a test of discrimination rather than of correctness.

And the test is a constructed one. Both passages were built for this comparison, from one progression and its transposition. A model that separates two artificial passages has demonstrated a capability rather than an accuracy.

What this does not settle

One thing is worth marking as deliberately not attempted. The rung’s question was whether an ordered model could tell an alternation from a simultaneity, and the answer is that a small one can. Whether an ordered model is a better key-finder in general is a different question, it needs a corpus with ground-truth key annotations, and this collection has neither the corpus nor the annotations.

What is available without a corpus is a capability test, which is what has been run. That is a weaker kind of result and it is the kind this collection can produce honestly, and it is worth saying plainly rather than letting a constructed passage look like an evaluation.

Where this ladder goes next

Eight rungs. Keys are neighbours and the map is computed; the key plan is the form; three distance measures disagree; a modulation takes a measurable number of chords to detect; the circle is a circle and the map is not; two keys at once are not in the vocabulary; two keys in alternation are followed only above a rate; and now a model that keeps the order can tell those last two apart, at eighty-four states and 7,056 transitions.

The rung after it is the one the assimilation failure names. Both models are wrong about fast alternation and one of them is wrong in an interesting way: the ordered model reads the second key as borrowed chords in the first, which is what a tonal analyst would write. So the two readings — a modulation and a set of borrowings — are the same evidence under two different priors on how often keys change, and the parameter that separates them is the one number this rung tuned. That would make the key-change cost the object of study rather than a nuisance, and it is measurable: how often keys actually change is a corpus statistic.

Part 8 of 21

One essay in the series on Key-relations. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 15.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

BitonalityCadenceInferenceKey-findingKrumhansl schmucklerModulationNull modelPitch-class profile