Harmony and voice leading

Two keys at once

Every key figure so far assumes one key is sounding, and the standard key-finder has no value it can return that means two. Play one progression in C and the same progression a major third away at the same time, and the model does not report uncertainty: it reports E minor, at a correlation of 0.886, against 0.959 for the same progression in one key. It is as confident as it ever is, and neither key it names is being played. Give it the missing hypothesis — pairs of key profiles rather than single ones — and it recovers both keys at every separation, all seven of seven.

Assumes: Three ways to measure how far a key is · How much evidence a modulation needs

The key-finding machinery this site uses is a correlation. Count how much of each pitch class a passage contains, compare that histogram against twenty-four stored profiles — twelve major keys and twelve minor — and report the best match.

It is a good model, it is the standard one, and every rung of this ladder has run on it. Keys are neighbours because their profiles overlap. A modulation takes a measurable number of bars because the histogram takes that long to move. Three ways of measuring key distance agree almost everywhere, and the one place they disagree is instructive.

Every one of those results assumes there is a key. The model’s twenty-four answers are twenty-four single keys, and there is nothing in the list that means two.

What it does with two

The test is direct. Take one progression — I–IV–V–I — play it in C major, and play it simultaneously in another major key, and move the second key round the circle of fifths.

Two keys at once, 4 steps apart, as the finder sees it. I – IV – V – I played simultaneously in C major and E major, whose tonics are 4 steps apart round the circle of fifths. The pitch-class histogram is correlated against all twenty-four key profiles; the eight best are drawn. The winner is E minor at r = 0.886, ahead of the next by 0.207, and it is neither of the two keys sounding. Adding the hypothesis the model does not have — a PAIR of keys — gives C major with E major at r = 0.943, which is the two keys that are playing.
Fig. 1 C major and E major at once, which is four steps round the circle. The finder’s best answer is E minor at r = 0.886, ahead of the next candidate by 0.207 — a decisive margin. Neither of the keys sounding is in the top two. This is what the model does when the question it was built for has no answer: it answers a different question well.

The comparison that makes it stark is the monotonal case. The same progression in C alone correlates with the C major profile at 0.959. So the bitonal passage is read at 0.886 — lower, but nothing like the collapse a reader would expect, and higher than plenty of perfectly ordinary single-key passages.

At one step apart it is worse than that. C and G together give 0.961, which is higher than either key alone. The model is more confident about a two-key passage than about a one-key one, and it is not wrong to be: C and G together are very nearly the G major scale, and the finder has correctly found the collection. What it has not found is that two things are happening.

Run that across the whole circle and the pattern is consistent rather than occasional. The single-key finder’s confidence in its winner never collapses — it stays between 0.86 and 0.96 at every one of the seven separations, which is the range perfectly ordinary single-key passages occupy — and at four of the seven the answer it gives names neither key that is playing. A model that cannot say “two keys” does not say it is unsure; it says something else, fluently.

The one place it does admit defeat

There is one separation where the model reports something a reader could act on.

Two keys at once, 6 steps apart, as the finder sees it. I – IV – V – I played simultaneously in C major and F♯ major, whose tonics are 6 steps apart round the circle of fifths. The pitch-class histogram is correlated against all twenty-four key profiles; the eight best are drawn. The winner is C major at r = 0.358, ahead of the next by 0.000. Adding the hypothesis the model does not have — a PAIR of keys — gives C major with F♯ major at r = 0.899, which is the two keys that are playing.
Fig. 2 C major and F♯ major together, a tritone apart, which is the maximum separation there is. The correlation collapses to 0.358 and there is an exact tie between the two keys sounding — the only case in the sweep where both are named and the only one where the confidence is low enough to be a warning. Two keys a tritone apart share nothing, so the histogram they produce is nearly flat, and a flat histogram correlates with everything equally badly.

The tie is worth dwelling on for a second reason: it is the only reading in the sweep that a downstream consumer could act on. Every other separation produces a single winner with a comfortable margin, and a margin is what a program checks. A pipeline that asked “is the key confidently determined?” would answer yes at four steps apart and no at six, which inverts the truth about how much is going on.

That is the model working properly, and it is worth seeing why it only works here. A correlation reports a shape match. Two keys close together produce a histogram shaped like a third key, and the model finds that third key. Two keys maximally far apart produce a histogram with no shape at all, and the model has nothing to find. The failure is loudest where the music is least ambiguous, and silent where the music is most ambiguous — the opposite of a useful error signal.

Why the tritone case is flat, and what the chord turns out to be

The collapse at six steps has a cause that can be counted rather than described. I–IV–V–I in C uses seven pitch classes; the same progression in F♯ uses the other five and two of the first seven. Together they use all twelve, once or twice each, and a histogram that is nearly flat correlates with every profile equally badly.

That is also the arithmetic behind the most famous bitonal object there is. A C major triad against an F♯ major triad — the Petrushka chord — is the six notes C, C♯, E, F♯, G and B♭, and those six are a subset of the octatonic collection. More than that: the six-note set is symmetric at the tritone. Transpose it by six semitones and it returns itself.

The collection the Petrushka chord lives in. The twelve semitones drawn as a cycle, with the notes of the scale filled in. The gaps between filled positions are the step pattern, and reading them round the circle is what makes the scale's asymmetry obvious.
Fig. 3 The octatonic collection, of which the C-against-F♯ chord is a six-note subset. A set with a transposition symmetry has fewer modes than notes and no preferred starting point — which is exactly what a key-finder is looking for and exactly what this one does not have. The reason the correlation collapses and the reason the chord sounds rootless are the same reason.

Run the finder on the six notes alone and it returns E minor at r = 0.315, tied with its runner-up to three decimals. A tie at a third of the correlation an ordinary phrase produces is the model saying, as clearly as its vocabulary allows, that there is nothing here of the shape it looks for.

The hypothesis it does not have

The repair is not a better correlation. It is a longer list of candidates.

Add every pair of the twenty-four profiles — 300 of them, counting each key with itself — sum the two profiles, and correlate the histogram against the sum. The pair model has exactly one more free parameter than the single-key model: the second key.

A model that cannot say two keys does not say it is unsure. I – IV – V – I played in C major and in a second major key at the same time, with the second key moved round the circle of fifths. For each separation: the correlation the standard key-finder gives its single best answer, and the correlation reached by the best PAIR of key profiles — a hypothesis the finder does not have. The pair recovers both keys that are sounding at every separation, 7 of 7. The single answer names neither of them at 4 of the 7, and its confidence does not fall when it is wrong: at four steps apart it reports E minor at r = 0.886, against 0.959 for the same progression in one key.
Fig. 4 The pair fit is the upper line, and it recovers both keys that are sounding at every one of the seven separations — including the tritone, where it reaches 0.899 against the single-key answer’s 0.358. The gain is largest exactly where the single-key model is worst, which is what a missing hypothesis looks like when it is supplied.

Two details of that are worth stating because they are what stop it being a trick. The pair profile is the plain sum of the two, not a weighted mixture — a weight would be a third parameter and would let the pair model win by fitting noise. And at zero separation the best pair is C major with itself, at exactly the single-key correlation.

That last check is one progression, and it is the one that passes

A single-key control at one point is not a specificity test, and running it properly changes what the pair model can be claimed to do. Handing it eight ordinary single-key progressions in C:

progression single key best pair pair names
I–IV–V–I 0.959 0.959 C major twice
I–IV–I–V–I 0.944 0.964 C major + E minor
I–vi–IV–V 0.935 0.974 C major + A minor
I–V–vi–IV 0.935 0.974 C major + A minor
ii–V–I 0.881 0.948 G major + D minor
I–ii–V–I 0.877 0.966 C major + G major
I–iii–IV–V–I 0.858 0.969 C major + E minor
vi–ii–V–I 0.817 0.963 G major + A minor

Seven of the eight get a second key that is not there, and the one that does not is the progression the essay already tested. I–IV–V–I is the most nearly key-defining sequence in the repertoire — three chords covering the whole diatonic set with the tonic twice — and it is the single best case for a one-key hypothesis. Checking there and stopping is checking where the answer was known.

The spurious pairs are not noise, which makes it worse rather than better. C major plus A minor for a progression containing vi, C major plus E minor for one containing iii: the pair model is finding a real property of the histogram, which is that a progression touching the relative minor looks partly like the relative minor. That is a good description of the music and it is not a second key, and nothing in the model’s output distinguishes the two readings.

What that leaves standing

The rung’s central finding does not depend on the pair model at all. That a single-key finder reports E minor at 0.886 for a passage in C and E, with a decisive margin over the runner-up, is a fact about the single-key model measured against a known ground truth; the pair model is not needed to establish it.

What the pair model was offered as is a diagnostic — evidence that the information needed to name both keys survives in the histogram. It still shows that, because it recovers both keys at all seven separations. What it cannot be is a detector, because it also reports two keys on seven of eight passages that have one. A test that fires on the positives and on almost all of the negatives has established that the signal is present and has established nothing about whether it can be found.

So the honest statement is narrower and is the one the essay’s own caveat was reaching for. The information is there; a model with one extra free parameter can always reach it; and one extra free parameter is enough to reach it when it is not there either. Separating the two cases needs a penalty for the second key, or a held-out test, and neither is in this construction.

Two keys at once, 1 step apart, as the finder sees it. I – IV – V – I played simultaneously in C major and G major, whose tonics are 1 step apart round the circle of fifths. The pitch-class histogram is correlated against all twenty-four key profiles; the eight best are drawn. The winner is G major at r = 0.961, ahead of the next by 0.183. Adding the hypothesis the model does not have — a PAIR of keys — gives C major with G major at r = 0.975, which is the two keys that are playing.
Fig. 5 C and G together, one step apart, and the finder’s winner is G major at 0.961 with a margin of 0.183 over the next candidate. That is higher than either key scores alone, and it is not a mistake: C and G together are very nearly the G major scale, and the finder has correctly found the collection. Keys a step apart share six of their seven notes, so two of them sounding at once produce a histogram shaped almost exactly like a single key — which is why the confusion at one step is a real ambiguity in the sound, and the confusion at four is not.

What a listener does, which is neither of these

None of this says a listener hears two keys. The evidence is that mostly they do not.

Bitonal writing is a small and deliberate repertoire — Stravinsky’s Petrushka chord, Milhaud’s Saudades do Brasil, Ives’s marching bands — and the reports of what it sounds like are consistent: not two keys, but one strange one. A C major triad against an F♯ major triad is heard as a chord, and the chord it is heard as has a name in the harmony books. So the single-key model’s answer is not a failure of realism; it is a fairly good model of a listener, and the pair model is a better model of the score.

The probe-tone profile, major against minor. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.
Fig. 6 The profiles the whole apparatus rests on: the major and minor key profiles from probe-tone experiments, in which listeners rated how well each of the twelve notes fitted after hearing a key established. Counting produced the hierarchy — these numbers are averages over listeners, and a listener who genuinely heard two keys at once would be averaged away in them.

What the experiment cannot resolve is exactly the question this rung raises. A profile measured by asking listeners to rate one note at a time against one established context has no way to represent a listener holding two contexts, and its averages would look identical whether nobody does it or half of everybody does.

There is a second reason to expect a listener to hear one key rather than two, and this site measured it. Establishing a tonic is not free: it costs a measurable share of the time spent on one note before a key-finder will name that note, and the same asymmetry has to be paid by whatever a listener does instead. Two keys sounding together are two claims on the same evidence, each weakening the other, and a mechanism that needs a fifth of a passage’s weight on one pitch class before it commits is not going to commit twice.

And the collection has an older result pointing the same way. Nothing in a pitch-class census knows which note is home — a set has no tonic until something in time supplies one. Bitonality asks a listener to supply two from one stream, and the supplying is the expensive part.

How fast two keys can alternate before the finder stops following. The share of bars a moving key-finder names correctly, once its reading is shifted back by its own lag, against how many bars each key holds for. One line per window. Below a block of three bars the second key is never named at all — 2 of the sweep's readings report a single key for the whole passage — and above about twice the window the tracking is over ninety per cent. The lag itself is about half the window: 0 bars at a window of 3, 0 bars at a window of 4, 3 bars at a window of 8.
Fig. 7 The same two keys taken in turn instead of at once, which is the object the histogram cannot tell from the other one. How well a moving finder follows an alternation, against how many bars each key holds: below a block of three bars the second key is never named at all, and above about twice the window the tracking is over ninety per cent. So the same two keys are trivially separable when they alternate slowly, indistinguishable from one key when they alternate fast, and — as the figures above show — read as a third key when they are simultaneous. Nothing about the content changes between those three cases. Only the arrangement in time does, and a histogram has no time in it.

Which computation produced the numbers

The histogram is the union of the two progressions’ pitch classes, each chord counted once. The profiles are Krumhansl and Kessler’s, correlated with Pearson’s r, which is the standard construction and the one every other figure in this ladder uses.

The pair search is exhaustive: all 300 unordered pairs of the twenty-four profiles, summed and correlated the same way, with the best reported. There is no fitting and no threshold anywhere in it.

The specificity test is the same search run on histograms with one key in them, eight ordinary progressions in C, chosen before their results were looked at and covering the common shapes — with and without vi, with and without ii, with the tonic once and twice. Nothing about them is adversarial; the point is that they are the material a key-finder is normally handed. The measure reported is not the correlation but whether the winning pair names two distinct keys, which is the only output a reader could act on.

The asymmetry between the two tests is worth stating plainly, because it is what a specificity test is for. The bitonal sweep has seven cases and the pair model gets seven right; the single-key control has eight cases and the pair model gets one right. A test that fires on every positive and on seven of eight negatives has a sensitivity of one and a specificity of one eighth, and a number quoted from the first without the second is not a measurement of anything.

The separations are steps round the circle of fifths rather than semitones, because that is the axis on which key distance is one-dimensional — among the twelve major keys, notes in common is a strict function of steps round the circle, so a sweep over steps is a sweep over overlap.

Whose music, and when

The key-finding model is a twentieth-century psychological one applied to eighteenth- and nineteenth-century tonal practice, and it is good at that. Polytonality is an early-twentieth-century device, used deliberately and sparingly, and the argument here is not that the model should have been built to handle it.

The device is also narrower than it sounds. What is usually written is not two keys of equal standing but a home key with a second layer over it, and the layer is normally a triad or an ostinato rather than a functioning tonality — which is why the results here are best read as an upper bound on the confusion. A real bitonal passage gives the finder more of one key than the other, and the finder then names that one, correctly, while missing the layer entirely.

The argument is about what a model’s confidence means. A model whose hypothesis space excludes the truth does not report low confidence; it reports high confidence in the best available falsehood — and the practical consequence is that a key-finder’s correlation cannot be read as a measure of tonal clarity. A passage the model reads at 0.886 might be an ordinary tonal phrase or might be two keys at once, and nothing in the number distinguishes them.

This site has now found that shape twice from opposite directions, and the two arrived independently. A metre induction handed a Balkan bar of nine ranks three candidates confidently and the right answer is not among them. Neither model is broken; both are complete over a space that does not contain the case.

What the picture cannot show

Both voices are the same progression. A real bitonal passage has two independent lines, and giving the two keys different progressions would change the histogram and every number here. The identical progression is chosen so that the only variable is the separation.

Register is discarded. The histogram counts pitch classes, so the two keys are on top of each other in a way no performance is. In practice bitonal writing separates the layers by register and timbre, which is most of what makes it legible, and none of that reaches this model.

Nothing is weighted by time. Each chord counts once. The site’s own window-based reading shows that when the histogram is taken matters as much as what is in it, and a bitonal passage in which the two keys take turns is a different object from one in which they are simultaneous.

The key plan of a whole movement is a different object. A classical movement’s shape is a journey through keys taken one at a time, and nothing in that argument is touched here: a plan is a sequence, and this rung is about simultaneity.

And the pair model is not proposed as a theory. It is a diagnostic: its job is to show that the information needed to name both keys is present in the histogram and that the single-key model discards it. A model with 300 candidates will out-fit one with twenty-four on almost anything — and the section above finds it does exactly that on seven of eight single-key progressions, so the zero-separation case that was offered as the control turns out to be the one progression in the set that passes it. The diagnostic reading survives; a detection reading does not.

Two keys alternating every 4 bars, and what the finder says. Above, which key is actually sounding in each bar of a 32-bar passage that alternates between two keys a fifth apart every 4 bars. Below, what a moving key-finder with a window of 4 bars reports. It names the right key in 72 per cent of bars as they stand, and in 72 per cent once the reading is shifted 0 bars earlier — so the finder is right or wrong rather than late.
Fig. 8 One alternation at four bars a key, bar by bar, which is where the failure is legible. The top row is which key is actually sounding across thirty-two bars; the bottom is what a four-bar window reports, and it is right in 72 per cent of them — wrong at every changeover and for a bar or two after. That is the same evidence problem the simultaneous case has, spread out: a modulation to a near key turns on one note the home key does not use, and in the bitonal case that note sounds at the same time as the note it replaces. The evidence for a modulation and the evidence for a bitonality are the same evidence, and only their arrangement in time tells them apart.

Where this ladder goes next

The rungs above have measured how far apart two keys are, how long a change between them takes, and which of three distance measures disagrees with the others. This one asks what the machinery does when two of them are true at once, and the answer is that the question is not in its vocabulary.

The next step is one this rung can point at and did not take. Run the moving key-finder over a passage that alternates between two keys bar by bar rather than sounding them together, and at every window this ladder has used — two bars, four, eight — the second key is never the answer at any point in the passage. An alternation faster than the window is, to this model, indistinguishable from a bitonality, because a histogram has no order in it and a window that spans both keys contains both keys’ notes whatever order they arrived in.

That is a rung rather than a remark, because the claim is checkable in the other direction too: there should be an alternation rate at which the reading starts tracking, it should be set by the window, and the window was measured three rungs ago. What is missing is not the machinery but the material — a passage whose alternation rate can be varied while everything else is held still, which is a construction rather than a piece, and the construction is the next rung’s job.

Part 6 of 21

One essay in the series on Key-relations. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Circle of fifthsEnumerationKey-findingKey-relationsModulationProbe-toneTonal hierarchyTonic