Harmony and voice leading

A count and a correlation

Cadence evidence is a count of ordered pairs and profile evidence is a correlation with a template, and the cadence essay refused to total them because they are not in the same units. Score each against its own chance and they are — standard deviations above chance are the same unit whatever produced them. Then the trouble moves: the obvious null does nothing at all to one of the two, because shuffling the bars leaves a pitch-class histogram exactly as it was.

Assumes: The cadence as evidence · Keys are neighbours, and the map is computed

The sixth rung of this ladder found that the ordered pairs a progression is made of carry key evidence a bag of notes cannot — two passages one note apart give the same histogram and different cadences — and then ran into the problem that stopped it:

Cadence evidence is a count and profile evidence is a correlation, and a key-finder that used both would have to put them in the same units — which is not a technical difficulty but the same question the closure ladder refused when it built the vector.

There is a way to put a count and a correlation in the same units and it needs no weights: score each against its own chance. How many standard deviations above what the same music would give at random — that is a dimensionless quantity, it means the same thing whatever produced it, and it is the standard move in every field that has to combine incommensurable evidence. It is also what the overlap check this site runs on its own essays does with a completely different pair of quantities.

Two kinds of evidence, each measured against its own chance. Every key as a point: across, how many standard deviations its cadence count is above what the same bars resampled would give; up, the same for its pitch-class profile correlation. The winner on cadences is key C at 6.2 standard deviations and on profile C at 1.1, and they agree. Standard deviations above chance are the same unit whatever produced them, which is the commensuration this figure is for — and it costs something: it assumes the two chances are equally interesting, which is a weighting in disguise. The obvious null does not work at all for one of the two: shuffling the bars leaves a pitch-class histogram exactly as it was, so its spread is zero and every key scores nothing against it.
Fig. 1 Every key as a point: across, how many standard deviations its cadence count is above what the same bars resampled would give; up, the same for its pitch-class profile correlation. The two agree — both name C — and they agree with very different confidence, 6.2 standard deviations against 1.1. A thousand draws build each null, which is enough for both numbers to have stopped moving.

Why standardising works, and what it costs

The reason a count and a correlation cannot be added is that they have different units and different scales: a cadence score might run from 0 to 30 and a correlation from −1 to 1, and any weighting of the two is a decision about how much a point of one is worth in the other.

Standardising removes the units by dividing each by its own spread. The question each becomes is: how surprising is this value, if the music were arranged at random? That question has the same meaning for both statistics, and the answers are comparable.

What it costs is stated rather than hidden. It assumes the two chances are equally interesting — that being three standard deviations above chance on a cadence count deserves the same weight as being three above chance on a correlation. That is a weighting, made invisible by being expressed as a convention, and it is the weakest step in this essay.

There is no way round it. The sixth rung’s position — report the vector and let a reader see the disagreement — remains the honest one, and this rung adds a defensible default total rather than a correct one.

Which chords of a key fit the key. The seven triads of the major scale ranked by how well their notes fit the Krumhansl–Kessler probe-tone profile for that mode, taking each chord's fit as the mean of its three notes' ratings. The tonic is first, at 5.31; the second-placed chord is vi at 4.80, which is 90 per cent of it.
Fig. 2 The profile half of the evidence: a pitch-class histogram correlated against Krumhansl and Kessler’s template for each of the twenty-four keys, which is what counting produced. It is a correlation, it lives in [−1, 1], and nothing about it is in the same units as a count of cadences.

The null that does nothing

Building the null is where the interesting thing happened, and it is a result rather than a nuisance.

The obvious null for a claim about order is to shuffle the order: take the same bars, permute them, and recompute. That is exactly right for cadence evidence, which is a statistic about ordered pairs — shuffling destroys the pairs and leaves everything else, which is the definition of a good null.

Applied to the profile correlation it does nothing at all. A pitch-class histogram is a bag of notes; the same bars in any order give the same bag; so the correlation is identical under every permutation, its spread across the null is zero to machine precision, and every key scores exactly zero standard deviations above chance.

That is not a bug. It is the sixth rung’s own finding, arriving as a property of the null instead of as a property of the model. A bag of notes has no order, so an order-shuffle cannot surprise it, and the two statistics cannot be scored against the same chance because they are sensitive to different things.

So a second null is needed for the profile — resampling the bars from the passage’s own vocabulary, which changes both order and content — and once there are two nulls, choosing which one to share is the same choice as choosing a weight. The commensuration problem has not been solved; it has moved one level down.

The test a histogram cannot pass. Two passages built from the same two keys: one takes them in turn and the other sounds them together. Over a window long enough to hold both, their pitch-class histograms are 97.9 per cent alike, so they are very nearly the same object to a profile model — and it names two keys in both. The ordered model names up to four for the alternation and one for the simultaneity, because a simultaneity has no sequence in it that any single key's grammar will not fit.
Fig. 3 The same blindness in its original form, from an earlier essay on key relations: two passages built from the same two keys, one alternating and one simultaneous, whose histograms are 98 per cent alike. A profile model cannot tell them apart because a histogram has no order in it. The vacuous null above is that same fact turning up in the statistics rather than in the model.

What the two kinds of evidence are actually sensitive to

Standardising makes the two comparable and the comparison is informative about the statistics rather than about the music.

Cadence evidence is sharp. On a thirty-two bar song it puts the right key more than six standard deviations above chance, with the runner-up well under two. That is a strong, well-separated verdict from a statistic that looks at a handful of ordered pairs.

Profile evidence is blunt. The same passage gives the right key a little over one standard deviation, and the runner-up below zero. It is the correct answer with much less confidence, from a statistic that looks at every note in the piece — and the ratio between the two, rather than either number, is what a reader should carry away, because a null of any given size estimates a ratio far better than it estimates a tail.

That ordering is worth pausing on, because the profile model uses far more of the data. It is blunt because a bag of notes from a piece in C is not very different from a bag of notes from a piece in G — they share six of seven pitch classes — whereas a cadence in C is quite different from a cadence in G.

So the two statistics are not two views of one thing. One is a low-information statistic over a large sample and the other is a high-information statistic over a small one, and the reason they should be combined is precisely that they fail in different circumstances: a passage with no cadences leaves the first with nothing, and a passage that stays inside one scale leaves the second unable to choose.

Five signals, computed separately, and no total. The five components of closure for 6 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.
Fig. 4 Where the same problem was refused before: the closure vector, five components of an ending measured on five different scales with no defensible total. The move this essay makes — standardise each against its own chance — is available there too, and it would run into the same difficulty, which is that a null for a harmonic component and a null for a rhythmic one are different objects.

The shape of the result, which is a general one

It is worth extracting the general form, because it applies well beyond key-finding and this collection keeps meeting it.

A null is a claim about what would be surprising. Choosing one is choosing which features of the data are the interesting ones, and a statistic is only scoreable against a null that moves it. So two statistics sensitive to different features cannot be scored against one null, and standardisation — which looks like a way of avoiding a weighting — turns out to require the same decision in a different vocabulary.

That is not an argument against standardising. It is an argument for saying which null was used, every time, and for reporting the vector alongside the total. Every figure in this essay carries both.

The same shape appears twice more on this site. The closure ladder’s vector has five components measured on five scales, and the nulls for a harmonic component and a rhythmic one are different objects. The three ways to measure how far a key is disagree with each other and there is no null that makes them commensurable either.

Three anchors, one problem, and it is not a technical one. It is that a musical claim usually rests on several kinds of evidence and there is no view from nowhere that says how much each is worth.

The pipeline and the joint search agree here. Every hypothesis the joint search considers for this passage, as a point: how well its phase explains the onsets against how well its chord explains the notes, with the key's own correlation and its agreement with the chord folded into the shading. The pipeline chooses the best phase first and is then committed — it reads C major7 in C major with the bar starting at slot 0. The joint search reads C major7 in C major at slot 0. The two searches are 176 hypotheses and 27648, a factor of 157.
Fig. 5 A fourth instance of the same thing, from an earlier essay: what counts as a chord depends on a segmentation, and the segmentation depends on the metre. Every statistic in this essay is computed over bars, and a bar is a decision. Standardising against a null takes the units out and leaves that decision exactly where it was.

Which computation produced the numbers

The passage is a thirty-two bar AABA scheme. The cadence statistic is the sixth rung’s: every adjacent pair of bars is read as roman numerals in each candidate key, and the closure components — root motion by fifth, arrival on the tonic, a leading tone resolving — are counted. The profile statistic is a pitch-class histogram correlated against the major and minor templates for each key, taking the better of the two.

The order null permutes the bars 160 times and takes the best-scoring key each time. The resampling null draws thirty-two bars at random from the roman numerals the passage actually uses, 160 times, and does the same. Each observed score is then expressed as standard deviations above the mean of the relevant null.

Every figure here now builds its nulls from a thousand draws rather than the hundred and sixty the first version used, which is the smallest change that makes the published numbers the converged ones. The cost is a second of build time per figure.

The generator is seeded, so the figure is the same drawing every time it is looked at. A null built from fresh random draws would move slightly whenever it was recomputed, and a chance that moves is not a chance anything can be scored against.

The vacuous-null result is a property of arithmetic and is not sampling noise: the profile null’s spread comes out at 1.6 × 10⁻¹⁵, which is a floating-point zero and not a small number.

How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is.
Fig. 6 What the combined verdict looks like bar by bar rather than as a scatter: the reading each statistic gives through the passage, with the standardised scores beside them. Where they agree the total is uninteresting and where they disagree the total is a weighting, which is the whole content of this essay stated as a picture.

Where the model stops

The passage is a scheme rather than a piece. Thirty-two bars of roman numerals with one chord each is a caricature of a song, and both statistics are being fed exactly the diet they were designed for: unambiguous triads at a regular rate, with no passing notes, no inversions and no ornamentation. Real music supplies all three, and every one of them adds pitch classes to the histogram without adding cadences — which would blunt the profile statistic further and leave the cadence statistic alone.

Two nulls is not a solution and is presented as one only in the narrow sense. The essay’s positive claim is that standardising can put a count and a correlation on one scale; its negative claim is that doing so requires choosing a null per statistic, and that the choice is a weighting. Both are true and the second is the larger.

The nulls are both wrong in the same way. Neither the shuffle nor the resample produces music. A permuted progression is not a plausible piece and a random draw from a vocabulary is not either, so “how surprising is this against chance” is being asked with a chance nobody would ever hear. A better null would be a generative model of tonal progressions, and this site has one — the root-motion weights of the ordered key-finder — which would make the null musical and would import that model’s own eight asserted parameters into this one.

The number of shuffles bounds what can be claimed, and by more than was first supposed — the section above raises it and the scores move. What survives is the comparison between the two statistics rather than either absolute value, and comparing two passages whose scores differ by less than half a standard deviation still needs a larger null than any figure here uses.

The two statistics were assumed to be dependent and are not. Computed from the same bars, they might have been expected to rise and fall together; measured under the null they correlate at 0.012. That is the one caveat on this list which the arithmetic simply removes.

Two numbers the caveats promised and did not give

Two of the limitations listed below are checkable in the same run that produced the figures, and both were left as assertions. Checked, one is worse than stated and one is not a problem at all.

The shuffle count was too small, and the error bar given for it was too narrow. The scores above are computed from a hundred and sixty draws, with the claim that a hundred and sixty resolves a standard deviation to about six per cent and a score of 5.6 to about ±0.4. Raising it says otherwise:

draws best cadence score best profile score
160 5.53 1.40
1,000 6.19 1.13
5,000 6.20 1.22

Both move by more than their own stated bound, and the cadence score does not wobble around a value — it converges upward, settling by a thousand draws and unchanged at five thousand. So a hundred and sixty draws was not merely noisy; it was biased, because the quantity being estimated is the mean and spread of a maximum over twelve keys and a small sample under-estimates the spread of a maximum. The right figures are 6.2 and 1.2, and the qualitative finding — the cadence statistic is several times sharper — is untouched and slightly strengthened.

The two statistics are, as near as makes no difference, independent. The last caveat worries that adding two standardised scores computed from the same bars double-counts what they share, and says the joint distribution would settle it. Over five thousand resample draws the correlation between the two statistics is r = 0.012.

That makes the sum’s variance 2.024 rather than 2.000, so the combined score is inflated by six parts in a thousand — smaller than the sampling error on any number in this essay, and far smaller than the weighting decision the whole rung is about. The double-counting is not a real objection.

And the reason it is not is the same fact this rung keeps arriving at from different directions. The two statistics decorrelate because one of them is order-blind. Under a null that randomises order and content together, the cadence count moves with the order and the histogram does not care about it, so the two vary almost independently even though they are computed from the same thirty-two bars. The property that makes one null vacuous is the property that makes the pair safe to add.

There is a third number in the same position — computed by the machinery already, reported nowhere — and it is worth printing because it says how much of the confidence is real.

How sure the reading is, bar by bar. Every earlier essay reports one best reading. A dynamic program that finds a best path has, by construction, the best score into every state at every bar — so the gap between the best reading and the best reading in any other key is already computed and has never been printed. Here it is, in bits, for the thirty-two-bar AABA. The mean margin is 1.14 bits and 13 of 32 bars are inside one bit of a rival reading, which is where a listener would be genuinely undecided. The reading itself names C, E, B, D, G; the margin says what that naming is worth, and at the weakest bar — bar 31, C over F — it is worth 0.07.
Fig. 7 The gap between the best reading and the best reading in any other key, bar by bar, in bits. A dynamic programme that finds a best path has the best score into every state at every bar by construction, so this quantity has been available all along and has never been drawn. It is not flat: the reading is decided in some bars and nearly a coin toss in others.

That the margin varies at all is the point. A single reported reading invites the assumption that a key-finder is equally sure everywhere, and it is not — which is the same lesson as the one above, arriving from the confidence side rather than the correlation side.

Whose music, and when

The two statistics belong to different traditions and the difference is not accidental.

The profile model comes from experimental psychology: the templates are ratings collected from listeners, and the statistic asks how well a passage’s note distribution matches what listeners say belongs in a key. It is a model of a perceptual state.

The cadence model comes from music theory: it counts the events an analyst would circle, and it asks whether the syntactic markers of a key are present. It is a model of a grammatical claim.

Those are not two estimators of one quantity, and the finding that they agree on a thirty-two bar song with very different confidence is what one would expect of two different questions with a common answer. Where they would disagree is exactly where the interesting music is: a passage that stays inside one scale and cadences somewhere else, which is what a great deal of modal writing after 1950 does deliberately, and which the profile statistic would call one key and the cadence statistic another.

That case is the one this rung has not run, and it is the one where a combined verdict would either earn its keep or be exposed.

What the picture cannot show

Whether a listener does either of these. The site has a measurement of how many chords a modulation takes to detect and nothing at all about which evidence a listener is using. Two statistics that agree on the answer say nothing about which one is the mechanism.

And it cannot show a passage with no key. Both statistics always return a best key, because both are maximisations. A passage that is genuinely atonal gets a winner with a low score and the score is what says so — which means the standardised value is doing double duty as a verdict and as a confidence, and nothing here separates “the key is C” from “there is a key and it is C”. A statistic that never returns no key cannot be used to find one, which is a limitation every model in this ladder shares.

Where this ladder goes next

Seven rungs. A progression is a path across a space with a geometry; it never comes home in just intonation; its rate has limits set outside music; the objects it connects are produced by a segmentation that depends on the metre; three decisions change two readings in five; the ordered pairs carry key evidence a bag cannot; and now the two kinds of evidence on one scale, with the commensuration problem moved rather than removed.

The rung after it is the case the last section names, and the passage built to make them disagree is where it goes: every figure here is run on a passage where the two statistics agree, which is the uninteresting case, because agreement between a low-information statistic and a high-information one tells nobody anything. What that rung produces is not a better key-finder but a measurement of how much the two kinds of evidence are worth against each other — the number the weighting this essay avoided would have needed.

Part 7 of 17

One essay in the series on progression. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CadenceEvidenceHarmonic analysisKey-findingProgressionSegmentationStatisticsTonal hierarchy