Harmony and voice leading

A chord, given a key and a predecessor

A chord's improbability has been priced from its root motion alone, which left one multiplication unmade: a chord is also improbable because its notes do not fit the key, and that number has been available since the probe-tone profile. Multiplied and renormalised, the two give a conditional distribution — and a distribution has an entropy, which is the quantity a surprise has to be read against and which a list of preferences cannot supply.

Assumes: Surprise is a number · Counting produced the hierarchy

Surprise is a number turned the chord that did not come from a description into a quantity: the information content of the chord that did arrive, minus the log of its probability, under this site’s own root-motion weights. It ended by naming the multiplication it had not done.

A chord’s improbability has two independent sources. Its root motion is one — a descending fifth is expected and a step down is not — and that is the third rung’s. How well its notes fit the key is the other, and this collection has had a measurement of it since counting produced the hierarchy: the probe-tone profile, which is how strongly a key specifies each of the twelve.

A listener presumably uses both. Multiplying them and renormalising over the seven destinations gives a conditional distribution — p of a chord given a key and a predecessor — which is a standard object, which this collection has both factors of, and which has one property the third rung’s model could not have.

The surprise of each chord, against the uncertainty it arrived into. The information content of each step — minus the log of its probability under a distribution that multiplies the root-motion weight by how well the destination triad's notes fit the key — with the entropy of the moment before it drawn behind. I – IV – V – I: I→IV 1.71 bits, IV→V 2.80 bits, V→I 1.60 bits, against a mean uncertainty of 2.63; I – IV – V – vi: I→IV 1.71 bits, IV→V 2.80 bits, V→vi 2.61 bits, against a mean uncertainty of 2.63. A surprise larger than the entropy it arrived into is an outcome the model was not expecting even given how uncertain it was; one below it is an outcome the model had already mostly bet on. An earlier essay produced the first of those numbers and had no way to produce the second, because a set of preferences is not a distribution and only a distribution has an entropy.
Fig. 1 An authentic cadence and a deceptive one, with each step’s surprise in bits and — the grey bars behind — the entropy of the moment it arrived into. The two progressions are identical until the last chord, where the authentic resolution costs 1.60 bits and the deceptive one 2.61. The difference is a shade over one bit, which is a doubling of improbability and is as clean a number as this subject produces.

What a distribution has that a preference list does not

The third rung’s model is a set of weights over root motions, normalised to give a probability for the chord that arrived. That is enough for a surprise and it is not enough for anything else, because the weights depend only on the interval and not on where in the key it lands.

Once the profile term is in, the distribution is different at every degree. From the tonic, a descending fifth reaches IV and a step reaches ii; both are strongly specified by the key, so the distribution is fairly flat and the listener is fairly uncertain. From the dominant, a descending fifth reaches the tonic — the most strongly specified triad there is — so the distribution is sharply peaked and the listener is fairly certain.

A distribution has an entropy, and the entropy is the thing a surprise has to be read against. Two bits of surprise arriving into a moment of two bits of uncertainty is an ordinary outcome; two bits arriving into a moment of half a bit is a shock. The third rung could report only the first number.

step        surprise   entropy of the moment
I → IV        1.71          2.62
IV → V        2.80          2.69
V → I         1.60          2.57
V → vi        2.61          2.57

The deceptive resolution’s 2.61 bits arrives into 2.57 bits of uncertainty — it uses almost exactly the uncertainty available. The authentic resolution’s 1.60 uses less than two thirds of it.

What the deceptive cadence costs, in one number

The gap is 1.01 bits, and it is worth being careful about what that means.

One bit is a factor of two in probability. Under this model, arriving at vi instead of I after a dominant is half as likely — not a tenth as likely, not a hundredth. That is a smaller number than the rhetoric around deceptive cadences suggests, and it is the right size: a deceptive cadence is a common device, appears in almost every large tonal form, and is recognised rather than found bewildering. A device that cost five bits would be one nobody could use twice.

It is also the same conclusion the chord that did not come reached in prose. That essay described the moment and could not price it; this prices it and gets a number consistent with the description, which is the direction of agreement worth having.

How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; I – IV – V – vi totals 7.3 bits over 3 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire.
Fig. 2 The earlier figure for the same progression: the surprise of each step under root motion alone, with no key term. The shape is the same and the numbers are not, which is what the multiplication does — and there is nothing behind these bars, because a set of preferences has no entropy.

Pricing the second factor

Any multiplication of two sources needs a weight, and inventing one is what this collection has been at some pains to learn not to do. So the profile term carries an exponent: at zero the model is the third rung’s exactly, and at one the two sources are weighted equally in log space.

Sweeping it from zero to three moves the numbers and does not move the ranking. The most surprising moment of the progression is IV→V at every exponent; the second is the final resolution; the third is I→IV. The ordering is stable across a twelve-fold change in how much the profile term is worth.

That is the fifth time this collection has priced an asserted parameter and the fourth time it has come back with the same verdict: the ordering is a result and the magnitude is an illustration. The pattern is consistent enough now to be worth stating as a working expectation rather than as a finding — a model built out of ordinal weights produces conclusions about order, and the numbers on the axis are there to be read against each other.

What the profile factor is worth, priced. The conditional model for I – IV – V – vi with the profile term raised to a stated power: at zero it is the earlier root-motion model exactly, and at one the two sources of improbability are weighted equally in log space. The mean surprise falls from 2.42 bits to 2.36 and the entropy of the moment falls from 2.65 to 2.51, so a stronger profile term makes the listener more certain and the arriving chord less surprising at once. The ranking of the moments does not change at any exponent drawn — it is moment 2, then moment 3, then moment 1 throughout — which is the same verdict the root-motion weights got when they were priced: the ordering is a result and the magnitude is an illustration.
Fig. 3 The profile term priced. Three quantities against the exponent: the entropy of the moment, the largest surprise in the progression, and the mean. All three move and the ranking of the moments does not, so the second factor changes how confident the model is and not what it is confident about.

The two factors pull in opposite directions

There is a structural feature of the multiplication that is not obvious in advance and shows up in the sweep.

A stronger profile term makes the listener more certain — the entropy of every moment falls, because the distribution concentrates on the well-specified degrees. It also makes the arriving chord less surprising on average, because the chords that actually arrive in a tonal progression are the well-specified ones.

So both curves fall together and the ratio between them is nearly flat. A model with a heavy profile term is a model of a listener who knows the style well and is therefore both more confident and less often startled, which is the right direction for a parameter that stands for stylistic knowledge — and it means the exponent cannot be fitted from surprise data alone, because raising it moves the prediction and the yardstick together.

The circle of fifths, which the model should find dull

A model of expectation earns its keep by being unsurprised where a listener is unsurprised, so the useful control is a passage everybody finds predictable.

The circle of fifths is that passage. Seven chords, each a descending fifth from the last, and after two of them nobody is in any doubt about the third. Under the conditional model every step is a descending fifth — the heaviest weight in the root-motion table — and the profile term varies over the seven degrees, so the surprise is low and uneven rather than low and flat.

The unevenness is the interesting part. The steps that land on well-specified degrees are the cheapest, and the step that lands on the diminished triad on the seventh degree is dearer than the rest even though its root motion is identical. That is the profile term doing exactly what it was added for: distinguishing two moves that the third rung’s model had to call the same.

The surprise of each chord, against the uncertainty it arrived into. The information content of each step — minus the log of its probability under a distribution that multiplies the root-motion weight by how well the destination triad's notes fit the key — with the entropy of the moment before it drawn behind. the circle of fifths: I→IV 1.71 bits, IV→vii° 2.09 bits, vii°→iii 1.91 bits, iii→vi 1.73 bits, vi→ii 2.01 bits, ii→V 1.99 bits, V→I 1.60 bits, against a mean uncertainty of 2.63. A surprise larger than the entropy it arrived into is an outcome the model was not expecting even given how uncertain it was; one below it is an outcome the model had already mostly bet on. An earlier essay produced the first of those numbers and had no way to produce the second, because a set of preferences is not a distribution and only a distribution has an entropy.
Fig. 4 The sequence every model of harmony has to find dull. Every step is the same root motion and the surprises are not equal, because the destinations are not equally at home in the key. The entropy bars behind them are lowest where the arriving degree is most strongly specified.

What a bit is worth here

It is worth pinning down the units, because “bits” invites an accuracy the model does not have.

The alphabet is seven diatonic triads. A listener with no expectation at all — a uniform distribution over the seven — has an entropy of 2.81 bits and every chord costs exactly that. The model’s mean entropy over an ordinary progression is about 2.6, so the whole apparatus of root motion and key fit buys about two tenths of a bit against knowing nothing.

That is a small number and it is honest. Seven diatonic triads is already an enormously constrained alphabet — the constraint of being in the key at all is most of what a listener knows, and it has been applied before the model starts. What the model adds is the ordering within that, and the ordering is what the figures are about.

It also puts the deceptive cadence’s one bit in proportion: a whole bit, against a total budget of two point eight, is a third of everything a listener could possibly be uncertain about at that moment. Read that way the device is large rather than small, and both readings are the same number.

The twenty-four triads, and the three moves between them. Every major and minor triad, placed on the single cycle that alternating two of the three transformations produces. Neighbours round the ring differ by one voice; the chords across the middle are the third transformation. Every edge's cost is the measured voice-leading distance, and none of them is more than two semitones.
Fig. 5 The alphabet, as a space, with the key’s own triads marked: six of the seven are consonant and appear here, and the diminished vii° is not a node on this graph at all. The moves between them are the edges. Every distribution in this essay is over the seven nodes of this graph, and the entropy of a moment is how spread the model’s bet is across them.

Which computation produced the numbers

The root-motion factor is ROOT_MOTION, the ordered key-finder’s own table, with the fix the essay that first read it as a distribution found in it: seven entries rather than eight, with the fourth-or-fifth class carrying the weight its comment always meant it to. That defect is worth remembering here because it was found by exactly this kind of reuse — a table read as a distribution rather than as a set of preferences.

The profile factor is the mean probe-tone rating of the destination triad’s three tones, normalised to the profile’s own maximum. Triads are built by degreeTriad on the diatonic set, so the term is a property of the degree rather than of a voicing.

The two are multiplied, the seven products are normalised to sum to one, and the surprise of the chord that arrived is minus the log of its share. The entropy is the usual sum over the seven.

Nothing here is fitted. The root-motion weights are ordinal and asserted, the profile is a measurement on listeners, and the exponent joining them is the thing the sweep prices.

What each factor buys, separately

The section above prices the exponent joining the two factors and never asks the prior question, which is how much either of them is worth on its own. Running each alone over the seven starting degrees, in mean entropy against a listener who knows only that the chord will be one of seven:

model mean entropy bits bought
nothing at all 2.807
root motion alone 2.647 0.161
key fit alone 2.793 0.015
both, multiplied 2.634 0.174

The second factor’s marginal contribution is 0.013 bits, seven and a half per cent of a total that was already small. The reason is in the numbers rather than in the idea: the fit factor runs from 0.549 on the diminished triad to 0.836 on the tonic, a range of 1.52 to one, while the motion weights run from 0.15 to 1.0, a range of 6.7 to one. The measurement on listeners is a much flatter function than the table of preferences it is multiplied by.

That looks damning and is not, because entropy is the wrong currency for what the term was added to do. From the dominant, root motion alone ranks the seven destinations I, vii, vi, iii, IV, ii, V; with the fit term at an exponent of one the ranking is I, vi, IV, vii, iii, ii, V. The diminished triad on the seventh degree falls from second place to fourth, and it falls for the reason the circle-of-fifths section gives: its root motion from V is a step and the step class is shared, so only a term that knows about the destination can separate it from the moves that share its interval. The factor is worth almost nothing in confidence and a great deal in order, and those are separate ledgers.

It also disposes of the independence worry the closing caveats raise. Measured over the forty-nine ordered pairs, the correlation between the two factors in log space is exactly zero — structurally, because the motion term is a function of the interval alone and the fit term of the destination alone, and across all pairs each destination is reached by each interval once. Over the twelve moves an ordinary tonal progression actually makes it is −0.09, very slightly negative. There is no shared component being double-counted, and the multiplication is doing what the section below says it does.

Why the multiplication and not a weighted sum

The two sources could have been combined the other way, and it is worth saying why they are not.

A weighted sum of two scores needs a weight in the units of the two things being added, which is the commensuration problem the progression ladder has met twice and this collection has now recorded as a standing hazard. A product of two probabilities needs no such weight: it is what independence means, and the normalisation that follows removes any common scale factor.

So the exponent this rung sweeps is not a weight in a sum. It is a statement about how much evidence the profile term carries relative to the motion term — a temperature rather than a mixing coefficient — and at one the two are simply multiplied, which is the assumption of independence stated plainly and then flagged as false in the section below.

That is the same move the orchestration anchor made for a different pair of quantities, and for the same reason: a problem that looks like it needs an exchange rate often does not, if it can be restated so that one of the two is a constraint or a factor rather than a term.

The probe-tone profile, major against minor. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.
Fig. 6 The second factor, which is a measurement rather than a model: how strongly a key specifies each of the twelve. Every triad’s fit in this essay is the mean of this curve at its three tones, and it is the only ingredient in the whole computation that came from listeners rather than from a table of preferences.

Where the model stops

Seven destinations is a very small alphabet. A real listener’s expectation is over chords with inversions, sevenths, applied dominants and chromatic alterations, and over modulations to other keys. Seven diatonic triads is the alphabet the third rung chose and this rung inherits, and every entropy quoted is an entropy over that alphabet — so the absolute bits are not comparable with any published figure computed over a larger one.

The key is given. Both factors need to know what the key is, and a listener does not have one handed to them; they infer it, with a delay and a margin, which is what the key-relations ladder is about. A model that inferred the key would have a third source of uncertainty and would produce larger entropies everywhere.

The two factors are treated as independent, and the section on what each buys shows they nearly are. The right object is still a joint distribution measured over a corpus, which is one of the counts this collection keeps recording that it does not have — but the double-counting this caveat used to warn about is not there to be found.

What the picture cannot show

It cannot show a listener being surprised. Information content is a property of a model, and every psychological claim attached to it is an additional hypothesis — that reaction time scales with bits, that a physiological response does, that a report of surprise does. Those are published claims with published evidence and none of it is here.

Nor can it show learning. The whole point of an expectation model is that expectations come from exposure, and both factors here are fixed. A listener hearing their thousandth deceptive cadence has a different distribution from one hearing their first, and nothing in the model has a way to move.

And it cannot show where in the bar the chord arrives. Which notes are the chord established that segmentation depends on the metre, and a chord arriving on a strong beat is a different event from the same chord arriving on a weak one. The model has an alphabet and an order and no clock.

Whose harmony, and when

Both factors are claims about the European common practice, roughly 1650 to 1900, and both were arrived at from it. The root-motion weights encode the descending-fifth preference that is the defining regularity of that repertoire; the probe-tone profile was measured on listeners raised in it.

So the model’s verdict on a deceptive cadence is a verdict about a device in the style the model is of, and that is not circular so much as narrow. A deceptive cadence in a repertoire where vi is not a weakly specified degree would cost fewer bits under this model and would presumably be a weaker device. That is a testable consequence, and one half of it can be tested here rather than promised, because the fit term is a measured profile and there is a second measured profile.

The surprise of each chord, against the uncertainty it arrived into. The information content of each step — minus the log of its probability under a distribution that multiplies the root-motion weight by how well the destination triad's notes fit the key — with the entropy of the moment before it drawn behind. I – IV – V – I: I→IV 1.70 bits, IV→V 2.66 bits, V→I 1.63 bits, against a mean uncertainty of 2.63; I – IV – V – vi: I→IV 1.70 bits, IV→V 2.66 bits, V→vi 2.72 bits, against a mean uncertainty of 2.63. A surprise larger than the entropy it arrived into is an outcome the model was not expecting even given how uncertain it was; one below it is an outcome the model had already mostly bet on. An earlier essay produced the first of those numbers and had no way to produce the second, because a set of preferences is not a distribution and only a distribution has an entropy.
Fig. 7 The same two progressions, the same root-motion table, and the same seven destinations — priced against the minor probe profile instead of the major one. Only the measurement of key strength has changed. The deceptive step costs 2.72 bits against the major profile’s 2.61, and the authentic one 1.63 against 1.60.

Both steps cost more, which looks at first like nothing but a flatter profile, and the ratio is where the content is. What prices V→vi is not how well vi fits but how well it fits relative to the alternatives, and against the minor profile the submediant’s fit falls from 0.755 to 0.612 while the tonic’s falls only from 0.836 to 0.720 — from 90 per cent of the tonic’s to 85. The submediant is more weakly specified under the second profile, not less.

So this is the prediction run in the direction the data actually offers, and it comes out the right way: where vi is more weakly specified, the deceptive cadence costs more. It is the contrapositive rather than the test that was asked for — no measured profile on this site makes vi well specified, so the antecedent of the original claim cannot be met with what is here, and that half stays owed to a corpus.

Where this ladder goes next

Four rungs. Counting produced the hierarchy; a chord did not come and the moment was described; the moment acquired a number; and now the number has an uncertainty behind it to be read against.

What is owed after this is the time axis. Every quantity here is attached to a chord change, and a listener’s expectation is a continuous thing that sharpens as a bar proceeds and collapses when the change arrives. The progression ladder has a model of how often the chord changes and the metre ladder has a model of where the strong beats are, so the ingredients for an expectation that rises and falls within a bar are here — and what would come out is not a list of surprises but a curve, which is the shape every published account of musical expectation is drawn as and this collection has never computed.

Part 4 of 11

One essay in the series on Tonal-expectation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CadenceExpectationInformationKey-findingProbe-toneProgressionTonal functionTonal hierarchy