A chord, given a key and a predecessor
Assumes: Surprise is a number · Counting produced the hierarchy
Surprise is a number turned the chord that did not come from a description into a quantity: the information content of the chord that did arrive, minus the log of its probability, under this site’s own root-motion weights. It ended by naming the multiplication it had not done.
A chord’s improbability has two independent sources. Its root motion is one — a descending fifth is expected and a step down is not — and that is the third rung’s. How well its notes fit the key is the other, and this collection has had a measurement of it since counting produced the hierarchy: the probe-tone profile, which is how strongly a key specifies each of the twelve.
A listener presumably uses both. Multiplying them and renormalising over the seven destinations gives a conditional distribution — p of a chord given a key and a predecessor — which is a standard object, which this collection has both factors of, and which has one property the third rung’s model could not have.
What a distribution has that a preference list does not
The third rung’s model is a set of weights over root motions, normalised to give a probability for the chord that arrived. That is enough for a surprise and it is not enough for anything else, because the weights depend only on the interval and not on where in the key it lands.
Once the profile term is in, the distribution is different at every degree. From the tonic, a descending fifth reaches IV and a step reaches ii; both are strongly specified by the key, so the distribution is fairly flat and the listener is fairly uncertain. From the dominant, a descending fifth reaches the tonic — the most strongly specified triad there is — so the distribution is sharply peaked and the listener is fairly certain.
A distribution has an entropy, and the entropy is the thing a surprise has to be read against. Two bits of surprise arriving into a moment of two bits of uncertainty is an ordinary outcome; two bits arriving into a moment of half a bit is a shock. The third rung could report only the first number.
step surprise entropy of the moment
I → IV 1.71 2.62
IV → V 2.80 2.69
V → I 1.60 2.57
V → vi 2.61 2.57
The deceptive resolution’s 2.61 bits arrives into 2.57 bits of uncertainty — it uses almost exactly the uncertainty available. The authentic resolution’s 1.60 uses less than two thirds of it.
What the deceptive cadence costs, in one number
The gap is 1.01 bits, and it is worth being careful about what that means.
One bit is a factor of two in probability. Under this model, arriving at vi instead of I after a dominant is half as likely — not a tenth as likely, not a hundredth. That is a smaller number than the rhetoric around deceptive cadences suggests, and it is the right size: a deceptive cadence is a common device, appears in almost every large tonal form, and is recognised rather than found bewildering. A device that cost five bits would be one nobody could use twice.
It is also the same conclusion the chord that did not come reached in prose. That essay described the moment and could not price it; this prices it and gets a number consistent with the description, which is the direction of agreement worth having.
Pricing the second factor
Any multiplication of two sources needs a weight, and inventing one is what this collection has been at some pains to learn not to do. So the profile term carries an exponent: at zero the model is the third rung’s exactly, and at one the two sources are weighted equally in log space.
Sweeping it from zero to three moves the numbers and does not move the ranking. The most surprising moment of the progression is IV→V at every exponent; the second is the final resolution; the third is I→IV. The ordering is stable across a twelve-fold change in how much the profile term is worth.
That is the fifth time this collection has priced an asserted parameter and the fourth time it has come back with the same verdict: the ordering is a result and the magnitude is an illustration. The pattern is consistent enough now to be worth stating as a working expectation rather than as a finding — a model built out of ordinal weights produces conclusions about order, and the numbers on the axis are there to be read against each other.
The two factors pull in opposite directions
There is a structural feature of the multiplication that is not obvious in advance and shows up in the sweep.
A stronger profile term makes the listener more certain — the entropy of every moment falls, because the distribution concentrates on the well-specified degrees. It also makes the arriving chord less surprising on average, because the chords that actually arrive in a tonal progression are the well-specified ones.
So both curves fall together and the ratio between them is nearly flat. A model with a heavy profile term is a model of a listener who knows the style well and is therefore both more confident and less often startled, which is the right direction for a parameter that stands for stylistic knowledge — and it means the exponent cannot be fitted from surprise data alone, because raising it moves the prediction and the yardstick together.
The circle of fifths, which the model should find dull
A model of expectation earns its keep by being unsurprised where a listener is unsurprised, so the useful control is a passage everybody finds predictable.
The circle of fifths is that passage. Seven chords, each a descending fifth from the last, and after two of them nobody is in any doubt about the third. Under the conditional model every step is a descending fifth — the heaviest weight in the root-motion table — and the profile term varies over the seven degrees, so the surprise is low and uneven rather than low and flat.
The unevenness is the interesting part. The steps that land on well-specified degrees are the cheapest, and the step that lands on the diminished triad on the seventh degree is dearer than the rest even though its root motion is identical. That is the profile term doing exactly what it was added for: distinguishing two moves that the third rung’s model had to call the same.
What a bit is worth here
It is worth pinning down the units, because “bits” invites an accuracy the model does not have.
The alphabet is seven diatonic triads. A listener with no expectation at all — a uniform distribution over the seven — has an entropy of 2.81 bits and every chord costs exactly that. The model’s mean entropy over an ordinary progression is about 2.6, so the whole apparatus of root motion and key fit buys about two tenths of a bit against knowing nothing.
That is a small number and it is honest. Seven diatonic triads is already an enormously constrained alphabet — the constraint of being in the key at all is most of what a listener knows, and it has been applied before the model starts. What the model adds is the ordering within that, and the ordering is what the figures are about.
It also puts the deceptive cadence’s one bit in proportion: a whole bit, against a total budget of two point eight, is a third of everything a listener could possibly be uncertain about at that moment. Read that way the device is large rather than small, and both readings are the same number.
Which computation produced the numbers
The root-motion factor is ROOT_MOTION, the ordered key-finder’s own table, with the fix the essay that first read it as a distribution found in it: seven entries rather than eight, with the fourth-or-fifth class carrying the weight its comment always meant it to. That defect is worth remembering here because it was found by exactly this kind of reuse — a table read as a distribution rather than as a set of preferences.
The profile factor is the mean probe-tone rating of the destination triad’s three tones, normalised to the profile’s own maximum. Triads are built by degreeTriad on the diatonic set, so the term is a property of the degree rather than of a voicing.
The two are multiplied, the seven products are normalised to sum to one, and the surprise of the chord that arrived is minus the log of its share. The entropy is the usual sum over the seven.
Nothing here is fitted. The root-motion weights are ordinal and asserted, the profile is a measurement on listeners, and the exponent joining them is the thing the sweep prices.
What each factor buys, separately
The section above prices the exponent joining the two factors and never asks the prior question, which is how much either of them is worth on its own. Running each alone over the seven starting degrees, in mean entropy against a listener who knows only that the chord will be one of seven:
| model | mean entropy | bits bought |
|---|---|---|
| nothing at all | 2.807 | — |
| root motion alone | 2.647 | 0.161 |
| key fit alone | 2.793 | 0.015 |
| both, multiplied | 2.634 | 0.174 |
The second factor’s marginal contribution is 0.013 bits, seven and a half per cent of a total that was already small. The reason is in the numbers rather than in the idea: the fit factor runs from 0.549 on the diminished triad to 0.836 on the tonic, a range of 1.52 to one, while the motion weights run from 0.15 to 1.0, a range of 6.7 to one. The measurement on listeners is a much flatter function than the table of preferences it is multiplied by.
That looks damning and is not, because entropy is the wrong currency for what the term was added to do. From the dominant, root motion alone ranks the seven destinations I, vii, vi, iii, IV, ii, V; with the fit term at an exponent of one the ranking is I, vi, IV, vii, iii, ii, V. The diminished triad on the seventh degree falls from second place to fourth, and it falls for the reason the circle-of-fifths section gives: its root motion from V is a step and the step class is shared, so only a term that knows about the destination can separate it from the moves that share its interval. The factor is worth almost nothing in confidence and a great deal in order, and those are separate ledgers.
It also disposes of the independence worry the closing caveats raise. Measured over the forty-nine ordered pairs, the correlation between the two factors in log space is exactly zero — structurally, because the motion term is a function of the interval alone and the fit term of the destination alone, and across all pairs each destination is reached by each interval once. Over the twelve moves an ordinary tonal progression actually makes it is −0.09, very slightly negative. There is no shared component being double-counted, and the multiplication is doing what the section below says it does.
Why the multiplication and not a weighted sum
The two sources could have been combined the other way, and it is worth saying why they are not.
A weighted sum of two scores needs a weight in the units of the two things being added, which is the commensuration problem the progression ladder has met twice and this collection has now recorded as a standing hazard. A product of two probabilities needs no such weight: it is what independence means, and the normalisation that follows removes any common scale factor.
So the exponent this rung sweeps is not a weight in a sum. It is a statement about how much evidence the profile term carries relative to the motion term — a temperature rather than a mixing coefficient — and at one the two are simply multiplied, which is the assumption of independence stated plainly and then flagged as false in the section below.
That is the same move the orchestration anchor made for a different pair of quantities, and for the same reason: a problem that looks like it needs an exchange rate often does not, if it can be restated so that one of the two is a constraint or a factor rather than a term.
Where the model stops
Seven destinations is a very small alphabet. A real listener’s expectation is over chords with inversions, sevenths, applied dominants and chromatic alterations, and over modulations to other keys. Seven diatonic triads is the alphabet the third rung chose and this rung inherits, and every entropy quoted is an entropy over that alphabet — so the absolute bits are not comparable with any published figure computed over a larger one.
The key is given. Both factors need to know what the key is, and a listener does not have one handed to them; they infer it, with a delay and a margin, which is what the key-relations ladder is about. A model that inferred the key would have a third source of uncertainty and would produce larger entropies everywhere.
The two factors are treated as independent, and the section on what each buys shows they nearly are. The right object is still a joint distribution measured over a corpus, which is one of the counts this collection keeps recording that it does not have — but the double-counting this caveat used to warn about is not there to be found.
What the picture cannot show
It cannot show a listener being surprised. Information content is a property of a model, and every psychological claim attached to it is an additional hypothesis — that reaction time scales with bits, that a physiological response does, that a report of surprise does. Those are published claims with published evidence and none of it is here.
Nor can it show learning. The whole point of an expectation model is that expectations come from exposure, and both factors here are fixed. A listener hearing their thousandth deceptive cadence has a different distribution from one hearing their first, and nothing in the model has a way to move.
And it cannot show where in the bar the chord arrives. Which notes are the chord established that segmentation depends on the metre, and a chord arriving on a strong beat is a different event from the same chord arriving on a weak one. The model has an alphabet and an order and no clock.
Whose harmony, and when
Both factors are claims about the European common practice, roughly 1650 to 1900, and both were arrived at from it. The root-motion weights encode the descending-fifth preference that is the defining regularity of that repertoire; the probe-tone profile was measured on listeners raised in it.
So the model’s verdict on a deceptive cadence is a verdict about a device in the style the model is of, and that is not circular so much as narrow. A deceptive cadence in a repertoire where vi is not a weakly specified degree would cost fewer bits under this model and would presumably be a weaker device. That is a testable consequence, and one half of it can be tested here rather than promised, because the fit term is a measured profile and there is a second measured profile.
Both steps cost more, which looks at first like nothing but a flatter profile, and the ratio is where the content is. What prices V→vi is not how well vi fits but how well it fits relative to the alternatives, and against the minor profile the submediant’s fit falls from 0.755 to 0.612 while the tonic’s falls only from 0.836 to 0.720 — from 90 per cent of the tonic’s to 85. The submediant is more weakly specified under the second profile, not less.
So this is the prediction run in the direction the data actually offers, and it comes out the right way: where vi is more weakly specified, the deceptive cadence costs more. It is the contrapositive rather than the test that was asked for — no measured profile on this site makes vi well specified, so the antecedent of the original claim cannot be met with what is here, and that half stays owed to a corpus.
Where this ladder goes next
Four rungs. Counting produced the hierarchy; a chord did not come and the moment was described; the moment acquired a number; and now the number has an uncertainty behind it to be read against.
What is owed after this is the time axis. Every quantity here is attached to a chord change, and a listener’s expectation is a continuous thing that sharpens as a bar proceeds and collapses when the change arrives. The progression ladder has a model of how often the chord changes and the metre ladder has a model of where the strong beats are, so the ingredients for an expectation that rises and falls within a bar are here — and what would come out is not a list of surprises but a curve, which is the shape every published account of musical expectation is drawn as and this collection has never computed.
Part 4 of 11
One essay in the series on Tonal-expectation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
CadenceExpectationInformationKey-findingProbe-toneProgressionTonal functionTonal hierarchy
- The passage built to make them disagree cadence, key-finding, probe-tone, progression, tonal function
- The cadence as evidence cadence, key-finding, progression, tonal function
- What a tonic costs in seconds expectation, key-finding, probe-tone, tonal hierarchy
- A progression is a path, and the map can be drawn cadence, progression, tonal function
- A short note is heard more in tune than it is expectation, probe-tone, tonal hierarchy
- How much of the reading arrives late expectation, information, key-finding