Surprise is a number
Assumes: The chord that did not come · Counting produced the hierarchy
The second rung of this ladder is about a deceptive cadence: a dominant that resolves somewhere other than the tonic, and what a listener does with the arrival that was prepared and did not come. It describes the moment carefully and it does not put a number on it, because it had none to put.
There is a standard number for exactly this, and it is one line. If a listener expects the next chord with probability p and something else arrives, the information content of what arrived is
— the standard measure of surprise, in the sense that a coin flip is one bit and an event with probability a thousandth is about ten. What it needs is p, and this site has had a model of p since the key ladder’s eighth rung: eight weights on root motion by each number of scale steps, used there to stop a key-finder re-analysing every bar.
What the weights are, and what they mean here
ROOT_MOTION is a weight for each number of scale steps a root can move by. Motion by a fourth or a fifth is the largest at 1.0, ascending seconds are 0.6, static motion is the smallest at 0.15, and the rest sit between.
The eighth rung of the key ladder used them as a transition score — a term in a dynamic program deciding which key a passage is in — and said what they are: “ordinal rather than measured, and the model’s verdicts are checked against changing them.”
Used as an expectation model they are a probability distribution over what comes next, normalised over the seven possible destinations. That is a small change of interpretation and a large change of what can be asked: a set of scores can be compared, and a distribution has an entropy, a most likely outcome, and an information content for every outcome.
So this rung is not a new model. It is the same eight numbers read as a probability instead of as a preference, and everything below is what that reading makes available.
Per step that is 1.8 bits against 2.1, so the sequence every model of harmony is supposed to find dull is not much duller than a cadence. The weights are a graph on the seven degrees with the descending fifth as its heaviest edge, and walking that edge repeatedly is exactly what a model with those weights should find least surprising — and it very nearly does not.
The eighth weight, which nothing had ever read
Doing the arithmetic turned up a defect in the table, and it is worth reporting before the numbers that rest on it.
There are seven possible destinations from a degree and the table had eight entries, because it named both -4 — a descending fifth, V to I — and 3, an ascending fourth, I to IV. Those are the same root motion: the roots are a fourth or a fifth apart and which one it is depends on the register, which a sequence of scale degrees does not carry.
Every reader of the table folds a step into a residue, and the fold sends −4 and +3 to the same place. So ROOT_MOTION["-4"] was never read by anything. Its value is 1.0, it is the largest number in the table, its own comment describes it as the strongest motion in the common practice — and every V–I the ordered key-finder has ever scored was scored at 0.4, the weight of an ascending third.
That is now fixed, with the fourth-or-fifth class carrying the weight the comment always meant. What it changes elsewhere is small and is worth stating: the eighth rung’s headline — that an ordered model tells an alternation from a simultaneity where a histogram cannot — is unaffected, because that comparison never depended on the weight’s size. What moves is the threshold the previous essay swept, which rises.
The deceptive cadence, priced
A perfect cadence is V to I, a motion by a fourth or a fifth, which is the weights’ largest value. Under the normalised distribution its probability is 0.28 and its surprise is 1.85 bits.
A deceptive cadence is V to vi, an ascending second, weight 0.55 — a perfectly ordinary motion in general, and one of the commonest in the table. Its probability is 0.15 and its surprise is 2.71 bits.
That is 0.86 bits more than the perfect cadence, and it is the number the model gives for the motion. It is also obviously too small for what a deceptive cadence does, and the reason it is too small is the model’s largest gap: the weights have no memory of what prepared them.
A deceptive cadence is surprising because of what came before the dominant — because the phrase was built to arrive, the metre placed the arrival, and a listener has been counting. The root-motion model sees a V and asks what follows a V, and that is a question about a chord rather than about a cadence.
So the number this rung produces is a lower bound on the surprise, and it is the part of the surprise attributable to harmony alone with everything about preparation removed. Given the size of what is left out, 2.71 bits against 1.85 is about the best a memoryless model could be expected to do.
What the weights are worth, which is the other half
A surprise computed from asserted numbers is only as good as the ordering they encode, so the second thing this rung does is price them.
The sensitivity is not uniform. The ordering of the two cadences is robust: for a deceptive cadence to be less surprising than a perfect one, the weight on an ascending second would have to exceed the weight on a descending fifth, which reverses the single most secure claim in tonal harmony. No plausible perturbation does it.
The size of the difference is not robust. It depends on the ratio of the two weights, which is 1.0 to 0.55 and which nobody measured. At a ratio of two to one the difference is 1.0 bits; at four to one it is 2.0. So the deceptive cadence costs somewhere between about half a bit and two bits more than the perfect one, and the model cannot say where in that range.
And the entropy of the whole distribution turns out to be the most robust thing here, which is the opposite of what it looked like. It depends on all seven numbers, so the expectation was that it would be the first thing to move. Perturb every weight independently by up to thirty per cent, two hundred thousand times, and the entropy stays between 2.459 and 2.735 bits — and those are the extremes, not a spread. Entropy is a smooth saturating function of a distribution that is already near uniform, so nothing short of collapsing a weight to zero moves it.
What is fragile is the ceiling, and the static-motion weight is where the fragility lives. There is no principled reason for it to be 0.15 rather than 0.05 or 0.3, and it sets the largest surprise the model can report:
| static weight | entropy | ceiling | floor | cost of seven steps |
|---|---|---|---|---|
| 0.02 | 2.54 bits | 7.44 | 1.79 | 12.6 to 52.1 |
| 0.05 | 2.57 | 6.13 | 1.81 | 12.7 to 42.9 |
| 0.15 | 2.65 | 4.58 | 1.85 | 12.9 to 32.1 |
| 0.30 | 2.70 | 3.64 | 1.91 | 13.3 to 25.5 |
| 0.50 | 2.73 | 2.98 | 1.98 | 13.9 to 20.9 |
The floor moves by a third of a bit across that whole range and the ceiling moves by more than five. So the range a passage can cost is very nearly a property of one asserted number, and the number is the one describing a chord that does not change root — which is the least interesting motion in the table and the one nobody would have thought to argue about.
Which gives the honest summary: the ordering is a result and the magnitudes are an illustration. That is the same sentence four consecutive phases of this collection have written about asserted parameters, and it is a sentence that would stop being necessary if somebody counted root motions in a corpus, which is a small thing to count.
So the number this essay has been computing is not the whole of what a listener is surprised by, and it is not even the larger half. A surprise is a number, and there are two of them at every chord change — which is the honest limit on treating harmonic surprise as the quantity.
What this connects that was not connected
Reading the weights as probabilities makes a connection this collection had not made, and it is between two ladders that have never referred to each other.
The probe-tone profiles are a distribution over pitch classes collected from listeners. ROOT_MOTION is a distribution over root motions asserted from theory. Both are models of what a listener expects, both are used on this site to compute how well something fits, and neither has ever been put beside the other.
They are answering different questions — one about which note belongs and one about which chord follows — and a full model of expectation would need both, multiplied. A chord is surprising if its root motion is unlikely and if its notes are unlikely in the key, and those are separable and separately computable with what this site holds.
That is the rung this anchor should take next, and it is worth saying that it would produce a quantity with a name: the conditional entropy of a chord given the key and the previous chord, which is the standard object in every information-theoretic account of music and which this collection has both halves of and has never assembled.
What a whole passage costs, and the floor underneath it
Totalling a passage’s surprise gives a quantity with a floor and a ceiling, and both are informative.
The floor is the descending-fifths cycle: seven steps at 1.85 bits each, 12.94 bits in all, which is the least surprising thing that visits every degree. Nothing a composer writes in this vocabulary can cost less per step.
The ceiling is a passage of static motions — the smallest weight at 0.15 — at 4.58 bits a step. So a seven-step passage under this model costs between 12.9 and 32 bits, and everything anybody has ever written is somewhere in that range.
That is a narrow range and its narrowness is a real finding about the model. Seven destinations is at most 2.81 bits of entropy if they were equally likely, and the weights come out at 2.647 — ninety-four per cent of the maximum. That is worth sitting with. A table written down to say that descending fifths are the strongest motion in the common practice spreads its weight from 1.0 to 0.15, a ratio of nearly seven to one, and buys with it a sixth of a bit of departure from knowing nothing at all. Under this model the whole vocabulary carries between 1.85 and 4.58 bits a chord. Against that, a chromatic style with twelve possible roots rather than seven has a ceiling nearly a bit higher, and one that also changes key has more again — which is the arithmetic reason chromatic harmony can sustain surprise over long spans and diatonic harmony has to spend its budget carefully.
Which computation produced the numbers
The probability of a root motion is its weight over the sum of the weights for all seven possible destinations, with a floor of 0.02 so that nothing has probability zero and no surprise is infinite. The floor is a convention and it matters only for motions the table does not name.
The information content of a step is the negative base-two logarithm of that probability. A passage’s total is the sum over its steps, its mean is that over the number of steps, and the peak is the step carrying the most.
The three passages drawn are a perfect cadence, a deceptive cadence and a complete descending-fifths cycle, each written as a sequence of scale degrees. They are constructed rather than found, which is the same limitation the previous rung had and is the reason the next section is short.
Where the model stops
The model has no memory and a cadence is entirely about memory. A first-order model looks at the previous chord and nothing else. Every account of why a deceptive cadence works involves the phrase that led to it, the metrical position of the arrival, and the listener’s count — none of which is a property of the previous chord.
It has no metre. The same V to vi at a phrase end and in the middle of a bar are the same transition here and are not the same event; the segmentation that produces chords depends on the metre and this model takes the chords as given.
And it has no key change. A model over scale degrees assumes a key, so a chord that is surprising because it is outside the key gets no surprise at all — it simply is not in the vocabulary. Which is exactly backwards for the chromatic repertoire where surprise is most of the point, and it is the largest single reason this rung is a lower bound.
A first-order model is the weakest one that is not trivial. Anything that used two previous chords would be a second-order model with 343 transitions to estimate, which is why nobody asserts one — and a deceptive cadence is precisely a second-order phenomenon, because what makes it deceptive is the pair that prepared it.
Nothing here is measured. Both the weights and the passages are constructed, so the figures are an illustration of an arithmetic rather than a measurement of a repertoire. The corpus that would fix it is small: root motions counted, per style, with their frequencies. It is the same corpus the key-relations ladder needs for its own parameter, which means one count would serve two anchors.
Whose music, and when
The root-motion weights encode a specific claim: that descending fifths are the strongest motion. That is a claim about the common practice, roughly 1650 to 1900, and it is the claim the pedagogy of that period is built on.
It travels badly. In modal jazz the commonest root motion is by second or by third and the descending fifth is comparatively rare, so a model with these weights would find that repertoire uniformly surprising — which is a statement about the model rather than about the music. The same is true of most rock harmony, where the I–IV–V–I cycle involves an ascending fourth that this table scores at 0.45.
So the surprise computed here is surprise relative to an eighteenth-century listener, which is the right thing to compute for an eighteenth-century deceptive cadence and the wrong thing to compute for almost anything else. Making it right for another repertoire is a matter of counting that repertoire’s root motions, which is the corpus above.
There is a deeper version of the same point. A listener’s expectations are learned from what they have heard, so whose weights these are is not a detail — it is the whole model. Consonance is half learned makes the same argument about a different quantity, and this rung’s weights are the same kind of object as that essay’s preference term.
What the picture cannot show
Whether surprise is what a listener experiences. Information content is a property of a model, and the leap from “this chord is improbable under this distribution” to “a listener is surprised” is a large one that a great deal of published work makes without pausing.
And it cannot show pleasure. The interesting thing about a deceptive cadence is not that it is unlikely but that it is good — that a composer chose it and that a listener enjoys being denied. There is no quantity in this essay that has a sign, and any account of why a particular surprise is welcome and another is merely wrong is not one this collection can currently begin.
Where this ladder goes next
Three rungs. Counting produced the hierarchy; a chord did not come and the moment was described; and now the moment has a number, which is a lower bound and which prices the eight weights it rests on.
The rung after it is the multiplication the last connection names. A chord’s improbability has two independent sources — its root motion, which is this rung’s, and how well its notes fit the key, which is the probe-tone profile’s — and a listener presumably uses both. Multiplying them gives a conditional entropy for a chord given a key and a predecessor, which is a standard object, which this collection has both factors of, and which would for the first time let a passage’s surprise be computed rather than illustrated.
Part 3 of 11
One essay in the series on Tonal-expectation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
CadenceDeceptive cadenceEntropyExpectationPriorProgressionRoot motionTonal hierarchy
- A count and a correlation cadence, progression, tonal hierarchy
- A short note is heard more in tune than it is expectation, prior, tonal hierarchy
- A final chord is not made loud by adding to it cadence, expectation
- An ending that can be heard coming cadence, expectation
- An ending that exists so a bigger one can cadence, expectation
- How much evidence a modulation needs expectation, tonal hierarchy