Harmony and voice leading

Surprise is a number

The chord that did not come was described rather than measured. Its measure is the information content of what did arrive, and a model of the probability has been to hand since the key-finding essays — eight root-motion weights, ordinal and stipulated. Reading them as a distribution prices a deceptive cadence at 2.71 bits against a perfect one's 1.85, and turns up the fact that the largest of the eight had never been read by anything.

Assumes: The chord that did not come · Counting produced the hierarchy

The second rung of this ladder is about a deceptive cadence: a dominant that resolves somewhere other than the tonic, and what a listener does with the arrival that was prepared and did not come. It describes the moment carefully and it does not put a number on it, because it had none to put.

There is a standard number for exactly this, and it is one line. If a listener expects the next chord with probability p and something else arrives, the information content of what arrived is

−log₂ p   bits

— the standard measure of surprise, in the sense that a coin flip is one bit and an event with probability a thousandth is about ten. What it needs is p, and this site has had a model of p since the key ladder’s eighth rung: eight weights on root motion by each number of scale steps, used there to stop a key-finder re-analysing every bar.

How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; I – IV – V – vi totals 7.3 bits over 3 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire.
Fig. 1 Each step’s information content, in bits, under the root-motion weights used here. A perfect cadence — dominant to tonic, a motion by fourth or fifth, which the weights make the likeliest there is — costs 1.85 bits. The deceptive resolution to the submediant costs 2.71, from the same dominant, which is 0.86 bits of surprise attributable to the harmony alone.

What the weights are, and what they mean here

ROOT_MOTION is a weight for each number of scale steps a root can move by. Motion by a fourth or a fifth is the largest at 1.0, ascending seconds are 0.6, static motion is the smallest at 0.15, and the rest sit between.

The eighth rung of the key ladder used them as a transition score — a term in a dynamic program deciding which key a passage is in — and said what they are: “ordinal rather than measured, and the model’s verdicts are checked against changing them.”

Used as an expectation model they are a probability distribution over what comes next, normalised over the seven possible destinations. That is a small change of interpretation and a large change of what can be asked: a set of scores can be compared, and a distribution has an entropy, a most likely outcome, and an information content for every outcome.

So this rung is not a new model. It is the same eight numbers read as a probability instead of as a preference, and everything below is what that reading makes available.

How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; the circle of fifths totals 12.9 bits over 7 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire.
Fig. 2 Each step’s information content — minus the log of the probability those root-motion weights give it — for a full circuit of the circle of fifths. It totals 12.9 bits over seven steps, against 6.4 for a perfect cadence over three.

Per step that is 1.8 bits against 2.1, so the sequence every model of harmony is supposed to find dull is not much duller than a cadence. The weights are a graph on the seven degrees with the descending fifth as its heaviest edge, and walking that edge repeatedly is exactly what a model with those weights should find least surprising — and it very nearly does not.

The eighth weight, which nothing had ever read

Doing the arithmetic turned up a defect in the table, and it is worth reporting before the numbers that rest on it.

There are seven possible destinations from a degree and the table had eight entries, because it named both -4 — a descending fifth, V to I — and 3, an ascending fourth, I to IV. Those are the same root motion: the roots are a fourth or a fifth apart and which one it is depends on the register, which a sequence of scale degrees does not carry.

Every reader of the table folds a step into a residue, and the fold sends −4 and +3 to the same place. So ROOT_MOTION["-4"] was never read by anything. Its value is 1.0, it is the largest number in the table, its own comment describes it as the strongest motion in the common practice — and every V–I the ordered key-finder has ever scored was scored at 0.4, the weight of an ascending third.

That is now fixed, with the fourth-or-fifth class carrying the weight the comment always meant. What it changes elsewhere is small and is worth stating: the eighth rung’s headline — that an ordered model tells an alternation from a simultaneity where a histogram cannot — is unaffected, because that comparison never depended on the weight’s size. What moves is the threshold the previous essay swept, which rises.

The deceptive cadence, priced

A perfect cadence is V to I, a motion by a fourth or a fifth, which is the weights’ largest value. Under the normalised distribution its probability is 0.28 and its surprise is 1.85 bits.

A deceptive cadence is V to vi, an ascending second, weight 0.55 — a perfectly ordinary motion in general, and one of the commonest in the table. Its probability is 0.15 and its surprise is 2.71 bits.

That is 0.86 bits more than the perfect cadence, and it is the number the model gives for the motion. It is also obviously too small for what a deceptive cadence does, and the reason it is too small is the model’s largest gap: the weights have no memory of what prepared them.

A deceptive cadence is surprising because of what came before the dominant — because the phrase was built to arrive, the metre placed the arrival, and a listener has been counting. The root-motion model sees a V and asks what follows a V, and that is a question about a chord rather than about a cadence.

So the number this rung produces is a lower bound on the surprise, and it is the part of the surprise attributable to harmony alone with everything about preparation removed. Given the size of what is left out, 2.71 bits against 1.85 is about the best a memoryless model could be expected to do.

How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; the circle of fifths totals 12.9 bits over 7 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire.
Fig. 3 Three passages priced. A perfect cadence at 6.41 bits over three steps, a deceptive one at 7.27, and a complete descending-fifths cycle through all seven degrees — which is the least surprising thing the model can be shown, at 1.85 bits a step and nothing else. That cycle’s total is the floor for a passage of that length, and everything a composer does is spending above it.

What the weights are worth, which is the other half

A surprise computed from asserted numbers is only as good as the ordering they encode, so the second thing this rung does is price them.

The sensitivity is not uniform. The ordering of the two cadences is robust: for a deceptive cadence to be less surprising than a perfect one, the weight on an ascending second would have to exceed the weight on a descending fifth, which reverses the single most secure claim in tonal harmony. No plausible perturbation does it.

The size of the difference is not robust. It depends on the ratio of the two weights, which is 1.0 to 0.55 and which nobody measured. At a ratio of two to one the difference is 1.0 bits; at four to one it is 2.0. So the deceptive cadence costs somewhere between about half a bit and two bits more than the perfect one, and the model cannot say where in that range.

And the entropy of the whole distribution turns out to be the most robust thing here, which is the opposite of what it looked like. It depends on all seven numbers, so the expectation was that it would be the first thing to move. Perturb every weight independently by up to thirty per cent, two hundred thousand times, and the entropy stays between 2.459 and 2.735 bits — and those are the extremes, not a spread. Entropy is a smooth saturating function of a distribution that is already near uniform, so nothing short of collapsing a weight to zero moves it.

What is fragile is the ceiling, and the static-motion weight is where the fragility lives. There is no principled reason for it to be 0.15 rather than 0.05 or 0.3, and it sets the largest surprise the model can report:

static weight entropy ceiling floor cost of seven steps
0.02 2.54 bits 7.44 1.79 12.6 to 52.1
0.05 2.57 6.13 1.81 12.7 to 42.9
0.15 2.65 4.58 1.85 12.9 to 32.1
0.30 2.70 3.64 1.91 13.3 to 25.5
0.50 2.73 2.98 1.98 13.9 to 20.9

The floor moves by a third of a bit across that whole range and the ceiling moves by more than five. So the range a passage can cost is very nearly a property of one asserted number, and the number is the one describing a chord that does not change root — which is the least interesting motion in the table and the one nobody would have thought to argue about.

Which gives the honest summary: the ordering is a result and the magnitudes are an illustration. That is the same sentence four consecutive phases of this collection have written about asserted parameters, and it is a sentence that would stop being necessary if somebody counted root motions in a corpus, which is a small thing to count.

Two surprises at every chord change. Each chord change of a short progression, with the two surprises a listener meets at it: how unlikely the chord is given the one before, and how unlikely the moment is given the metre. They are not the same size and they are not the same shape — the identity surprise runs from 1.85 to 2.85 bits and the timing surprise from 2.02 to 4.52, and the timing one is the larger at every event here. The pale outline is the total if the two are perfectly dependent, which is the smaller of the two ends of the family swept in the next figure. A listener meets one event, so what they are surprised by is somewhere between the outline and the bar.
Fig. 4 Each chord change with the two surprises a listener meets at it: how unlikely the chord is given the one before, and how unlikely the moment is given the metre. The identity surprise runs from 1.85 to 2.85 bits and the timing surprise from 2.02 to 4.52, and the timing one is the larger at every event here.

So the number this essay has been computing is not the whole of what a listener is surprised by, and it is not even the larger half. A surprise is a number, and there are two of them at every chord change — which is the honest limit on treating harmonic surprise as the quantity.

What keeping the order costs, on a logarithmic axis. The two models' sizes. A histogram model carries 24 hypotheses — twelve major keys and twelve minor — and no transitions at all, because it has nothing to transition between. An ordered model over (key, degree) carries 84 states and 7,056 transitions between them. That is a factor of 3.5 in hypotheses and an infinite factor in transitions, and it is the trade named earlier: what a model can distinguish against how many hypotheses it must carry.
Fig. 5 The model the weights were built for, from an earlier essay on key finding: eighty-four states and 7,056 transitions, against a profile model’s twenty-four hypotheses and none. Every one of those transitions is scored by the eight numbers this essay is reading as a probability — which is why a defect in one of them was worth finding, and why it went unnoticed for a phase.

What this connects that was not connected

Reading the weights as probabilities makes a connection this collection had not made, and it is between two ladders that have never referred to each other.

The probe-tone profiles are a distribution over pitch classes collected from listeners. ROOT_MOTION is a distribution over root motions asserted from theory. Both are models of what a listener expects, both are used on this site to compute how well something fits, and neither has ever been put beside the other.

They are answering different questions — one about which note belongs and one about which chord follows — and a full model of expectation would need both, multiplied. A chord is surprising if its root motion is unlikely and if its notes are unlikely in the key, and those are separable and separately computable with what this site holds.

That is the rung this anchor should take next, and it is worth saying that it would produce a quantity with a name: the conditional entropy of a chord given the key and the previous chord, which is the standard object in every information-theoretic account of music and which this collection has both halves of and has never assembled.

What a whole passage costs, and the floor underneath it

Totalling a passage’s surprise gives a quantity with a floor and a ceiling, and both are informative.

The floor is the descending-fifths cycle: seven steps at 1.85 bits each, 12.94 bits in all, which is the least surprising thing that visits every degree. Nothing a composer writes in this vocabulary can cost less per step.

The ceiling is a passage of static motions — the smallest weight at 0.15 — at 4.58 bits a step. So a seven-step passage under this model costs between 12.9 and 32 bits, and everything anybody has ever written is somewhere in that range.

That is a narrow range and its narrowness is a real finding about the model. Seven destinations is at most 2.81 bits of entropy if they were equally likely, and the weights come out at 2.647 — ninety-four per cent of the maximum. That is worth sitting with. A table written down to say that descending fifths are the strongest motion in the common practice spreads its weight from 1.0 to 0.15, a ratio of nearly seven to one, and buys with it a sixth of a bit of departure from knowing nothing at all. Under this model the whole vocabulary carries between 1.85 and 4.58 bits a chord. Against that, a chromatic style with twelve possible roots rather than seven has a ceiling nearly a bit higher, and one that also changes key has more again — which is the arithmetic reason chromatic harmony can sustain surprise over long spans and diatonic harmony has to spend its budget carefully.

The seven chords of a key, by distance from home. Each triad of the major scale placed at a radius equal to how far its voices must move from the tonic chord. The chords that feel closest to home are the ones that are closest, in the plain arithmetic sense.
Fig. 6 The space the walk happens in, from an earlier essay on progressions: the seven degrees with a geometry on them, and one path across it. Those essays are about the shape of such a path; this one is about its cost, and the two are the same object measured with different instruments.

Which computation produced the numbers

The probability of a root motion is its weight over the sum of the weights for all seven possible destinations, with a floor of 0.02 so that nothing has probability zero and no surprise is infinite. The floor is a convention and it matters only for motions the table does not name.

The information content of a step is the negative base-two logarithm of that probability. A passage’s total is the sum over its steps, its mean is that over the number of steps, and the peak is the step carrying the most.

The three passages drawn are a perfect cadence, a deceptive cadence and a complete descending-fifths cycle, each written as a sequence of scale degrees. They are constructed rather than found, which is the same limitation the previous rung had and is the reason the next section is short.

Where the model stops

The model has no memory and a cadence is entirely about memory. A first-order model looks at the previous chord and nothing else. Every account of why a deceptive cadence works involves the phrase that led to it, the metrical position of the arrival, and the listener’s count — none of which is a property of the previous chord.

It has no metre. The same V to vi at a phrase end and in the middle of a bar are the same transition here and are not the same event; the segmentation that produces chords depends on the metre and this model takes the chords as given.

And it has no key change. A model over scale degrees assumes a key, so a chord that is surprising because it is outside the key gets no surprise at all — it simply is not in the vocabulary. Which is exactly backwards for the chromatic repertoire where surprise is most of the point, and it is the largest single reason this rung is a lower bound.

A first-order model is the weakest one that is not trivial. Anything that used two previous chords would be a second-order model with 343 transitions to estimate, which is why nobody asserts one — and a deceptive cadence is precisely a second-order phenomenon, because what makes it deceptive is the pair that prepared it.

Nothing here is measured. Both the weights and the passages are constructed, so the figures are an illustration of an arithmetic rather than a measurement of a repertoire. The corpus that would fix it is small: root motions counted, per style, with their frequencies. It is the same corpus the key-relations ladder needs for its own parameter, which means one count would serve two anchors.

Whose music, and when

The root-motion weights encode a specific claim: that descending fifths are the strongest motion. That is a claim about the common practice, roughly 1650 to 1900, and it is the claim the pedagogy of that period is built on.

It travels badly. In modal jazz the commonest root motion is by second or by third and the descending fifth is comparatively rare, so a model with these weights would find that repertoire uniformly surprising — which is a statement about the model rather than about the music. The same is true of most rock harmony, where the I–IV–V–I cycle involves an ascending fourth that this table scores at 0.45.

So the surprise computed here is surprise relative to an eighteenth-century listener, which is the right thing to compute for an eighteenth-century deceptive cadence and the wrong thing to compute for almost anything else. Making it right for another repertoire is a matter of counting that repertoire’s root motions, which is the corpus above.

There is a deeper version of the same point. A listener’s expectations are learned from what they have heard, so whose weights these are is not a detail — it is the whole model. Consonance is half learned makes the same argument about a different quantity, and this rung’s weights are the same kind of object as that essay’s preference term.

What the picture cannot show

Whether surprise is what a listener experiences. Information content is a property of a model, and the leap from “this chord is improbable under this distribution” to “a listener is surprised” is a large one that a great deal of published work makes without pausing.

And it cannot show pleasure. The interesting thing about a deceptive cadence is not that it is unlikely but that it is good — that a composer chose it and that a listener enjoys being denied. There is no quantity in this essay that has a sign, and any account of why a particular surprise is welcome and another is merely wrong is not one this collection can currently begin.

Where this ladder goes next

Three rungs. Counting produced the hierarchy; a chord did not come and the moment was described; and now the moment has a number, which is a lower bound and which prices the eight weights it rests on.

The rung after it is the multiplication the last connection names. A chord’s improbability has two independent sources — its root motion, which is this rung’s, and how well its notes fit the key, which is the probe-tone profile’s — and a listener presumably uses both. Multiplying them gives a conditional entropy for a chord given a key and a predecessor, which is a standard object, which this collection has both factors of, and which would for the first time let a passage’s surprise be computed rather than illustrated.

Part 3 of 11

One essay in the series on Tonal-expectation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

CadenceDeceptive cadenceEntropyExpectationPriorProgressionRoot motionTonal hierarchy