Harmony and voice leading

Eighty-one chords, one number

A dominant seventh has eighty-one arrangements inside three octaves and their roughness spans a factor of five and a half. The tonal-expectation model gives every one of them the same 3.51 bits, because its states are scale degrees and there is no register anywhere in them. Conditioning the surprise on the voicing costs no corpus — and the arithmetic says the conditioning belongs beside the probability rather than inside it, for three reasons that can each be computed.

Assumes: Where the chord actually lands · A chord is a register

Every quantity on this ladder is a set of pitch classes at a position in a bar. Surprise is a number prices a chord by its root motion; a chord given a key and a predecessor multiplies that by how well its notes fit the key; where the chord actually lands prices when it arrives. Not one of those quantities knows where the chord is.

A chord is not a set of pitch classes. It is voiced, in a range, with a note at the bottom, and a player has to decide all of that before anybody hears anything.

Eighty-one chords the expectation model cannot tell apart. Every voicing of a dominant seventh on G inside the three octaves above its own root, placed by how rough it is and how far its outer voices are apart, and coloured by which member of the chord is at the bottom. The roughness runs from 0.269 to 1.449, a factor of 5.4, computed from each voicing's own spectrum under Plomp and Levelt's roughness model. The identity surprise the expectation model assigns is 3.51 bits for every one of the 81, because it is a function of a scale degree and its predecessor and there is no register anywhere in it. What separates them is spacing rather than inversion: roughness falls as the outer voices spread apart, correlating -0.42 with the span, and is indifferent to which member of the chord is at the bottom at 0.02. The seventh in the bass is not what makes a chord rough; a fourth and a third packed together at the bottom of the range is.
Fig. 1 Every arrangement of a dominant seventh inside three octaves above its root, by roughness and by how far the outer voices are apart. The identity surprise is one number for the whole scatter, and the roughness spans a factor of five and a half.

Eighty-one arrangements, roughness from 0.269 to 1.449, and 3.51 bits for every one of them. That is not an approximation the model makes; it is a consequence of what its states are. A scale degree has no register in it, so a dominant seventh in close position at the top of a treble stave and the same chord spread over two octaves with its seventh underneath are the same event, exactly, to everything this ladder computes.

The ingredients for a repair are already here. A chord is a register enumerates the voicings and computes the roughness of each from its own spectrum under Plomp and Levelt’s model rather than from the chord’s name. So conditioning the identity surprise on voicing needs no corpus, no recording and no listener. What it needs is a decision about whether the conditioning belongs in the probability or beside it, and that decision is what this rung is.

What the roughness is actually about

Before deciding where the term goes, it is worth knowing what it measures, because the obvious answer is wrong.

The obvious answer is the bass note. It is the one the ear supplies a root from — the root an ear supplies is the collection’s account of that — and it is the note an orchestration manual’s low-interval-limit table is written about. Over the eighty-one arrangements above, roughness correlates 0.02 with which member of the chord is at the bottom. It is indifferent to the inversion.

What it correlates with is the spacing, at −0.42, and the sign is the surprising part: a chord gets smoother as its outer voices move apart. The roughest arrangement in the scatter is the textbook one — root, third, fifth and seventh packed into a tenth at the bottom of the range — and the smoothest is wide at the bottom and close at the top, which is the spacing of the harmonic series itself.

Every voicing of a dominant seventh over G2, least rough first. All 81 arrangements of the same three pitch classes within 3 octaves from 98 Hz, scored for roughness. The best is spaced 19 then 9 then 6 semitones — wide below, close above — and the worst is the chord in close position at the bottom of the range, 5.4 times rougher with exactly the same notes in it.
Fig. 2 The same eighty-one arrangements as pitch strips, least rough first. The best of them are not close position at the bottom; they are the spacing an orchestrator would call open, which is the spacing the harmonic series has.

That is the earlier essay’s finding arriving on a different ladder, and it matters for the decision ahead. The debt that named this rung named the seventh in the bass as the case that would separate two events. The arithmetic says the bass is not what separates them at all, and that the quantity available here is about spacing and register rather than about which note is underneath.

The seven continuations, four times over

The comparison the ladder actually needs is not one chord’s voicings but seven candidates’ — the seven degrees a passage could go to next, each in the voicing a player would reach for from where the hands already are.

The same seven chances, four different sounds. The seven continuations from the tonic, each with the probability the expectation model gives it and the roughness of the voicing a player would reach for at four registers a bodily octave apart. The bars on the left are the probabilities and they are identical at every register — the same seven numbers, to the last bit, because the model's states are scale degrees. The columns on the right are the roughnesses, and they fall from a mean of 2.370 at the bottom to 0.280 at the top — a factor of 8.5. A listener meets the second column as well as the first, and the expectation model has only ever computed the first.
Fig. 3 The seven continuations from the tonic, with the probability the model gives each and the roughness of the nearest voicing at four registers an octave apart. The left column is one set of numbers; the right column is four.

The probabilities are identical at every register, to the last bit, and they have to be: the model’s states are scale degrees and a bodily transposition changes no degree. The roughnesses fall from a mean of 2.370 two octaves below the written range to 0.280 an octave above it — a factor of 8.5 — which is the same steep fall with pitch that a chord is a register measured at a factor of fifteen across the piano.

So a listener meets two columns and the ladder has only ever computed one. The question is what to do about the second.

Putting it inside the probability

The first option is the one the debt leans toward. Reweight the seven candidates by how rough each would sound, renormalise, and read the arrived-at chord’s surprise off the new distribution. The chord’s improbability then has three sources rather than two — its root motion, its fit to the key, and how hard it is to hear where it is being played — and the model that comes out is the same shape as a chord given a key and a predecessor, which multiplied two sources and renormalised over the same seven destinations.

There is one thing to settle before it can be tried, and it turns out to be the whole argument. A term added inside a log-probability has an arbitrary scale. Multiplying by the exponential of minus some coefficient times roughness is not a model until the coefficient is chosen, and there is no measurement anywhere in this collection that fixes it.

The non-arbitrary way to choose it is to ask what strength makes the new term worth as much as a term the model already has. The profile term spreads the seven candidates by a certain amount in log space; set the coefficient so that the roughness term spreads them by the same amount, and the new source is exactly as loud as the one beside it. That is a defensible choice and it is computable.

The coefficient that is not a number. The strength a roughness term would need in order to be worth as much as the profile term the model already carries — the value that gives it the same spread across the seven candidates — plotted against register. It runs from 1.45 at the bottom of the range to 3.35 at the top, a factor of 2.3, because roughness and its spread across the candidates both fall steeply as the music rises while the profile term does not move at all. A coefficient inside a probability is one number, so any single choice misprices the term everywhere but one octave. The lower line is how much of the roughness the profile term already carries: the two agree 0.42, 0.53, 0.59, 0.60 across the four registers, so a third of what the new term would add is already in the distribution.
Fig. 4 The coefficient that would make a roughness term worth one profile term, against register. It is not a number: it more than doubles over four octaves, because roughness and its spread across the candidates both fall steeply with pitch while the profile term does not move.

There is a way to dodge that, and it is worth naming so that it can be rejected explicitly. The coefficient could be made a function of register rather than a constant — fitted per octave, or written as a decreasing function of the mean roughness. That is no longer a third source of improbability; it is a normalisation chosen to keep the third source at a fixed loudness, which is a way of saying that the term’s own scale carries no information and only its ranking does. A factor whose strength has to be re-derived at every pitch is not a factor in a probability. It is a comparison, and a comparison belongs beside.

It is not a number. The coefficient runs from 1.45 two octaves down to 3.35 an octave up, a factor of 2.3, and the reason is structural rather than accidental: roughness falls steeply with register and so does its spread across the seven candidates, while the profile term is a property of pitch classes and does not move at all. A coefficient inside a probability has to be a single constant. Any single choice therefore misprices the term everywhere but one octave, and the mispricing is not small — at the value calibrated for a bass register the term is worth less than half a profile term in the treble.

The second thing the figure reports is how much of the new term is already in the distribution. Across the four registers the roughness of a candidate agrees with its profile fit at correlations of 0.42, 0.53, 0.59 and 0.60. Between a fifth and a third of the variance the new factor would contribute is variance the model already has, because the chords that fit a key well are largely the chords built on its consonant degrees, and consonance is what roughness measures.

And what it buys, which is very little

The third test is the one that ought to be decisive on its own: run the model both ways and see whether anything moves.

Two numbers that are not one number. Every arrival in I – vi – IV – V – I at four registers: how surprising the chord was, against how rough the voicing it arrived in is. The vertical position of a point is fixed by the scale degree and its predecessor, so the four registers give four points at the same height and different roughnesses — the horizontal lines are single events that the expectation model has been reporting as one number. Over the 16 arrivals the two quantities correlate 0.03, which is nothing: knowing how unexpected a chord was says nothing about how hard it is to hear. Folding the second into the first at the strength the previous figure calibrates moves the passage's bill by 1.9%, 2.2%, 2.0%, 0.1% and changes which continuation is most expected at 1 of the 16 arrivals. So the conditioning belongs beside the probability rather than in it.
Fig. 5 Sixteen arrivals — a four-chord progression at four registers — placed by how surprising the chord was against how rough the voicing it arrived in was. Each dashed line joins one arrival played in four different places, which is one point to everything computed so far.

At its own calibrated strength the roughness term moves the passage’s expectation bill by 1.9, 2.2, 2.0 and 0.1 per cent at the four registers, and it changes which continuation is most expected at one of the sixteen arrivals. A term that has to be calibrated by three separate arguments and then moves the answer by two per cent is not doing the work it was added for.

It is worth being precise about what “moves the bill by two per cent” means, because a small number is easy to dismiss for the wrong reason. The bill is a sum of four surprises and it changes by about a fifth of a bit; a fifth of a bit is well inside the range over which the collection’s own weights are asserted rather than measured, so the change is smaller than the uncertainty in the numbers being changed. Surprise is a number makes that point about the root-motion table itself — the ordering is the result and the magnitudes are the illustration — and a new factor that moves the total by less than the table’s own slack has not been shown to do anything.

The one arrival where the ranking does move is the one at the top of the range, where the roughness spread across candidates is smallest in absolute terms and the calibrated coefficient is therefore largest. That is the opposite of where a listener would expect voicing to matter most, and it is a consequence of the calibration rather than of the music — one more sign that the term is being forced into a shape the quantity does not have.

That is the case against putting it inside. The case for putting it beside is on the same figure and it is stronger.

Over those sixteen arrivals the surprise of the chord and the roughness of the voicing it arrived in correlate 0.03. They are, to a very good approximation, independent. Knowing that a chord was unexpected tells a listener nothing about whether it was hard to hear, and knowing that it was rough tells nothing about whether it was expected.

Two independent numbers are worth more as two numbers than as one. Folding the second into the first at any coefficient discards exactly the information that made it worth having: the pair splits the sixteen arrivals into four quadrants — the expected and smooth, the expected and rough, the surprising and smooth, the surprising and rough — and a single reweighted probability puts all four on one line.

Which is a decision about what a probability is for

The three computed reasons all point the same way and there is a fourth that is not computed and should be said anyway, because it is the one that would still hold if the numbers had come out differently.

A voicing is chosen by the performer. A scale degree, a key and a metrical position are chosen by whoever wrote the piece. If the roughness of a voicing enters the probability, then the information content of a progression stops being a property of the progression: the same four bars carry different bits played by a string quartet and by a piano, and a conductor who asks the second violins to take a note up an octave has changed how surprising the harmony is.

That is not obviously wrong — it is a defensible position about what a listener’s model conditions on — but it is a large claim, and nothing in this collection supports it. Reported beside the probability, the same arithmetic makes a much smaller and much better-attested claim: this chord was this surprising, and it arrived sounding like this. The first number is about the composition and the second is about the performance, and a reader can see which is which.

That distinction is the same one two surprises and one event had to make about identity and timing, and it was resolved there in the other direction — those two were added, under a swept correlation, because both are about the same decision by the same person at the same moment. Timing and identity are both the composer’s. Voicing is not.

Whose music this is a claim about

The voicings above are four-part, in the ranges of a chorale — bass, tenor, alto and soprano, the four choral compasses this collection has used since what the rules cost priced the prohibitions in them, and which why the exercise is in four parts showed to be the smallest count at which those prohibitions cost anything — and the nearest-voicing rule that picks one is a rule about a keyboard player’s hands rather than about four singers.

That matters for how far the numbers travel. In the common-practice repertoire the four-part chorale texture is the case in which spacing is most constrained and least expressive: the ranges are fixed, the parts do not cross, and an editor can predict most voicings from the chord and the previous voicing. It is therefore the texture in which the register conditioning has the least to add, and the two-per-cent figure above should be read as a floor.

The textures where it would have most to add are the ones this collection has not modelled at all: orchestral scoring, where the same chord is spread over five octaves and the doublings are the composition; close-position jazz voicings in the middle of the piano, where the spacing is what distinguishes one arranger from another; and the guitar, where the voicing is very nearly forced by the instrument’s tuning and is therefore not a free variable at all. The low-interval-limit tables that the roughness model turns out to derive are written about the first of those and not about any of the others.

Which computation produced the numbers

The identity surprise is the fourth rung’s, unchanged: a root-motion weight times a probe-tone fit, renormalised over the seven degrees of the key, and its logarithm in bits. The root-motion weights are this collection’s own ordinal table and the profile is Krumhansl and Kessler’s measured probe-tone ratings.

The voicings come from the four-part enumerator, which places every pitch class of the chord in every position in range and keeps the arrangements in which all of them sound. A register is those four ranges moved bodily by a stated number of semitones, so nothing about the relative spacing is changed by moving the passage — which is what makes the comparison a comparison.

The voicing a candidate would be played in is the one nearest the previous chord’s voicing, summing absolute semitone motion over the four parts. That is a minimal assumption and it is the same quantity the shortest move minimises, with the counterpoint rules switched off.

Roughness is Plomp and Levelt’s, summed over every pair of partials of every pair of notes, with a string-like spectrum. The scatter of eighty-one arrangements uses the wider enumerator that allows any octave placement within a span, which is why it includes arrangements no four-part writer would choose.

The surprise of each chord, against the uncertainty it arrived into. The information content of each step — minus the log of its probability under a distribution that multiplies the root-motion weight by how well the destination triad's notes fit the key — with the entropy of the moment before it drawn behind. I – vi – IV – V – I: I→vi 2.68 bits, vi→IV 2.68 bits, IV→V 2.80 bits, V→I 1.60 bits, against a mean uncertainty of 2.64. A surprise larger than the entropy it arrived into is an outcome the model was not expecting even given how uncertain it was; one below it is an outcome the model had already mostly bet on. An earlier essay produced the first of those numbers and had no way to produce the second, because a set of preferences is not a distribution and only a distribution has an entropy.
Fig. 6 The identity surprise as it has always been computed, on the progression used above. Every number on it is a function of a degree and its predecessor, and this essay adds nothing to any of them — it adds a second number beside each.

Where the account stops

Roughness is not the only thing a voicing does. A wide spacing is also quieter per part, easier to hear the individual lines in, and slower to fuse — the spectrum that will not fuse is the collection’s account of fusion, and none of it is in the number used here.

The nearest voicing is not the played one. Real players choose voicings for reasons that include the melody, the instrument’s mechanics and what comes next, and the minimal-motion rule is a stand-in for all of that. Where the rule is wrong the roughness attached to a candidate is the roughness of a chord nobody would have played.

And the independence is measured on one progression. Sixteen arrivals is not a sample, and the correlation of 0.03 is a demonstration that the two quantities can be independent rather than a measurement of how independent they are in music. What would settle it is the corpus this anchor keeps recording as owed — with the voicings in it, which the recorded version of that debt does not ask for and now should.

The seven chords of a key, by distance from home. Each triad of the major scale placed at a radius equal to how far its voices must move from the tonic chord. The chords that feel closest to home are the ones that are closest, in the plain arithmetic sense.
Fig. 7 The same progression as a walk over the seven triads placed by voice-leading distance, which is the map drawn at the outset. A voicing is what turns each of those points into a sound, and the map has never had one.

Where this ladder goes next

Nine rungs: a measured hierarchy of pitch stability, a chord that fails to arrive, a surprise given a number, that number conditioned on a key and a predecessor, expectation drawn as a curve, two surprises added under a swept correlation, the third term the curve had all along, the curve read at the beats a chord actually lands on, and now the register the whole ladder has been missing.

What is owed after this is the pair as a statistic. Two numbers per arrival is the answer this rung arrives at and it is not yet a measurement of anything: what a reader wants to know is which quadrant real music lives in, and whether composers put rough voicings under surprising chords or under expected ones. That is a joint distribution over two continuous quantities and it needs analysed harmony with its voicings, which is a sharper request than the corpus this anchor has been recording — a set of pitch-class progressions would not answer it, and every corpus this collection has asked for so far has been exactly that.

The second debt is arithmetic and can be paid immediately: the roughness of a voicing is not the only register-dependent quantity available, and the other one is already computed. Where the chord actually lands prices when a chord arrives; roughness prices where it arrives; and a chord’s loudness at a register decides both how rough it is and how well it masks what surrounds it. The masking machinery is in this collection and has never been pointed at a progression. A version of the pair with a third coordinate would say whether a rough arrival is rough because of its spacing or because of its level, and those are separable in the model even though a player cannot separate them at the instrument.

The third is the one this rung declines to do and should be named so that somebody can argue with the refusal. Conditioning the probability on voicing is refused here on three computed grounds and one uncomputed one, and the uncomputed one — that a probability should condition on what the composer chose and not on what the performer chose — is a position rather than a result. A listener who has never seen a score has no access to that distinction at all, and a model of that listener would be right to fold the two together. The test that would settle it is a listener’s, not an arithmetic one: play the same progression in two voicings and ask which chord was expected. Nothing here can do that, and it is the honest place for this ladder to stop.

Part 9 of 11

One essay in the series on Tonal-expectation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Critical bandwidthExpectationInformationOctave equivalenceRegisterRoughnessSurpriseVoicing