Harmony and voice leading

Three decisions that constrain each other

Every model here decides one thing at a time — the key from the pitch classes, the metre from the onsets, the chords from the metre — and an earlier essay ended by saying a listener does all three at once. Resolving them jointly costs a hundred and fifty-seven times the search and changes the reading of two passages in five. It never once changes the key, and the reason it cannot is the reason the whole account is built the way it is.

Assumes: Which notes are the chord · The time signature is a claim

The previous rung put the other three on a footing and left one thing owed. A progression is a list of chords; before there is a list, something has to decide which of the sounding notes are chord tones and which are passing; and the thing that decides is the metre. So harmonic analysis is a function of a variable that is not harmony.

The rung ended by naming the joint problem. Every model on this site decides one thing at a time — the key from the pitch classes, the metre from the onsets, the chords from the metre — and a listener does all three at once, with each constraining the others. A model that resolved them jointly would be a considerably better model and a considerably harder one to draw.

It is one of those things. Which one is the rung.

What the joint search changes, and what it never changes. Over 552 constructed passages of eight slots with rests, how often the joint reading differs from the pipeline's. The chord differs in 29 per cent and the barline in 31, with both differing in 20. The key differs in 0 per cent — never — because the key is read from a pitch-class histogram, which does not know where the bar starts or which notes are chord tones. Two of the three decisions are entangled and the third is not.
Fig. 1 Over five hundred and fifty-two constructed passages, how often the joint reading differs from the pipeline’s. The chord changes in three passages in ten and the barline in three in ten, with both changing in one in five. The key changes in none of them, ever, and that is not a small number in a noisy measurement — it is a structural zero.

The pipeline, written out

The sequence every model on this site follows, when it is followed all the way through, is three stages and each one throws away what the next might have wanted.

Stage one, the metre. Take the onsets — just the times, no pitches — and score each candidate barline position by how well a periodic accent pattern explains the onsets. That is metreFit, which the rhythm ladder has used since the metre rung, and which is the model a time signature can be checked against.

Stage two, the chords. Weight each note by the metrical position the winning barline gives it, then score every chord at every root by how much of the weighted note mass it accounts for and how much of itself was actually sounded. That is segmentFit, and it is what the previous rung is built on.

Stage three, the key. Build a pitch-class histogram and correlate it against twenty-four rotated probe-tone profiles. That is the key-finder this ladder’s neighbour has been running for four rungs.

Each stage commits. The metre is chosen before any pitch has been looked at; the chord is chosen inside a metre that cannot now be revised; and the key arrives last, with no power to object to either.

The joint search, and what it costs

The joint version uses the same three scoring functions and changes nothing about them. What changes is that every combination is scored together and the best total wins.

The total is the product of the three, not the sum, and the choice is deliberate: a product refuses a hypothesis that is bad on any one of the three, which is what “constraining each other” means. A sum lets a triple with an excellent chord and an impossible metre outscore a triple that is good at everything.

The cost is a factor of a hundred and fifty-seven. The pipeline evaluates eight barline positions, then a hundred and forty-four chord-and-root pairs, then twenty-four keys: a hundred and seventy-six evaluations, because each stage’s answer is fixed before the next begins. The joint search evaluates their product: eight times a hundred and forty-four times twenty-four, which is twenty-seven thousand six hundred and forty-eight.

That is the “considerably harder” made into a number, and it is the mild version. The passage here is eight slots long, and a model that also searched over bar lengths, over subdivision trees, and over sequences of chords rather than one chord per window would multiply again by each — by roughly three, four and a hundred and forty-four:

what is searched hypotheses against the pipeline
the pipeline 176
the joint search here 27,648 ×157
plus three bar lengths 82,944 ×471
plus four subdivision trees 331,776 ×1,885
plus a chord per half-bar 47,775,744 ×271,453

The last row is the honest size of the thing this rung is a miniature of, and it is why nobody writes the joint model. A quarter of a million times the pipeline’s work, for eight quavers.

The product, which turns out not to matter

The choice of a product over a sum was argued for above and never tested, and testing it is one line: score every hypothesis both ways and count the disagreements.

On 552 passages the two rules choose the same triple every time. Not usually — every time, zero disagreements. The argument for the product is sound and the passages this rung is built on do not contain a case where it bites, because a hypothesis bad enough on one term to need vetoing is already losing on the other two.

That is worth recording as a null rather than quietly dropping, because it says where the argument’s force actually is. The product matters when one term can be near zero while the others are excellent, and in this construction the metre term is nearly flat — the same flatness that lets the chord overturn the barline — so nothing ever gets near zero on it.

Which term does the overturning

Ablating the joint search’s own terms, against its full product on the same passages, says how the disagreement with the pipeline is divided. Dropping the key term changes the joint answer in 12 per cent of passages and dropping the metre term in 16. So of the disagreement the rung reports, the key’s veto accounts for about a quarter and the metre’s own contribution for about a third, with the chord term — which both models share — doing the rest.

That is a more useful decomposition than the headline, because it prices the one thing the pipeline structurally cannot do. The key term can never change the key, and it changes the chord in one passage in eight.

The pipeline and the joint search agree here. Every hypothesis the joint search considers for this passage, as a point: how well its phase explains the onsets against how well its chord explains the notes, with the key's own correlation and its agreement with the chord folded into the shading. The pipeline chooses the best phase first and is then committed — it reads C major7 in C major with the bar starting at slot 0. The joint search reads C major7 in C major at slot 0. The two searches are 176 hypotheses and 27648, a factor of 157.
Fig. 2 The hypothesis space for the earlier example — the scale as eight quavers — with each hypothesis placed by how well its barline explains the onsets and how well its chord explains the notes. Every barline position scores the same on the horizontal axis, because a note on every slot gives the metre nothing to prefer. The pipeline and the joint search land in the same place here, and they land there for different reasons: the pipeline because it broke a tie arbitrarily and the joint search because the tie was broken by the harmony.

Where they disagree

On five hundred and fifty-two constructed passages — eight slots, some of them rests, notes drawn from one diatonic collection — the two readings differ in thirty-nine per cent of cases.

The chord differs in twenty-nine per cent. The barline differs in thirty-one. Both differ in twenty.

The mechanism is visible in the individual cases and it runs in the direction the previous rung predicted. A passage whose onsets weakly prefer one barline gets that barline from the pipeline, and the chord it then finds is whatever those metrical weights favour. The joint search will trade a little of the metre’s score for a great deal of the chord’s, and the trade is often available because the metre score is nearly flat over barline positions in any passage with notes on most of its slots — which is most music.

The commonest single mechanism is the key vetoing a chord. The pipeline’s third stage has no power to reject; the joint search’s key term multiplies a chord’s score down when the chord’s root is outside the key it names. In one worked case the pipeline reads F major seventh and the joint search reads A minor seventh, on the same notes at the same barline, because the key is E minor and F natural is not in it.

The pipeline and the joint search do not agree. Every hypothesis the joint search considers for this passage, as a point: how well its phase explains the onsets against how well its chord explains the notes, with the key's own correlation and its agreement with the chord folded into the shading. The pipeline chooses the best phase first and is then committed — it reads F major7 in E minor with the bar starting at slot 0. The joint search reads A minor7 in E minor at slot 0. The two searches are 176 hypotheses and 27648, a factor of 157.
Fig. 3 A passage where they part. The pipeline commits to a barline, then reads the chord that barline favours; the joint search takes a slightly worse barline for a much better chord, and the key term does the deciding. Every point in the cloud is one hypothesis, and the two marked ones are the two answers.

Both of those readings run on the same eight numbers, so it is worth seeing the weights themselves before asking which stage should be allowed to overturn which.

Syncopation against a bar of 8. An 8-step pattern with 3 onsets, against the metrical weights of its bar. A position's weight is zero on the downbeat and one lower at each level down the subdivision tree, drawn here as the depth of the bar hanging beneath it. A note on a weak position followed by a rest on a stronger one costs the difference. the metre scores 0 — no note here is followed by a rest on a stronger position.
Fig. 4 The weights the segmentation runs on: zero on the downbeat and one lower at each level down the subdivision tree, which is Longuet-Higgins and Lee’s definition. Everything a metre contributes to a harmonic analysis passes through this eight-number vector, and rotating a passage against it is the whole of what moving the barline does.
One pattern, four metres. The same 8-step onset pattern read under 3 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are in four 0, in eight -2, in three 2, so in three wins. Nothing about the sound differs between these readings; the bar line is supplied by the listener.
Fig. 5 And the first stage in its own right: candidate metres scored against a bare onset list, which is where the pipeline’s commitment is made. The scores here are close together, and closeness is the condition under which the second stage can overturn the first — a passage whose onsets strongly prefer one barline gives the joint search nothing to trade.

Before the three are put together it is worth seeing the grid that the second of them runs on, because every weight in this essay descends from it.

The grid that decides which notes count. Longuet-Higgins and Lee's metrical weights for a bar of 8, with 8 notes on them. Zero on the downbeat and one less at each level down, so the first and fifth quavers carry more weight than the second and sixth and far more than the odd-numbered ones. Nothing about the notes themselves distinguishes them; the grid is the whole of the difference, and it is imposed by the barline rather than read off the sound.
Fig. 6 Longuet-Higgins and Lee’s metrical weights for a bar of eight with eight notes on it: zero on the downbeat and one less at each level down, so the first and fifth quavers carry more weight than the second and sixth and far more than the odd-numbered ones.

Nothing about the notes distinguishes them here — they are eight identical events — so every difference in the drawing is supplied by the metre and none of it by the music. That is the sense in which the segmentation decision depends on the metre decision, and it is why the two cannot be made in either order without the other already being made.

The zero

The number that is worth more than the other three is the one that is not there.

The key never changes. Not rarely: never, in five hundred and fifty-two passages.

The reason is structural and it is the thing the whole ladder should be read through. The key is inferred from a pitch-class histogram, and a histogram is invariant under everything the other two decisions are about. Move the barline and the histogram is unchanged, because a histogram has no bar in it. Choose a different chord and the histogram is unchanged, because the chord is a conclusion about the notes and not an input.

So there is nothing for the joint search to do to the key. Two of the three decisions are entangled with each other and the third is coupled to them in one direction only: the key can constrain the chord, and the chord cannot inform the key.

The same eight notes, read three ways. The scale, eight quavers, scored against every triad and seventh at every root. Barred as written the best reading is C major7 at 0.850; with the barline one quaver later it is D minor7 at 0.850. With no metre — every note weighted the same — 4 readings tie at 0.625 and the passage has no best analysis at all. The notes are identical in all three. What changed is where the bar starts, which is not a fact about harmony.
Fig. 7 The earlier figure, which is the entangled pair in isolation. The same eight notes read three ways: barred as written the best reading is C major seventh; with the barline one quaver later it is D minor seventh; with no metre at all four readings tie. The histogram behind all three is the same histogram, and the key-finder gives the same answer to all three.

That is a limitation of the model and not a fact about hearing. A listener plainly does use the chords to decide the key — a single V–I settles a key faster than a bag of its pitch classes ever could, and the whole apparatus of tonicisation is about chords declaring a key. How much evidence a modulation needs measured that in the histogram’s own terms and found a lag; a chord-based model would not have one. The histogram model cannot represent that, and the joint search cannot rescue it, because a joint search over stages one and two and a stage three that is blind to both is still blind.

The order of the stages is itself a claim

There is a version of the pipeline nobody runs and it is worth naming, because its absence is informative.

Nothing in the three stages forces the order metre-then-chords-then-key. The key could be taken first — it is the one stage whose input needs neither of the others — and then used to restrict which chords are candidates, and only then the metre chosen to suit. That pipeline is a different pipeline and it gives different answers, and it is not obviously worse.

The order that is used is the order in which the evidence is cheapest, not the order in which it is strongest. Onsets are the easiest thing to extract from a signal, pitch classes the next, and chords a conclusion rather than an observation — so the pipeline is a processing convenience that has hardened into an account of what a listener does.

The joint search has no order at all, which is its one unambiguous virtue. It is also why its cost is a product rather than a sum: an ordered pipeline pays for its stages one after another and an unordered search pays for all of them at once.

What would have to change

The fix is not a bigger search. It is a different third stage: a key model whose evidence is a sequence of chords rather than a bag of notes.

That model exists in the literature — it is what a grammar or a hidden Markov model over harmonic function does — and it is a different object from anything on this site, because every method here begins by discarding order. The cost is the same trade the pair model paid when it went from twenty-four hypotheses to three hundred: a model that can distinguish more must carry more hypotheses, and the number grows with what it wants to notice.

What can be said now, from the arithmetic above, is how much of the problem such a model would inherit. The chord and the barline are already entangled at thirty-nine per cent, so a key model that fed back into them would be joining a search that is already unstable in two dimensions, and its own answers would move the other two around in turn. The joint problem is not three loosely coupled problems; it is one problem, and this rung has drawn two-thirds of it.

Which computation produced the numbers

Three functions, none of them new.

metreFit scores a barline position: an onset on a strong position counts three, a strong position with nothing on it counts minus two, an onset off the grid counts minus one. Rotating the onset list is how the barline is moved, and the scores are normalised across the positions available in a given passage so that the product is meaningful.

segmentFit scores a chord: coverage times parsimony, where coverage is the share of the metrically weighted note mass the chord accounts for and parsimony is the share of the chord’s own notes that were sounded. The parsimony factor exists because coverage alone picks the diminished seventh every time.

keyCorrelations scores a key: Pearson’s r between the pitch-class histogram and each of twenty-four rotated Krumhansl–Kessler profiles.

The one new number is the key’s veto, and it is a stated parameter rather than a measurement: a chord whose root is outside the winning key’s scale has its total multiplied by 0.6. Any value below one produces the same qualitative result and a different disagreement rate; at 1.0 the key term does nothing but scale, and the disagreement rate falls to the metre-versus-chord trade alone. The essay says so because the number is doing work.

The passages are constructed rather than sampled from music: eight slots, each carrying a note with probability 0.62, pitches drawn uniformly from one diatonic collection, seeded so the figure is the same every time it is drawn. That is a strong assumption about the material and the last section returns to it.

Whose music, and when

The pipeline this essay takes apart is not a straw man; it is how harmonic analysis is taught and how most software does it. A student is told the metre, given the barlines on the page, and asked what the chords are — which is stages two and three with stage one supplied by the notation. A time signature is a claim and a barline is that claim written down, so notated music hands the analyst the first decision for free, in the way a notation always fixes some coordinates and leaves the rest, and hides the fact that it was a decision.

Where the repertoire makes it visible is exactly where the notation is contested: a hemiola, where the page says three and the accents say six; a syncopated passage whose barlines a performer overrides; unbarred plainchant and the unmeasured prelude, where nothing is supplied at all. In those places the analyst is doing the joint problem by hand, and the disagreement rate above is a rough measure of how often the answer depends on which order the decisions are taken in.

What the picture cannot show

The passages are not music. Random notes from a diatonic collection have no voice leading, no repetition and no cadences, so the disagreement rate of thirty-nine per cent is a rate over a population no composer has ever written in. What it measures is how loosely the two scores constrain each other in general, not how often a real analysis would change. The schemes this site carries would be the honest population and they are twelve.

One chord per passage. The search finds the single best chord for eight slots, where a real passage has a chord per bar or per half-bar, and the joint problem over a sequence of chords is combinatorially much worse and structurally different — the chords constrain each other too.

The product is a choice. Multiplying three scores treats them as independent likelihoods, and they are not: a barline that fits the onsets well tends to give a chord that fits the notes well, so the product double-counts the agreement between them.

And the key’s veto is a parameter, not a model. A proper joint model would have the key and the chord in one probability, and the 0.6 above is a stand-in for a relation nobody here has estimated.

Where this ladder goes next

Five rungs. A progression is a path across a space with a geometry; it never comes home in just intonation; its rate is a variable with limits set outside music; the objects it connects are produced by a segmentation that depends on the metre; and now, taking the three decisions together changes two readings in five and cannot touch the third.

The rung after it is the one the zero demands. Every model in this ladder infers a key from a bag of notes, and the thing that most obviously declares a key is a cadence — an ordered pair of chords, in which the order is the whole content. A model that scored the sequence rather than the bag could be built from parts this site already has: the closure ladder’s five-component cadence vector is exactly an ordered-pair measurement, and it has never been used as evidence for a key. Whether it beats the histogram, and by how much, is a computation — and the interesting case is the one this rung could not construct, where the notes say one key and the cadences say another. An ending is five components with no total, so the evidence would arrive as a vector rather than a number, which is a harder thing to put into a search and a more honest one.

Part 5 of 17

One essay in the series on progression. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 9.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AmbiguityChord toneHarmonic analysisInferenceKey-findingMetreMetrical weightSegmentation