Three decisions that constrain each other
Assumes: Which notes are the chord · The time signature is a claim
The previous rung put the other three on a footing and left one thing owed. A progression is a list of chords; before there is a list, something has to decide which of the sounding notes are chord tones and which are passing; and the thing that decides is the metre. So harmonic analysis is a function of a variable that is not harmony.
The rung ended by naming the joint problem. Every model on this site decides one thing at a time — the key from the pitch classes, the metre from the onsets, the chords from the metre — and a listener does all three at once, with each constraining the others. A model that resolved them jointly would be a considerably better model and a considerably harder one to draw.
It is one of those things. Which one is the rung.
The pipeline, written out
The sequence every model on this site follows, when it is followed all the way through, is three stages and each one throws away what the next might have wanted.
Stage one, the metre. Take the onsets — just the times, no pitches — and score each candidate barline position by how well a periodic accent pattern explains the onsets. That is metreFit, which the rhythm ladder has used since the metre rung, and which is the model a time signature can be checked against.
Stage two, the chords. Weight each note by the metrical position the winning barline gives it, then score every chord at every root by how much of the weighted note mass it accounts for and how much of itself was actually sounded. That is segmentFit, and it is what the previous rung is built on.
Stage three, the key. Build a pitch-class histogram and correlate it against twenty-four rotated probe-tone profiles. That is the key-finder this ladder’s neighbour has been running for four rungs.
Each stage commits. The metre is chosen before any pitch has been looked at; the chord is chosen inside a metre that cannot now be revised; and the key arrives last, with no power to object to either.
The joint search, and what it costs
The joint version uses the same three scoring functions and changes nothing about them. What changes is that every combination is scored together and the best total wins.
The total is the product of the three, not the sum, and the choice is deliberate: a product refuses a hypothesis that is bad on any one of the three, which is what “constraining each other” means. A sum lets a triple with an excellent chord and an impossible metre outscore a triple that is good at everything.
The cost is a factor of a hundred and fifty-seven. The pipeline evaluates eight barline positions, then a hundred and forty-four chord-and-root pairs, then twenty-four keys: a hundred and seventy-six evaluations, because each stage’s answer is fixed before the next begins. The joint search evaluates their product: eight times a hundred and forty-four times twenty-four, which is twenty-seven thousand six hundred and forty-eight.
That is the “considerably harder” made into a number, and it is the mild version. The passage here is eight slots long, and a model that also searched over bar lengths, over subdivision trees, and over sequences of chords rather than one chord per window would multiply again by each — by roughly three, four and a hundred and forty-four:
| what is searched | hypotheses | against the pipeline |
|---|---|---|
| the pipeline | 176 | — |
| the joint search here | 27,648 | ×157 |
| plus three bar lengths | 82,944 | ×471 |
| plus four subdivision trees | 331,776 | ×1,885 |
| plus a chord per half-bar | 47,775,744 | ×271,453 |
The last row is the honest size of the thing this rung is a miniature of, and it is why nobody writes the joint model. A quarter of a million times the pipeline’s work, for eight quavers.
The product, which turns out not to matter
The choice of a product over a sum was argued for above and never tested, and testing it is one line: score every hypothesis both ways and count the disagreements.
On 552 passages the two rules choose the same triple every time. Not usually — every time, zero disagreements. The argument for the product is sound and the passages this rung is built on do not contain a case where it bites, because a hypothesis bad enough on one term to need vetoing is already losing on the other two.
That is worth recording as a null rather than quietly dropping, because it says where the argument’s force actually is. The product matters when one term can be near zero while the others are excellent, and in this construction the metre term is nearly flat — the same flatness that lets the chord overturn the barline — so nothing ever gets near zero on it.
Which term does the overturning
Ablating the joint search’s own terms, against its full product on the same passages, says how the disagreement with the pipeline is divided. Dropping the key term changes the joint answer in 12 per cent of passages and dropping the metre term in 16. So of the disagreement the rung reports, the key’s veto accounts for about a quarter and the metre’s own contribution for about a third, with the chord term — which both models share — doing the rest.
That is a more useful decomposition than the headline, because it prices the one thing the pipeline structurally cannot do. The key term can never change the key, and it changes the chord in one passage in eight.
Where they disagree
On five hundred and fifty-two constructed passages — eight slots, some of them rests, notes drawn from one diatonic collection — the two readings differ in thirty-nine per cent of cases.
The chord differs in twenty-nine per cent. The barline differs in thirty-one. Both differ in twenty.
The mechanism is visible in the individual cases and it runs in the direction the previous rung predicted. A passage whose onsets weakly prefer one barline gets that barline from the pipeline, and the chord it then finds is whatever those metrical weights favour. The joint search will trade a little of the metre’s score for a great deal of the chord’s, and the trade is often available because the metre score is nearly flat over barline positions in any passage with notes on most of its slots — which is most music.
The commonest single mechanism is the key vetoing a chord. The pipeline’s third stage has no power to reject; the joint search’s key term multiplies a chord’s score down when the chord’s root is outside the key it names. In one worked case the pipeline reads F major seventh and the joint search reads A minor seventh, on the same notes at the same barline, because the key is E minor and F natural is not in it.
Both of those readings run on the same eight numbers, so it is worth seeing the weights themselves before asking which stage should be allowed to overturn which.
Before the three are put together it is worth seeing the grid that the second of them runs on, because every weight in this essay descends from it.
Nothing about the notes distinguishes them here — they are eight identical events — so every difference in the drawing is supplied by the metre and none of it by the music. That is the sense in which the segmentation decision depends on the metre decision, and it is why the two cannot be made in either order without the other already being made.
The zero
The number that is worth more than the other three is the one that is not there.
The key never changes. Not rarely: never, in five hundred and fifty-two passages.
The reason is structural and it is the thing the whole ladder should be read through. The key is inferred from a pitch-class histogram, and a histogram is invariant under everything the other two decisions are about. Move the barline and the histogram is unchanged, because a histogram has no bar in it. Choose a different chord and the histogram is unchanged, because the chord is a conclusion about the notes and not an input.
So there is nothing for the joint search to do to the key. Two of the three decisions are entangled with each other and the third is coupled to them in one direction only: the key can constrain the chord, and the chord cannot inform the key.
That is a limitation of the model and not a fact about hearing. A listener plainly does use the chords to decide the key — a single V–I settles a key faster than a bag of its pitch classes ever could, and the whole apparatus of tonicisation is about chords declaring a key. How much evidence a modulation needs measured that in the histogram’s own terms and found a lag; a chord-based model would not have one. The histogram model cannot represent that, and the joint search cannot rescue it, because a joint search over stages one and two and a stage three that is blind to both is still blind.
The order of the stages is itself a claim
There is a version of the pipeline nobody runs and it is worth naming, because its absence is informative.
Nothing in the three stages forces the order metre-then-chords-then-key. The key could be taken first — it is the one stage whose input needs neither of the others — and then used to restrict which chords are candidates, and only then the metre chosen to suit. That pipeline is a different pipeline and it gives different answers, and it is not obviously worse.
The order that is used is the order in which the evidence is cheapest, not the order in which it is strongest. Onsets are the easiest thing to extract from a signal, pitch classes the next, and chords a conclusion rather than an observation — so the pipeline is a processing convenience that has hardened into an account of what a listener does.
The joint search has no order at all, which is its one unambiguous virtue. It is also why its cost is a product rather than a sum: an ordered pipeline pays for its stages one after another and an unordered search pays for all of them at once.
What would have to change
The fix is not a bigger search. It is a different third stage: a key model whose evidence is a sequence of chords rather than a bag of notes.
That model exists in the literature — it is what a grammar or a hidden Markov model over harmonic function does — and it is a different object from anything on this site, because every method here begins by discarding order. The cost is the same trade the pair model paid when it went from twenty-four hypotheses to three hundred: a model that can distinguish more must carry more hypotheses, and the number grows with what it wants to notice.
What can be said now, from the arithmetic above, is how much of the problem such a model would inherit. The chord and the barline are already entangled at thirty-nine per cent, so a key model that fed back into them would be joining a search that is already unstable in two dimensions, and its own answers would move the other two around in turn. The joint problem is not three loosely coupled problems; it is one problem, and this rung has drawn two-thirds of it.
Which computation produced the numbers
Three functions, none of them new.
metreFit scores a barline position: an onset on a strong position counts three, a strong position with nothing on it counts minus two, an onset off the grid counts minus one. Rotating the onset list is how the barline is moved, and the scores are normalised across the positions available in a given passage so that the product is meaningful.
segmentFit scores a chord: coverage times parsimony, where coverage is the share of the metrically weighted note mass the chord accounts for and parsimony is the share of the chord’s own notes that were sounded. The parsimony factor exists because coverage alone picks the diminished seventh every time.
keyCorrelations scores a key: Pearson’s r between the pitch-class histogram and each of twenty-four rotated Krumhansl–Kessler profiles.
The one new number is the key’s veto, and it is a stated parameter rather than a measurement: a chord whose root is outside the winning key’s scale has its total multiplied by 0.6. Any value below one produces the same qualitative result and a different disagreement rate; at 1.0 the key term does nothing but scale, and the disagreement rate falls to the metre-versus-chord trade alone. The essay says so because the number is doing work.
The passages are constructed rather than sampled from music: eight slots, each carrying a note with probability 0.62, pitches drawn uniformly from one diatonic collection, seeded so the figure is the same every time it is drawn. That is a strong assumption about the material and the last section returns to it.
Whose music, and when
The pipeline this essay takes apart is not a straw man; it is how harmonic analysis is taught and how most software does it. A student is told the metre, given the barlines on the page, and asked what the chords are — which is stages two and three with stage one supplied by the notation. A time signature is a claim and a barline is that claim written down, so notated music hands the analyst the first decision for free, in the way a notation always fixes some coordinates and leaves the rest, and hides the fact that it was a decision.
Where the repertoire makes it visible is exactly where the notation is contested: a hemiola, where the page says three and the accents say six; a syncopated passage whose barlines a performer overrides; unbarred plainchant and the unmeasured prelude, where nothing is supplied at all. In those places the analyst is doing the joint problem by hand, and the disagreement rate above is a rough measure of how often the answer depends on which order the decisions are taken in.
What the picture cannot show
The passages are not music. Random notes from a diatonic collection have no voice leading, no repetition and no cadences, so the disagreement rate of thirty-nine per cent is a rate over a population no composer has ever written in. What it measures is how loosely the two scores constrain each other in general, not how often a real analysis would change. The schemes this site carries would be the honest population and they are twelve.
One chord per passage. The search finds the single best chord for eight slots, where a real passage has a chord per bar or per half-bar, and the joint problem over a sequence of chords is combinatorially much worse and structurally different — the chords constrain each other too.
The product is a choice. Multiplying three scores treats them as independent likelihoods, and they are not: a barline that fits the onsets well tends to give a chord that fits the notes well, so the product double-counts the agreement between them.
And the key’s veto is a parameter, not a model. A proper joint model would have the key and the chord in one probability, and the 0.6 above is a stand-in for a relation nobody here has estimated.
Where this ladder goes next
Five rungs. A progression is a path across a space with a geometry; it never comes home in just intonation; its rate is a variable with limits set outside music; the objects it connects are produced by a segmentation that depends on the metre; and now, taking the three decisions together changes two readings in five and cannot touch the third.
The rung after it is the one the zero demands. Every model in this ladder infers a key from a bag of notes, and the thing that most obviously declares a key is a cadence — an ordered pair of chords, in which the order is the whole content. A model that scored the sequence rather than the bag could be built from parts this site already has: the closure ladder’s five-component cadence vector is exactly an ordered-pair measurement, and it has never been used as evidence for a key. Whether it beats the histogram, and by how much, is a computation — and the interesting case is the one this rung could not construct, where the notes say one key and the cadences say another. An ending is five components with no total, so the evidence would arrive as a vector rather than a number, which is a harder thing to put into a search and a more honest one.
Part 5 of 17
One essay in the series on progression. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 9.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
AmbiguityChord toneHarmonic analysisInferenceKey-findingMetreMetrical weightSegmentation
- A count and a correlation harmonic analysis, key-finding, segmentation
- A modulation and a borrowing are one number apart harmonic analysis, key-finding, segmentation
- How much of the reading arrives late ambiguity, inference, key-finding
- The chords are a weak witness to the barline harmonic analysis, metre, segmentation
- The margin the dynamic program already had ambiguity, key-finding, segmentation
- A bass line is not a list of roots inference, key-finding