The repeat that is not in the notes
Assumes: A piece is mostly itself again · The boundary is where the neighbourhood changes
Every picture in this ladder is a square of cells, and every cell is the same arithmetic: write each bar as a twelve-dimensional vector with the chord’s notes at one and the rest of the key at a fraction, and take the cosine between two of them.
That vector was chosen once. The first rung needed a way to compare two bars, picked one that worked, and explained in careful detail why the key had to be in it — with chord tones alone the matrix goes to a two-tone chequerboard and there is nothing to read. Every rung since has inherited the choice. The boundary operator, the compression ratio, the recurrence period, the causal version, the memory discount and the threshold that decides what a repeat is are all statements about that vector, and the ladder has never asked which of them are also statements about music.
A bag of pitch classes is not the only thing a bar can be. It is one of five the machinery already here supports.
Four other things a bar can be
The chord’s own notes alone. The same vector with the key weight at zero — what the first rung rejected, kept here because rejecting it was a decision about legibility rather than about correctness.
Tonic, subdominant or dominant. Three dimensions instead of twelve, plus the bar’s own key beside them at the same weight the pitch-class vector gives its key. This throws away almost everything: a I and a vi become the same object, and so do a IV and a ii.
How far the root moved. A twelve-dimensional one-hot on the interval from the previous bar’s root, and nothing else. A bar has no identity at all in this encoding — only a relationship to the bar before it. Two passages match when their root motions match, wherever they are and in whatever key.
The motion between the chords. This one is not a feature vector, and that is why it is here. The cheapest total voice motion from one chord to another is a property of the pair: there is no map from a chord to a point in a space whose distances are those costs. It goes straight into the matrix as a similarity, normalised by the largest cost the piece contains.
What the encoding was hiding
At a key weight of 0.7, two bars in the same key already agree on seven of their twelve dimensions before either chord is looked at. The cosine between any two bars of a piece that stays in one key is therefore high, and the differences between chords live in the last two decimal places.
The mean off-diagonal similarity puts a number on that saturation and it moves fast: 0.49 at a key weight of zero, 0.67 at 0.2, 0.87 at 0.5 and 0.94 at 0.7. So the last two tenths of the dial spend a quarter of the available range of the measure, and everything above about 0.5 is compressing the whole piece into the top six per cent of a similarity axis. A cosine has no more resolution up there than anywhere else; what it has is fewer distinct values in play.
Three of the six schemes behave this way. The thirty-two-bar song, the verse-chorus and the sixteen-bar period all sit above 0.93 mean similarity under the pitch-class encoding, and the boundary operator finds nothing at all in any of them — not a boundary in the wrong place, no boundaries. Under the function encoding it finds every one in two of the three and two of three in the last.
That is not a small correction to a rung; it is a different verdict on the operator. The rung that built it reported that it worked and that what it could not find said more than what it could. The finding stands and its cause moves: what it could not find was not a limit of local novelty detection, it was a saturated similarity measure.
It is worth separating those two verdicts carefully, because the difference between them is a difference in what the collection should do next. The operator has a limit is a fact to be recorded and worked around. The operator was starved of contrast is a defect to be repaired, and the repair is a number rather than a redesign. The rung’s own conclusion — that the absences were informative — turns out to have been informative about the measure and not about the music, which is the harder of the two things to notice from inside.
What survives, and it is one thing
Run all three of the ladder’s measurements under all five encodings on all six schemes, and they separate cleanly.
The recurrence period survives. The strongest peak of the lag profile is the same number under every encoding on five of the six schemes — twelve bars for the blues, sixteen for the rondo and the verse-chorus, eight for the sixteen-bar period, one for the ostinato. The only disagreement is the thirty-two-bar song, where the voice-leading encoding reports sixteen and the other four report four. So the rung that found how long until it comes back found something about the music: the diagonal sums do not care what a bar is made of, only that bars a fixed distance apart resemble each other in whatever the measure is.
The block structure survives with different numbers. Every encoding puts blocks in the same places; what changes is the contrast. Even root motion, which knows nothing about any bar, produces the sections — because a section is partly a pattern of moves.
The boundary operator does not survive. Its recall runs from 0 to 1 across encodings on the same piece, and the ordering is consistent. Averaged over the five schemes that have boundaries: 0.93 for function, 0.83 for chord tones and for voice-leading motion, 0.50 for root motion, and 0.18 for the bag of pitch classes every rung used. A local operator that measures contrast is only as good as the contrast it is given, and this one was given almost none.
Which raises a question the five-way comparison does not: the pitch-class encoding and the chord-tone encoding are the same vector with a dial between them, so what does the dial do? Sweeping it:
| key weight | mean off-diagonal similarity | boundary recall |
|---|---|---|
| 0 | 0.49 | 0.83 |
| 0.2 | 0.67 | 0.78 |
| 0.3 | 0.75 | 0.68 |
| 0.5 | 0.87 | 0.28 |
| 0.7 | 0.94 | 0.17 |
The recall is a monotone function of one number, and the ladder’s setting is at the far end of it. There is no boundary between two encodings here: 0.83 at weight zero is the chord-tone row, 0.17 at 0.7 is the pitch-class row, and everything between is available. Anywhere at or below 0.3 recovers two-thirds or more of the boundaries.
So the honest description of what went wrong is narrower and more embarrassing than the encoding was inherited. The encoding was not inherited; a parameter was, and it was set to 0.7 by a rung that needed the matrix picture to be legible — its own stated reason — and the setting that makes the picture legible is the setting that saturates the similarity and blinds the operator. The two requirements are in direct opposition, and nothing in the ladder said so because nothing in the ladder had both requirements in front of it at once.
That also means the repair is available without changing any machinery. The boundary operator wants a low key weight and the matrix picture wants a high one, and they are different figures; nothing requires them to share a value.
The ordering also holds at other kernel widths. On the rondo at kernels of two, four and eight bars the function encoding finds all four boundaries every time; the pitch-class encoding finds three, two and two. So this is not an artefact of the width the operator’s own rung settled on.
Root motion, which is a different kind of encoding
The root-motion feature is worth its own look, because it is the only one of the five in which a bar has no identity.
This is the encoding that answers a question the transposition rung put and could not fully settle. That rung built a dial between an absolute reading and a key-relative one, and found the trade it was supposed to make was not the trade it made. Root motion is the limit of that dial taken all the way: a passage returning a fifth higher is identical here, by construction, and the cost is that a passage returning with the same roots and different chord qualities is identical too.
The verse-chorus is where that pays. Its two sections are built from the same four chords in different rotations, so the pitch-class encoding cannot separate them at all — and root motion finds all three boundaries at 50 per cent precision, because the order of the moves differs even though the chords do not.
The case every measure has to fail on
The ostinato is in the corpus as the control: four bars, eight times, nothing else. A measure of structure that reports structure here is broken.
Every encoding agreeing on the degenerate case and disagreeing on the real ones is the right shape for a robustness check. A measure that behaved differently on the ostinato would be reporting a property of its own arithmetic.
The measurement that cannot be compared at all
There is a fourth of the ladder’s measurements and it behaves in a fourth way, which is worth a section because the reason is instructive.
The compression rung put a number on how much of a piece is a repeat of an earlier part of itself by feeding the bars to a coder and counting the bits. To run it under a different encoding, the bars have to become symbols, and each encoding supplies its own alphabet: the rondo has ten distinct bar symbols as pitch-class-and-key, seven as bare roman numerals, nine as function-and-key, and six as root motions.
The ratios that come out are 1.135, 1.108, 1.041 and 0.958. Read naively that says the rondo is most repetitive as root motion — and it says nothing of the kind, because a coder with an alphabet of six needs fewer bits per symbol than one with an alphabet of ten before it has compressed anything. That rung said so itself, in as many words: the absolute number is a property of the coder and the alphabet, and only a ratio between two sequences encoded the same way is a statistic.
So the compression measure is not robust and is not fragile either. It is undefined across encodings, and its own rung had already said why. That is a third kind of answer to the question this rung asks, and finding it is the reason to ask the question of every measurement rather than of the interesting ones.
Which computation produced the numbers
The matrices are the same similarityMatrix the ladder has always used, over different feature vectors; the voice-leading matrix is built directly from the cheapest total motion between the two chords, divided by the largest such cost in the piece and subtracted from one.
The boundary scores are Foote’s novelty at a kernel of four bars, with peaks taken above 0.03, scored against the section letters in the scheme’s own definition with a tolerance of one bar. Those are the settings the rung that built the operator used, kept deliberately: the comparison here is between encodings and not between settings.
The recurrence periods are read at a key weight of 0.18 rather than 0.7, for the same reason — that is the weight the lag rung used, and it had already established the weight as a setting worth naming.
One defect turned up in the machinery while this was being computed and is worth recording. The site’s voiceLeading minimises total motion over every assignment of voices and requires the two chords to be the same size, which its docstring said and nothing enforced: hand it a triad and a seventh chord and it indexed past the end of the shorter one, and returned a cost of NaN. Every existing caller passes equal sizes, so no figure was wrong. It now handles unequal sizes by doubling — the smaller chord lends a note to a second voice, minimised over which note, which is what a realisation does — and a triad to a dominant seventh comes out at four semitones.
Whose music, and when
The six schemes are conventions rather than pieces, and that has not changed. A twelve-bar blues is a family with named variants; a thirty-two-bar song form is a plan; the rondo’s letters are a shape rather than a transcription. Everything here is a claim about what an analysis recovers from a plan, not about what it recovers from a recording.
What does travel is the methodological half, and it travels well beyond music. A local contrast detector inherits the dynamic range of its similarity measure. Give it a measure that saturates and it will find nothing and report that the material has no boundaries — which is exactly what happened here, in three schemes out of six, for eight rungs.
What the picture cannot show
None of the five encodings has any rhythm in it. Every bar is one chord and every bar is the same length. A passage that returns with the same harmony and a different rhythm is a repeat that none of these matrices can see, and a passage that returns with the same rhythm and different harmony is invisible to all five. That is the largest thing missing, it is missing from this rung as much as from the eight before it, and the corpus is what makes it missing: these schemes carry a roman numeral per bar and nothing else.
The function assignment is a table. Applied dominants are called dominants, the diminished triad on the seventh degree is called a dominant, and the mediant is called a tonic — the ordinary functional readings, which are contested at the edges and which the encoding does not derive from anything.
Five encodings is not the space of encodings. Registral position, melodic contour, texture and dynamics are all things a bar has and none of them is here, because the schemes do not carry them.
And the boundary scores are at one kernel width. The operator’s own rung swept three, and a different width changes the recall figures — not the ordering across encodings, which was checked at two and eight as well, but the numbers.
The ladder ends here
repetition closes at nine rungs. The model is a piece as a sequence of comparable units and a form as the pattern of resemblances between them, and the ladder now bounds it in every variable it has: what the matrix is (one), what operator reads an edge out of it (two), what number can be put on the whole (three), what the diagonals say (four), what a listener has available at the time (five), how transposition is handled (six), how memory discounts it (seven), where a repeat stops being one (eight), and now what a unit is encoded as.
Name a variable the model has, and it is on that list — with one exception that is named rather than hidden: the unit is a bar in all nine, and it is a bar because the corpus is written in bars. That is a property of the data rather than a free parameter of the model, and it is the reason the closure is honest.
What is not on the list belongs elsewhere. Where a phrase ends rather than a section is phrase; what makes an ending an ending is closure; and the reason the key plan shows up in these matrices at all is key-relations.
Part 9 of 9
One essay in the series on repetition. The essays either side of this one:
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Description lengthNoveltyRepetitionSection boundarySegmentationSelf-similarityTonal functionVoice-leading
- An ending that can be heard coming description length, novelty, section boundary, segmentation
- A cycle cannot cadence novelty, repetition
- A return is shorter than its first hearing description length, repetition
- The margin the dynamic program already had segmentation, tonal function
- The passage built to make them disagree segmentation, tonal function
- The resolution the metric cannot find tonal function, voice-leading