Form and structure

The repeat that is not in the notes

Eight earlier essays compare bars by writing each one as a bag of pitch classes and taking a cosine. Nothing chose that encoding — the first used it and the other seven inherited it. Encode the same six schemes four other ways and one of the three findings survives untouched, one survives with different numbers, and one turns out to have been a statement about the encoding all along: the boundary operator finds none of the section edges in three schemes as a bag of pitch classes and every one of them as tonic, subdominant and dominant.

Assumes: A piece is mostly itself again · The boundary is where the neighbourhood changes

Every picture in this ladder is a square of cells, and every cell is the same arithmetic: write each bar as a twelve-dimensional vector with the chord’s notes at one and the rest of the key at a fraction, and take the cosine between two of them.

That vector was chosen once. The first rung needed a way to compare two bars, picked one that worked, and explained in careful detail why the key had to be in it — with chord tones alone the matrix goes to a two-tone chequerboard and there is nothing to read. Every rung since has inherited the choice. The boundary operator, the compression ratio, the recurrence period, the causal version, the memory discount and the threshold that decides what a repeat is are all statements about that vector, and the ladder has never asked which of them are also statements about music.

A bag of pitch classes is not the only thing a bar can be. It is one of five the machinery already here supports.

Four other things a bar can be

The chord’s own notes alone. The same vector with the key weight at zero — what the first rung rejected, kept here because rejecting it was a decision about legibility rather than about correctness.

Tonic, subdominant or dominant. Three dimensions instead of twelve, plus the bar’s own key beside them at the same weight the pitch-class vector gives its key. This throws away almost everything: a I and a vi become the same object, and so do a IV and a ii.

How far the root moved. A twelve-dimensional one-hot on the interval from the previous bar’s root, and nothing else. A bar has no identity at all in this encoding — only a relationship to the bar before it. Two passages match when their root motions match, wherever they are and in whatever key.

The motion between the chords. This one is not a feature vector, and that is why it is here. The cheapest total voice motion from one chord to another is a property of the pair: there is no map from a chord to a point in a space whose distances are those costs. It goes straight into the matrix as a similarity, normalised by the largest cost the piece contains.

verse and chorus, every bar against every other bar. A self-similarity matrix of 32 bars of verse and chorus. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.7. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 1 Verse and chorus as it has always been drawn: bars as bags of pitch classes, key weight 0.7. Almost every cell is above the drawing threshold, which is why the figure is nearly solid. The mean similarity between bars more than two apart is 0.97.
verse and chorus, encoded as tonic, subdominant or dominant. A self-similarity matrix of 32 bars of verse and chorus. Each cell is the cosine similarity of two bars encoded as tonic, subdominant or dominant, in place of the bag of pitch classes every other figure uses, with the bar's own key weighted 0.7 beside it. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 2 The same thirty-two bars as tonic, subdominant and dominant. Mean off-diagonal similarity 0.56, and the eight-bar blocks are visible as blocks rather than as a suggestion. Nothing has been added; three-quarters of the information has been thrown away, and the picture is better.

What the encoding was hiding

At a key weight of 0.7, two bars in the same key already agree on seven of their twelve dimensions before either chord is looked at. The cosine between any two bars of a piece that stays in one key is therefore high, and the differences between chords live in the last two decimal places.

The mean off-diagonal similarity puts a number on that saturation and it moves fast: 0.49 at a key weight of zero, 0.67 at 0.2, 0.87 at 0.5 and 0.94 at 0.7. So the last two tenths of the dial spend a quarter of the available range of the measure, and everything above about 0.5 is compressing the whole piece into the top six per cent of a similarity axis. A cosine has no more resolution up there than anywhere else; what it has is fewer distinct values in play.

What survives a change of encoding: verse and chorus. The same 32 bars of verse and chorus under five encodings, scored on the three things measured here measures. Mean off-diagonal similarity says how alike the piece looks to the arithmetic. Recall and precision are the novelty operator's boundaries against the 3 the section plan has, at a kernel of four bars. The period is the strongest peak of the lag profile, in bars. Under the bag of pitch classes every other figure uses, the piece is 90 per cent self-similar and the operator finds 0 per cent of the boundaries; under how far the root moved it finds 100 per cent. The period is the quantity that does not move.
Fig. 3 Verse and chorus under all five encodings. The bag of pitch classes reports the piece as 97 per cent self-similar and finds none of the three section boundaries. The chord-tone and function encodings find two of three, and root motion finds all three. The recurrence period is sixteen bars under every one of the five.

Three of the six schemes behave this way. The thirty-two-bar song, the verse-chorus and the sixteen-bar period all sit above 0.93 mean similarity under the pitch-class encoding, and the boundary operator finds nothing at all in any of them — not a boundary in the wrong place, no boundaries. Under the function encoding it finds every one in two of the three and two of three in the last.

That is not a small correction to a rung; it is a different verdict on the operator. The rung that built it reported that it worked and that what it could not find said more than what it could. The finding stands and its cause moves: what it could not find was not a limit of local novelty detection, it was a saturated similarity measure.

It is worth separating those two verdicts carefully, because the difference between them is a difference in what the collection should do next. The operator has a limit is a fact to be recorded and worked around. The operator was starved of contrast is a defect to be repaired, and the repair is a number rather than a redesign. The rung’s own conclusion — that the absences were informative — turns out to have been informative about the measure and not about the music, which is the harder of the two things to notice from inside.

What survives a change of encoding: twelve-bar blues. The same 36 bars of twelve-bar blues under five encodings, scored on the three things measured here measures. Mean off-diagonal similarity says how alike the piece looks to the arithmetic. Recall and precision are the novelty operator's boundaries against the 8 the section plan has, at a kernel of four bars. The period is the strongest peak of the lag profile, in bars. Under the bag of pitch classes every other figure uses, the piece is 85 per cent self-similar and the operator finds 63 per cent of the boundaries; under the chord's own notes alone it finds 100 per cent. The period is the quantity that does not move.
Fig. 4 The twelve-bar blues, where the pitch-class encoding does better — 38 per cent of the eight boundaries, at perfect precision — because the blues changes key-foreign notes at every section and the vector can see them. Chord tones, function and voice-leading motion each find all eight, at perfect precision. Root motion finds two.

What survives, and it is one thing

Run all three of the ladder’s measurements under all five encodings on all six schemes, and they separate cleanly.

The recurrence period survives. The strongest peak of the lag profile is the same number under every encoding on five of the six schemes — twelve bars for the blues, sixteen for the rondo and the verse-chorus, eight for the sixteen-bar period, one for the ostinato. The only disagreement is the thirty-two-bar song, where the voice-leading encoding reports sixteen and the other four report four. So the rung that found how long until it comes back found something about the music: the diagonal sums do not care what a bar is made of, only that bars a fixed distance apart resemble each other in whatever the measure is.

The block structure survives with different numbers. Every encoding puts blocks in the same places; what changes is the contrast. Even root motion, which knows nothing about any bar, produces the sections — because a section is partly a pattern of moves.

The boundary operator does not survive. Its recall runs from 0 to 1 across encodings on the same piece, and the ordering is consistent. Averaged over the five schemes that have boundaries: 0.93 for function, 0.83 for chord tones and for voice-leading motion, 0.50 for root motion, and 0.18 for the bag of pitch classes every rung used. A local operator that measures contrast is only as good as the contrast it is given, and this one was given almost none.

Which raises a question the five-way comparison does not: the pitch-class encoding and the chord-tone encoding are the same vector with a dial between them, so what does the dial do? Sweeping it:

key weight mean off-diagonal similarity boundary recall
0 0.49 0.83
0.2 0.67 0.78
0.3 0.75 0.68
0.5 0.87 0.28
0.7 0.94 0.17

The recall is a monotone function of one number, and the ladder’s setting is at the far end of it. There is no boundary between two encodings here: 0.83 at weight zero is the chord-tone row, 0.17 at 0.7 is the pitch-class row, and everything between is available. Anywhere at or below 0.3 recovers two-thirds or more of the boundaries.

So the honest description of what went wrong is narrower and more embarrassing than the encoding was inherited. The encoding was not inherited; a parameter was, and it was set to 0.7 by a rung that needed the matrix picture to be legible — its own stated reason — and the setting that makes the picture legible is the setting that saturates the similarity and blinds the operator. The two requirements are in direct opposition, and nothing in the ladder said so because nothing in the ladder had both requirements in front of it at once.

That also means the repair is available without changing any machinery. The boundary operator wants a low key weight and the matrix picture wants a high one, and they are different figures; nothing requires them to share a value.

The ordering also holds at other kernel widths. On the rondo at kernels of two, four and eight bars the function encoding finds all four boundaries every time; the pitch-class encoding finds three, two and two. So this is not an artefact of the width the operator’s own rung settled on.

How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 6 cases. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.483 and the smallest 0.000. 4 of 5 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.
Fig. 5 The measurement that does not move. Each scheme’s lag profile — every diagonal of the matrix summed and plotted against how far off the main diagonal it sits — reports a period, and it is the same period under all five encodings for five of the six. This is the most robust result of all and nothing had said so, because nothing had varied the thing that could have broken it.

Root motion, which is a different kind of encoding

The root-motion feature is worth its own look, because it is the only one of the five in which a bar has no identity.

rondo, ABACA, encoded as how far the root moved. A self-similarity matrix of 40 bars of rondo, ABACA. Each cell is the cosine similarity of two bars encoded as how far the root moved, in place of the bag of pitch classes every other figure uses. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 6 The rondo as root motion. Every cell compares two moves rather than two bars, so the diagonal blocks are places where the same succession of intervals happens twice — which is what a returning refrain is, and is also what a sequence is. The refrain’s three appearances are the three blocks; the two episodes are not blank but are made of different moves.

This is the encoding that answers a question the transposition rung put and could not fully settle. That rung built a dial between an absolute reading and a key-relative one, and found the trade it was supposed to make was not the trade it made. Root motion is the limit of that dial taken all the way: a passage returning a fifth higher is identical here, by construction, and the cost is that a passage returning with the same roots and different chord qualities is identical too.

The verse-chorus is where that pays. Its two sections are built from the same four chords in different rotations, so the pitch-class encoding cannot separate them at all — and root motion finds all three boundaries at 50 per cent precision, because the order of the moves differs even though the chords do not.

The case every measure has to fail on

The ostinato is in the corpus as the control: four bars, eight times, nothing else. A measure of structure that reports structure here is broken.

a four-bar ostinato, encoded as tonic, subdominant or dominant. A self-similarity matrix of 32 bars of a four-bar ostinato. Each cell is the cosine similarity of two bars encoded as tonic, subdominant or dominant, in place of the bag of pitch classes every other figure uses, with the bar's own key weighted 0.7 beside it. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 7 The four-bar ostinato under the function encoding. Mean off-diagonal similarity 1.000, no boundaries found, period one bar. All five encodings return exactly this, to three decimal places — which is the reassuring half of the result, and the reason it is worth showing the encoding that changes everything else changing nothing here.

Every encoding agreeing on the degenerate case and disagreeing on the real ones is the right shape for a robustness check. A measure that behaved differently on the ostinato would be reporting a property of its own arithmetic.

The measurement that cannot be compared at all

There is a fourth of the ladder’s measurements and it behaves in a fourth way, which is worth a section because the reason is instructive.

The compression rung put a number on how much of a piece is a repeat of an earlier part of itself by feeding the bars to a coder and counting the bits. To run it under a different encoding, the bars have to become symbols, and each encoding supplies its own alphabet: the rondo has ten distinct bar symbols as pitch-class-and-key, seven as bare roman numerals, nine as function-and-key, and six as root motions.

The ratios that come out are 1.135, 1.108, 1.041 and 0.958. Read naively that says the rondo is most repetitive as root motion — and it says nothing of the kind, because a coder with an alphabet of six needs fewer bits per symbol than one with an alphabet of ten before it has compressed anything. That rung said so itself, in as many words: the absolute number is a property of the coder and the alphabet, and only a ratio between two sequences encoded the same way is a statistic.

So the compression measure is not robust and is not fragile either. It is undefined across encodings, and its own rung had already said why. That is a third kind of answer to the question this rung asks, and finding it is the reason to ask the question of every measurement rather than of the interesting ones.

Which computation produced the numbers

The matrices are the same similarityMatrix the ladder has always used, over different feature vectors; the voice-leading matrix is built directly from the cheapest total motion between the two chords, divided by the largest such cost in the piece and subtracted from one.

The boundary scores are Foote’s novelty at a kernel of four bars, with peaks taken above 0.03, scored against the section letters in the scheme’s own definition with a tolerance of one bar. Those are the settings the rung that built the operator used, kept deliberately: the comparison here is between encodings and not between settings.

The recurrence periods are read at a key weight of 0.18 rather than 0.7, for the same reason — that is the weight the lag rung used, and it had already established the weight as a setting worth naming.

One defect turned up in the machinery while this was being computed and is worth recording. The site’s voiceLeading minimises total motion over every assignment of voices and requires the two chords to be the same size, which its docstring said and nothing enforced: hand it a triad and a seventh chord and it indexed past the end of the shorter one, and returned a cost of NaN. Every existing caller passes equal sizes, so no figure was wrong. It now handles unequal sizes by doubling — the smaller chord lends a note to a second voice, minimised over which note, which is what a realisation does — and a triad to a dominant seventh comes out at four semitones.

Whose music, and when

The six schemes are conventions rather than pieces, and that has not changed. A twelve-bar blues is a family with named variants; a thirty-two-bar song form is a plan; the rondo’s letters are a shape rather than a transcription. Everything here is a claim about what an analysis recovers from a plan, not about what it recovers from a recording.

What does travel is the methodological half, and it travels well beyond music. A local contrast detector inherits the dynamic range of its similarity measure. Give it a measure that saturates and it will find nothing and report that the material has no boundaries — which is exactly what happened here, in three schemes out of six, for eight rungs.

What the picture cannot show

None of the five encodings has any rhythm in it. Every bar is one chord and every bar is the same length. A passage that returns with the same harmony and a different rhythm is a repeat that none of these matrices can see, and a passage that returns with the same rhythm and different harmony is invisible to all five. That is the largest thing missing, it is missing from this rung as much as from the eight before it, and the corpus is what makes it missing: these schemes carry a roman numeral per bar and nothing else.

The function assignment is a table. Applied dominants are called dominants, the diminished triad on the seventh degree is called a dominant, and the mediant is called a tonic — the ordinary functional readings, which are contested at the edges and which the encoding does not derive from anything.

Five encodings is not the space of encodings. Registral position, melodic contour, texture and dynamics are all things a bar has and none of them is here, because the schemes do not carry them.

And the boundary scores are at one kernel width. The operator’s own rung swept three, and a different width changes the recall figures — not the ordering across encodings, which was checked at two and eight as well, but the numbers.

The ladder ends here

repetition closes at nine rungs. The model is a piece as a sequence of comparable units and a form as the pattern of resemblances between them, and the ladder now bounds it in every variable it has: what the matrix is (one), what operator reads an edge out of it (two), what number can be put on the whole (three), what the diagonals say (four), what a listener has available at the time (five), how transposition is handled (six), how memory discounts it (seven), where a repeat stops being one (eight), and now what a unit is encoded as.

Name a variable the model has, and it is on that list — with one exception that is named rather than hidden: the unit is a bar in all nine, and it is a bar because the corpus is written in bars. That is a property of the data rather than a free parameter of the model, and it is the reason the closure is honest.

What is not on the list belongs elsewhere. Where a phrase ends rather than a section is phrase; what makes an ending an ending is closure; and the reason the key plan shows up in these matrices at all is key-relations.

Boundaries found by a local operator, at three kernel widths. Foote's checkerboard novelty computed on the self-similarity matrix of rondo, ABACA, at kernel widths of 2, 4, 8 bars. The dashed verticals are where the encoding's sections actually change; nothing about them enters the computation. A peak is a place where the bars before resemble each other, the bars after resemble each other, and the two groups do not resemble each other.
Fig. 8 The operator this essay reassessed, at the three kernel widths its own essay used, on the scheme it was built for. It does reasonably here — three of the rondo’s four boundaries at the narrow kernel, two at the others — and it finds nothing whatever in three of the other five schemes. The reason for the difference is not in this picture. It is in the encoding underneath it, which is the same in all six.

Part 9 of 9

One essay in the series on repetition. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Description lengthNoveltyRepetitionSection boundarySegmentationSelf-similarityTonal functionVoice-leading