The boundary is where the neighbourhood changes
Assumes: A piece is mostly itself again
A self-similarity matrix shows a reader where the blocks are. It does not say where they end. That last step — turning a picture into a list of bar numbers at which one thing stops and another starts — is done by an eye, and an eye is exactly the instrument this site tries to keep out of its arguments.
It can be done arithmetically instead, by a construction of Jonathan Foote’s from 2000 that is small enough to describe in a sentence. Slide a square window along the diagonal of the matrix. Add up the similarity in the two quadrants that lie on the diagonal — the past against itself, and the future against itself — and subtract the two that lie off it, which compare past with future. Where the result is large, the recent past is internally consistent, the near future is internally consistent, and the two do not resemble each other. That is a boundary, and no part of the computation has any notion of what a section is.
The result is a curve with a value for every bar, and the peaks of that curve are candidate boundaries.
What the curve finds, and how much of it
The hero figure runs that operator over the rondo at three window widths, and draws the encoding’s real section boundaries as dashed verticals for comparison. The dashes are not an input. They are the answer sheet.
At every width the curve has two clear peaks, at the eighth and sixteenth bars. Those are the two places where the refrain gives way to the first episode and the first episode gives way back. The operator found them by asking a question about eight bars at a time, in a piece it knows nothing else about, and it put them within one bar of the truth.
There are two other section boundaries in that piece, at bars 24 and 32, and the curve does not find either of them.
That is the interesting half.
The AABA case is the hard one, because the boundary between the two A sections is a boundary between identical material. A checkerboard operator finds a change of neighbourhood and there is none there, which is the clearest statement of what this measure can and cannot see.
The first blindness: nothing changed
The second episode of this rondo is in the relative minor. That is not an exotic choice — it is one of the two commonest destinations in the whole of tonal practice, standing beside the dominant in almost every eighteenth- and nineteenth-century movement that leaves home at all.
And a major key and its relative minor contain the same seven notes. Not six of seven, as the dominant does; all seven. Key distance measured in shared tones puts them at zero.
So when the music moves from the refrain to the second episode, the pitch-class content does not change. The bars before the boundary and the bars after it are built from one collection. The two halves of the window resemble each other exactly as much as each resembles itself, the checkerboard sums to nothing, and the operator reports no boundary — correctly, given what it was shown.
What tells a listener that the music has arrived in the relative minor is not which notes are present but which one is being treated as home: where the cadences land, which note is held, what the bass does. That is a matter of emphasis and of the hierarchy a listener builds from what they have been hearing, and not one bit of it is in a bag of pitch classes.
The second blindness: it changed the same way it changed before
The other failure is sharper, and it is structural rather than a matter of encoding.
The eighth bar of an AABA plan is the seam between the first A and the second A, and the second A is a repeat of the first. So the window before that seam holds bars 5 to 8 and the window after holds bars 9 to 12 — which are, note for note, bars 1 to 4. The two halves are not merely similar. They are the same music.
The checkerboard is a difference detector. At the boundary between a passage and its own literal repeat, it is looking at the one place in the piece where it is guaranteed by construction to find nothing.
This is worth stating as a general result rather than as a quirk of one plan, because it is the exact complement of what the matrix does best. A literal repeat is the brightest object a similarity matrix contains: a stripe as strong as the diagonal, unmistakable, running the whole length of the repeated passage. It is also the dimmest object a novelty curve contains.
The two methods are blind to opposite things, which is why a segmentation built on either alone is unreliable and one built on both is worth having. The stripe says “this passage has happened before”; the peak says “something changed here”. Neither statement implies the other, and a form is made of both.
The pairing is not a compromise between two imperfect tools. It is a division of labour that follows from what the two operations are. A stripe is a comparison between two distant regions and is indifferent to whether anything happened in between; a checkerboard is a comparison between two adjacent regions and is indifferent to whether either has occurred before. There is no single window that does both, because they are questions about different distances.
A false positive is not a false alarm
The same figure has features where there is no boundary at all — two of them, at bars 18 and 22, and they are the largest things on the page.
Those bars are inside the bridge, which in this plan is a chain of applied dominants: a major triad with a minor seventh on each of four degrees in turn, each one the dominant of the next. Every two bars the pitch content shifts by a long way, because each of those chords brings in notes from outside the key.
So the operator is behaving exactly as specified. The neighbourhood really does change at bar 18, and it changes more there than at any labelled boundary in the piece. What the operator lacks is not sensitivity but a sense of scale: it cannot know that a two-bar change inside a section is a smaller event than an eight-bar change between sections, because it was never told how long a section is.
That is a real fact about the music and not a defect in the measurement. The bridge of a thirty-two-bar song is more internally varied than the seam between its sections, and any listener who has followed one knows this — the bridge is where the interest is. Reporting it as the most eventful part of the piece is not wrong. It is only unhelpful if the question was “where are the sections”.
The width is the question being asked
Which brings the argument to the parameter that has been sitting in every figure so far.
A narrow kernel asks a question about the last two bars and the next two. A wide one asks about the last eight and the next eight. Those are different questions and they have different right answers, and the blues has structure at both scales: four-bar groups inside twelve-bar choruses.
That figure is the strongest version of the caution. A listener never has the slightest difficulty telling a verse from a chorus; the distinction is carried by the singer’s register, by what the drums are doing, by the arrival of everything at once. None of it is harmony, and none of it is available to a method that has been handed a list of chords. A measurement can only be as discriminating as its description, and a chord scheme is a thin description of a song.
This is the same shape as an argument the site has already made about metre. A beat is inferred rather than received, and the same pattern of onsets supports several readings at once — the tresillo can be heard in four or in three, and both readings are available in the signal. What decides is which one a listener commits to, and a listener commits at a rate, which is the preferred tempo.
The kernel width is the same commitment, made explicit and made adjustable. A piece does not have a set of boundaries. It has a set of boundaries at each scale, and choosing the scale is not something the arithmetic can do.
The control, which returns nothing
This is the refusal the method needs. An operator that produced a plausible-looking list of section boundaries for a piece that has none would be producing plausible-looking lists for everything, and there would be no way to tell from the output which ones meant anything.
It also foreshadows a real problem, which gets an essay of its own. A great deal of music is built exactly like that figure: a cycle, repeated, with no chord changes at all. Every method in this field returns zero on it, and yet such music plainly has shape. What it uses instead of harmonic boundaries is a different quantity entirely, and one that is just as computable.
Kernel width is a claim about how far a neighbourhood extends, and the right width is the size of the sections being looked for. That is a circularity worth naming: the operator finds boundaries at the scale it is told to look at, so a scan at one width is a hypothesis rather than a measurement.
Where the model came from, and what it was for
Foote’s paper is about audio, not about chord symbols. The feature vectors in it are spectral — a description of the sound of a short frame, a fraction of a second long — and the matrix is thousands of frames on a side rather than forty. The method was proposed for finding structure in recordings, where no score exists and nobody has labelled anything, and it is still one of the standard tools for that job.
Running it on a symbolic encoding, as here, makes the curves far cleaner than any real one. A recording’s matrix is noisy: performance varies, instruments enter and leave, the same chord played twice does not produce the same spectrum. Every peak in the figures above would be a broader, lumpier thing on real audio, and the peak-picking that is trivial here is most of the difficulty there.
So these figures overstate how easy the job is, and they are worth reading as a demonstration of what the operator is rather than of how well it performs. The blindnesses, though, transfer intact. A spectral matrix is even less able to distinguish a key from its relative minor than a pitch-class one, and the seam between a passage and its literal repeat has no novelty in any feature space whatever.
Why a taper, which is not a detail
One implementation decision inside the operator is worth exposing, because leaving it out produces a result that looks like a finding and is an artefact.
The kernel described at the top is a plain square: plus one in two quadrants, minus one in the other two, sharp edges all round. A sharp-edged window is a box filter, and a box filter has sidelobes — it responds not only where the signal changes but at a fixed distance either side of the change, with the opposite sign, and then again further out.
Run the plain version over any of these schemes and the curve grows a secondary bump a few bars on each side of every genuine peak. Those bumps are large enough to survive peak-picking, and they land in the middle of sections. A reader would then be looking at a figure reporting section boundaries several bars into the middle of a section, with no way of telling from the picture that they were an echo of the real one.
So the kernel used here is tapered with a Gaussian, which is the standard remedy and costs one line. It is recorded in the generator with the reason attached, because the failure it prevents is the kind that gets written up as a discovery about music.
How large the failure would have been is worth measuring rather than asserting, and it depends entirely on the piece. Run both kernels over the blues at a width of four and the plain one produces peaks at bars 9, 21 and 33 that the tapered one does not report at all — four bars after each of its genuine peaks at 5, 17 and 29, which is exactly one half-kernel away, and at 41 per cent of the largest peak on the page. Anything at 41 per cent survives peak-picking, and bars 9, 21 and 33 are in the middle of the blues’ four-bar groups rather than at the edge of anything.
On the rondo the same comparison gives spurious peaks at one per cent, which is nothing. The difference is the material: three literally identical choruses give a box filter a great deal to ring on, and a rondo whose sections differ gives it almost none. So the taper is invisible on the pieces where it does not matter and decisive on the pieces where it does, which is the worst possible arrangement for noticing that it is needed.
The largest number in every one of these curves is an artefact
There is a second implementation fact that no figure mentions and that is larger than anything the taper does. A window centred near the end of a piece has half of itself hanging off the end, so its two forward quadrants are nearly empty and the checkerboard sums to whatever the past quadrant contains.
The result is a rise at the right-hand edge of every curve, and it is enormous:
| bars from the end | 4 | 3 | 2 | 1 |
|---|---|---|---|---|
| thirty-two-bar song | 0.044 | 0.098 | 0.219 | 0.464 |
| rondo | 0.040 | 0.098 | 0.216 | 0.463 |
Two schemes of different lengths, in different keys, with different section plans, produce the same four numbers to three decimal places — which is the proof that they are about the window rather than about the music. Against interior maxima of 0.042 and 0.065, the last bar of each piece scores seven to eleven times the largest genuine boundary in it.
Every reading in this essay is therefore a reading of the interior, and the figures are drawn over ranges that exclude the ramp. That is the right thing to do and it is a decision, not a property of the operator: a curve read end to end says the most eventful moment of a thirty-two-bar song is bar thirty-two, in every song ever written.
What this cannot show
The curve has no units. Its values are similarities averaged over a window, so a peak of 0.7 in one figure and a peak of 0.7 in another are not comparable unless the encodings and the widths match. Everything said above is about the shape of a curve and about which peaks are present, never about the size of one.
Nor does a peak say what kind of boundary it is. The operator that finds the entry of the rondo’s first episode and the operator that finds the arrival of the subdominant in a blues are the same operator returning the same kind of number, and nothing in the output distinguishes “a new section begins” from “the harmony moved”. Deciding that is the work, and it is done by a person.
And the whole apparatus is silent about the one thing a listener is most certain of, which is where a piece ends. An ending is not a change of neighbourhood. It is often the least surprising bar in the piece.
Where the ladder goes
Two directions open from here, and they are almost opposites.
One is to stop looking for edges and count instead: how much of a piece is a repeat of an earlier part of itself, measured in bits, which turns a picture into a number that two pieces can be compared on — and which has an awkward property that is worth meeting head-on.
The other is that ending. The reason no operator here can find one is that closure is not a local property of the music at the point where it happens; it is a conjunction of separate signals, each of which can be present without the others, and the site’s own solver has already refuted the usual account of how they combine.
Part 2 of 9
One essay in the series on repetition. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 12.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
HierarchyModulationNoveltyRelative minorRepetitionSection boundarySegmentationSelf-similarity
- A family of readings hierarchy, segmentation
- A modulation and a borrowing are one number apart modulation, segmentation
- A return has to be remembered repetition, self-similarity
- The ceiling is thirteen bars with names modulation, relative minor
- The circle is a circle, and the map is not modulation, relative minor
- The level the tempo chooses hierarchy, segmentation