Form and structure

A piece is mostly itself again

Take a piece of music, encode each bar as the notes sounding in it, and compare every bar with every other bar. The picture that comes out has blocks and stripes in it, and those blocks and stripes are the form — arrived at by arithmetic that has never heard of an exposition, a chorus or a refrain.

Assumes: A progression is a path, and the map can be drawn

Almost everything a listener hears has been heard before, a few seconds earlier, in the same piece. That is not a criticism and it is not a genre observation. It is the most robust statistical fact about music there is, and it is the one thing every account of musical form is an account of.

The problem is that the accounts are usually given in the wrong direction. A textbook says that this movement is a rondo, that this song is verse and chorus, and that this chorus has thirty-two bars in an AABA plan — and the evidence for each claim is that somebody with a score has said so. That is a perfectly good way to teach and a poor way to argue, because the labels arrive before the thing they are labels for.

thirty-two-bar AABA, as a strip of time. thirty-two-bar AABA laid out one cell per bar, coloured by section, with the roman numeral in each bar. the A section's turnaround is the ii-V every variant keeps; the bridge is a chain of applied dominants. At 108 beats a minute in 4/4 the whole of it lasts 71 seconds.
Fig. 1 The encoding this essay works from, and the whole of it. Thirty-two bars, one roman numeral each, coloured by section. Nothing here is a transcription of a recording — the thirty-two-bar song form is a convention with named variants, and this is the commonest of them. Everything below is computed from this row of symbols and from nothing else.

There is a way to run the argument in the other direction, and it is old, cheap and almost embarrassingly direct. Encode each bar as a description of what is sounding in it. Compare every bar with every other bar. Draw the answer.

Every bar against every other bar

The hero figure above is that comparison. Bar ii across, bar jj down, and the darkness of the cell is how similar the two bars are.

The similarity is a cosine. Each bar becomes a vector of twelve numbers, one for each pitch class, with the chord’s own notes set to 1 and the other notes of the key set to a smaller weight. Two bars sounding the same chord give parallel vectors and a similarity of 1. Two bars sharing nothing give a similarity near zero. The whole computation is a dot product and two square roots, and it is done 1,024 times.

What comes out is not a smooth wash. It has structure — visible, sharp-edged structure, in places a reader can point at. And nothing in the arithmetic knows what a section is, what a chorus is, or that the thirty-two bars are supposed to divide four ways. The section strips along the top and left edges of the figure are drawn beside the matrix rather than into it, precisely so the agreement can be checked rather than assumed.

The one decision inside the vector deserves its number stated, because it is doing more work than it looks like it is. A bar’s chord tones are set to 1 and the other notes of its key to 0.5. Drop that second weight to zero and two bars sharing no chord tone become exactly orthogonal: the tonic triad and the supertonic triad, which share not one note of the seven, would come out as far apart as any two objects can be, and so would the tonic triad and a chord from a distant key. The measure would then have one answer for “different chord” and the same answer for “different planet”, and every distinction between them would be gone.

With the key in the vector at half weight, two chords in one key stay nearer each other than either is to a chord from somewhere else, and the matrix acquires a middle range to read. That is a modelling choice and not a discovery. It is stated here, it is printed in the figure, and it is the parameter that gets varied later in this essay to make the point that it decides what the picture is about.

Two kinds of sameness, and they look different

The useful thing about the matrix, and the reason it is worth a figure rather than a number, is that it separates two things that ordinary language calls by one word.

A square block on the diagonal is a passage that stays like itself. Every bar in it resembles every other bar in it, so the whole square fills in. That is what a homogeneous stretch of music looks like — a pedal point, a static harmony, a groove.

A stripe parallel to the diagonal, offset from it is something else entirely: a passage that resembles a passage from earlier, bar for bar, in the same order. The offset is how long ago. A stripe eight cells up and to the right of the diagonal says that bars 9 to 16 repeat bars 1 to 8.

Those two are not variations of one phenomenon. The first says the music is not going anywhere; the second says it has been here before. Both are called repetition and only one of them is a return.

twelve-bar blues, every bar against every other bar. A self-similarity matrix of 36 bars of twelve-bar blues. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.18. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 2 Three choruses of a twelve-bar blues, thirty-six bars in all, with the key weight turned right down so that the picture is about chords rather than about keys. The stripes are unmissable and they are twelve cells apart, because that is how long a chorus is. Nobody supplied the number twelve. It is the offset at which the stripes appear.

The blues is the clean case because it is nothing but return. Twelve bars, then the same twelve bars, then the same twelve again — and the matrix says so in the only way it can, by putting bright diagonal stripes at an offset of twelve.

This is the first thing worth taking from the field. The period of a piece’s repetition is a measurable quantity, and it is measured as the offset of the strongest off-diagonal stripe. A listener who says “and now it comes round again” is reporting the same offset.

What the control looks like

An argument that a computation finds structure is worth nothing unless the computation can also fail to find it. So here is the case with no structure to find.

a four-bar ostinato, every bar against every other bar. A self-similarity matrix of 32 bars of a four-bar ostinato. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.5. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 3 Eight repetitions of a four-bar ostinato — the same chord throughout, thirty-two bars of it. Every bar resembles every other bar exactly, so the matrix is a single solid square with no block, no stripe and no edge anywhere inside it. The mean similarity off the diagonal is 1.00, which is the highest number the measure can return, and the picture carries no information at all.

This is the important negative result, and it is worth being precise about what it shows. The ostinato is not low on the measure. It is at the ceiling: every pair of bars is maximally similar. What is absent is not similarity but contrast, and structure turns out to be a claim about contrast rather than about repetition.

A measure that reported the ostinato as the most highly structured music in the collection would have been useless, and a reader is entitled to check that it does not. It reports it as a solid square, which is the correct answer and is also visibly nothing.

Where the blocks and the stripes disagree

The two kinds of sameness can be played against each other, and the case where they diverge is the one that decides how much the matrix is really doing.

verse and chorus, every bar against every other bar. A self-similarity matrix of 32 bars of verse and chorus. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.5. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 4 Two verses and two choruses, eight bars each. The verse is one four-chord loop stated twice; the chorus is a rotation of the same four chords. The stripes at an offset of sixteen bars are the second verse against the first and the second chorus against the first. The blocks are weaker than in the blues, and the reason is that this material never sits still for long enough to make one.

Two things in that figure are worth reading carefully.

First, the sixteen-bar stripe is strong, because verse two really is verse one. Second, the verse and the chorus are made of the same four chords in a different rotation, which means their similarity is high even though a listener has no trouble at all telling them apart. The matrix cannot separate them, because at the level of what notes are sounding in a bar they genuinely are near neighbours.

That is the honest limitation, and it arrives early enough to be a feature. A pitch-class encoding sees pitch classes. It does not see register, dynamics, orchestration, whether anybody is singing, or which of the two passages the drums come in on. In much popular music the verse and the chorus are distinguished by exactly those things and by almost nothing that this figure can measure.

verse and chorus, as a strip of time. verse and chorus laid out one cell per bar, coloured by section, with the roman numeral in each bar. eight bars each, the chorus keeping one progression while the verse moves under it. At 120 beats a minute in 4/4 the whole of it lasts 64 seconds.
Fig. 5 The encoding behind the matrix above. Verse and chorus are eight bars each and are built from the same four chords in a different rotation, which is why the matrix cannot separate them and a listener has no difficulty at all. Everything that distinguishes them in a recording — who is singing, what the drums are doing, how loud it is — is absent from this row of symbols.

Near-repeats, which are the interesting ones

Exact repetition produces a stripe of exactly the same brightness as the diagonal. What happens when a passage comes back changed is more informative.

two eight-bar phrases, every bar against every other bar. A self-similarity matrix of 16 bars of two eight-bar phrases. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.18. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 6 Sixteen bars in two eight-bar phrases that share their first six bars and differ in their last two — the standard antecedent-and-consequent shape. The stripe at an offset of eight is bright for six cells and breaks at the seventh, which is the point at which the second phrase stops agreeing with the first. The break is the cadence, and it is visible as a gap in a stripe before anything has been said about cadences.

The stripe that stops is the most useful single object in this figure. It says: this passage is a repeat of that one, up to a point, and the point is here. That is what a musician means by a varied repeat, and it is what a phrase that answers another phrase is built out of.

The site’s chord-distance map puts chords in a space by how far the voices must move between them; this puts bars in a space by how much they share. The two are related and are not the same: one is about the cost of a transition and the other about the identity of a state.

The same picture at a different scale

The weight given to the key’s notes as against the chord’s own is not a detail. It decides which structure the matrix reports.

A rondo plan is forty bars, ABACA, eight bars a section: the refrain at home, the first episode in the dominant and the second in the relative minor. Those key changes are part of the encoding rather than an addition to it, which is why the matrix has something to find in a rondo that it does not have in a verse and chorus.

rondo, ABACA, every bar against every other bar. A self-similarity matrix of 40 bars of rondo, ABACA. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.7. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 7 The same forty bars with the key weighted heavily. Now the picture is about where the music IS rather than which chord is sounding, and the three refrains resemble each other strongly while the two episodes resemble neither. The stripes at offsets of sixteen and thirty-two bars are the second and third refrains against the first.

Turn the key weight down and the matrix reports the bar-by-bar progression. Turn it up and it reports the sections. Both pictures are true of the same forty bars, which is not a defect in the method but the most useful thing it has to say: a piece does not have a structure, it has structures, and which one appears depends on what is being measured.

A listener does the same thing without noticing. Somebody following the harmony of a passage and somebody following its key are attending at two levels, and they will disagree about where the piece divides.

There is a further reason the rondo is the right example for this. Its two episodes are not equally far from home. The first is in the dominant, one step along the chain of fifths, sharing six of its seven notes with the refrain’s key; the second is in the relative minor, which shares all seven. Keys are neighbours in a computable sense, and a matrix built on pitch-class content is measuring exactly that neighbourliness — which means it can see the first episode leave and it cannot see the second one leave at all, because nothing about the pitch content has changed.

That is not a small caveat. The relative minor is one of the two commonest destinations in the whole of tonal music, and a method built on which notes are sounding is structurally blind to it. What tells a listener that the music has gone to the relative minor is which of those seven notes is being treated as home — a matter of emphasis, cadence and duration rather than of content — and none of that is in the vector.

It is also measurable, and the size of it is worse than “blind” suggests. Average every cell inside each section-pair block of the rondo at the key weight of 0.7 that the figure above uses:

refrain dominant episode relative-minor episode
refrain 0.972 0.846 0.946
dominant episode 0.846 0.972 0.823
relative-minor episode 0.946 0.823 0.938

The refrain resembles itself at 0.972 and resembles the dominant episode at 0.846 — a gap of 0.126, which is the contrast that makes the blocks visible. Against the relative-minor episode the gap is 0.026, a fifth as large.

And read the bottom-right corner. The relative-minor episode resembles itself at 0.938 and resembles the refrain at 0.946. It is more like the music it is supposed to be contrasting with than it is like itself — which is not a faint block or a missed edge but a section placed on the wrong side of the boundary. Anything that thresholded this matrix into segments would put the second episode inside the refrain.

The number does not improve by turning the parameter down, either. At a key weight of 0.18, where the matrix is nearly all chord, the refrain’s self-similarity is 0.671 and its similarity to the relative-minor episode is 0.591 — a gap of 0.080 against the dominant episode’s 0.067, so the ordering actually reverses and neither gap is large. There is no setting at which the relative minor separates the way the dominant does, because the thing that would separate it is not in the encoding at any weight.

The offset is the piece’s own clock

One more quantity falls out of the matrix without being asked for, and it connects this field to the one next door.

The offset at which the brightest off-diagonal stripe sits is the repetition period, and in the blues figure it is twelve bars. In the verse-and-chorus figure it is sixteen. In the rondo, sixteen and thirty-two. Those are the lengths at which the music comes round, and they are read off a picture rather than counted from a score.

A repetition period is a cycle, and a cycle drawn as a circle rather than a line is how this site has been treating rhythm since its first phase. The two are the same object at different time scales: a four-step drum pattern that comes round in two seconds and a thirty-two-bar chorus that comes round in seventy. The reason they get separate treatment is not that they are different kinds of thing but that only one of them fits inside the window a listener can hold at once — which is the subject of the phrase essay, and the boundary between rhythm and form as this site draws it.

thirty-two-bar AABA, every bar against every other bar. A self-similarity matrix of 32 bars of thirty-two-bar AABA. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.18. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 8 The hero figure’s thirty-two bars with the key weight turned down from 0.5 to 0.18, so that the vector is almost entirely the chord’s own three notes. The blocks dissolve and the bar-by-bar progression appears in their place — the ii–V turnarounds, the repeats at an offset of eight. The music has not changed; the question has.

The second time is not the same time

There is a fact about repetition that no matrix can hold, and it is worth putting next to one so that the gap is obvious.

Bars 9 to 16 of the thirty-two-bar plan are identical to bars 1 to 8. The matrix says so, correctly, by drawing a stripe of exactly the brightness of the diagonal: as far as the encoding is concerned the two passages are the same object. As far as a listener is concerned they are not, and the difference is not subtle. The first eight bars were new. The second eight are a repeat, and being a repeat is the most salient thing about them.

That difference lives entirely in the listener, which puts it beside every other quantity on this site that does. A metre is inferred rather than received; a tonal hierarchy is built by counting what has been heard; an interval is sorted into a category the sound does not contain. The status of a passage as a return belongs to the same list. The air is the same on both passes.

It has a practical consequence for everything below. A measure of structure computed from the notes is measuring the material a listener uses to construct a form, and not the form. That is worth having — the material is real, it is public, and two people can check it — but it is one side of the transaction. The other side is memory, and memory is why a stripe in a matrix and a return in a piece of music are not the same event.

What this cannot show, stated plainly

The matrix is a computation on an encoding, and it inherits every limitation the encoding has. The encodings used here are chord schemes: they carry the roman numeral sounding in each bar and the key that numeral is relative to, and nothing else.

So the figures above cannot see a melody, and a very great deal of musical repetition is melodic. They cannot see rhythm, and the return of a rhythmic figure over new harmony is one of the commonest ways music refers back to itself. They cannot see texture, instrumentation or dynamics. And they operate on a bar, which is far too coarse for anything that repeats faster than that.

Every one of those is fixable in principle by choosing a richer feature — the technique is indifferent to what goes in the vector, and the version of this figure used on real recordings puts a spectral description in it rather than a chord. What none of it fixes is the deeper point, which is that the picture is only ever as good as the description fed to it.

Where the ladder goes

The matrix is a picture, and a picture is not a segmentation. It shows a reader where the blocks are; it does not say where one ends and the next begins, and the eye that finds those edges is doing work the computation has not done.

That work can be automated too, by an operator that walks along the diagonal asking a purely local question — and the two boundaries it turns out to be blind to are precisely the ones this figure sees best, which makes the pair of methods more useful than either.

Beyond that, the amount of repetition in the matrix can be reduced to a number of bits, which allows two pieces to be compared rather than looked at. And the unit that all of this repetition is built out of — the phrase — turns out to have a length in seconds rather than in bars, which is a fact about the listener rather than about the music.

Part 1 of 9

One essay in the series on repetition. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 17.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

NotationProgressionRefrainRepetitionSegmentationSelf-similarityStrophicTransposition