Form and structure

How long until it comes back

A self-similarity matrix has a second reading that nobody looks for. Add up each diagonal instead of walking along one, and out falls repetition as a function of how long ago — a period, in bars, with no segmentation, no kernel width and no bar numbers anywhere in the answer. Five of the six schemes here report the length a listener would have named. The sixth reports something better.

Assumes: A piece is mostly itself again

Every measurement of musical form so far on this site has produced an answer with a bar number in it. The matrix shows a stripe that starts at bar 9. The boundary operator reports a peak at bar 16. Even the bit count is a count over a stated stretch of bars.

There is a quantity in the same data that has no bar number in it at all, and it is closer to what a listener means by form than any of them.

Along the diagonal, and across the diagonals

The self-similarity matrix is symmetric, so all its information sits in the triangle above the main diagonal, and there are two ways to read that triangle.

Foote’s checkerboard walks along the main diagonal, asking a local question at each step: do the bars behind resemble each other, do the bars ahead resemble each other, and do the two groups resemble each other. It produces a value per bar. It is an operator, and it has a window, and choosing the window is choosing which level of the hierarchy gets reported.

The other reading takes each diagonal whole. The diagonal offset by twelve holds every comparison between a bar and the bar twelve bars later — bar 1 against bar 13, bar 2 against bar 14, and so on to the end. Average it, and the result is a single number: how much this piece resembles itself twelve bars later. Do that for every offset and the result is a curve against lag.

twelve-bar blues, every bar against every other bar. A self-similarity matrix of 36 bars of twelve-bar blues. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.18. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.
Fig. 1 The matrix the hero figure is computed from — three choruses of a twelve-bar blues, with the key weight turned down so the picture is about chords. The stripes twelve cells off the diagonal are the second and third choruses against the first. The hero figure is this array of numbers averaged one diagonal at a time, which is why it has no bar numbers in it.

That difference is not a refinement. It is a change of subject. A novelty peak is a statement about a position — something happened here. A lag peak is a statement about a distance — this music comes round every so often, and the question of where it comes round does not arise. One is local and the other is global, and neither can be recovered from the other.

The arithmetic, and two decisions inside it

The computation is three lines. For each offset LL, take the mean of Mi,i+LM_{i,i+L} over every ii for which the pair exists, and plot it against LL.

Two decisions inside those three lines are doing real work, and both are visible on the page.

The number plotted is not the mean, it is the mean minus the piece’s own baseline. Cosine similarity on twelve-dimensional pitch-class vectors has a high floor: at the key weight used here every diagonal of every scheme has a mean between 0.62 and 0.97, so a plot of the raw means is a plot of six nearly identical large numbers. Subtracting the matrix’s own mean off-diagonal similarity gives an excess, and the excess has a property the raw mean does not: it is exactly zero, at every lag, for music that repeats no better than chance.

Lags run to half the length and no further. The diagonal at lag N1N-1 contains exactly one pair — the first bar against the last. Its “mean” is a single number, and it saturates at 1.00 for any piece whose first and last bars share a chord, which is nearly all of them. Uncapped, every scheme in this collection reports a period of N1N-1 bars, and that is a fact about averaging over one cell rather than about music. The cap is standard practice in autocorrelation and it is stated here because leaving it out produces a confident, wrong, publishable-looking result.

The hero figure is a twelve-bar blues under that treatment, and the answer is a spike at twelve. Nobody supplied the number twelve. There is no chorus in the computation, no section, no bar count.

Six schemes, and the one that refuses

An answer that is right on the easy case is worth nothing until it has been run on the hard ones.

How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 6 cases. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.483 and the smallest 0.000. 4 of 5 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.
Fig. 2 All six schemes on one vertical scale, so that the heights can be compared and not only the shapes. The dashed vertical on each row is that scheme’s own repeat unit — the highest common factor of its section lengths and its chorus length, derived rather than looked up — and it is drawn for checking, not supplied to the computation. Four of the five schemes with any structure at all put their strongest lag on that line or on a multiple of it.

The blues peaks at twelve, the rondo at sixteen, the verse-and-chorus plan at sixteen, the sixteen-bar period at eight. All four are the length a musician would have named, arrived at without naming anything.

The thirty-two-bar song form reports four.

That is not the answer the ladder was expecting. An AABA plan repeats its A section after eight bars — the second A is note for note the first — and eight is the number this measurement was supposed to return.

Why four is the right answer

The refusal survives inspection, which is what makes it worth keeping.

The thirty-two bars the four-bar answer is measured on are I, vi, ii, V in bars 1 to 4 and the same four again in bars 5 to 8, so the AABA plan’s shortest genuine period is four rather than eight — which is what the profile reports and is not what a reader expecting the section length would predict.

How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 1 case. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.302 and the smallest 0.302. 0 of 1 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.
Fig. 3 The thirty-two-bar plan alone, with the row scaled to its own contents. The strongest lag is four bars at an excess of 0.117; sixteen is at 0.107, twelve at 0.105 and eight at 0.087. Four candidates within a third of each other is what a piece with several clocks looks like, and it is the opposite of the blues, whose second-strongest lag is a fifteenth of its first.

The A section is not eight bars of unrepeated material. It is a four-bar turnaround stated twice, with the second statement’s last two chords thickened into sevenths. So the diagonal at lag four is drawing on every A section in the piece — six four-bar units, most of them nearly identical — while the diagonal at lag eight has to average the A-against-A comparisons together with eight bars of bridge against eight bars of A, and those are as unlike as anything in the piece.

Averaging is the whole method and averaging is what decides this. Lag eight has one perfect stretch and two poor ones; lag four has many good ones. The measurement is not wrong about the AABA plan. It is reporting a genuine period that is shorter than the one the letters name, and the letters were never an input.

And averaging is a fourth decision, which was not examined

Three decisions inside the arithmetic have been named and defended above — the baseline subtraction, the lag cap, the key weight. There is a fourth sitting in the word average, and it is the one that decides this section’s answer.

Open the two diagonals and count what is in them. Lag four holds 28 comparisons: 7 of them are effectively identical, 14 are above 0.9, and its median is 0.850, the highest of the four candidates. Lag eight holds 24: 9 are effectively identical — more than lag four has, out of a smaller diagonal — and 12 are at or below the piece’s own baseline, with a median of 0.606. That is the essay’s verbal account arriving as numbers, and it is more lopsided than the prose suggested: lag eight is not a weak diagonal, it is a bimodal one, more than a third literal repeat and half of it nothing at all.

So the ranking depends on which summary is taken of the same twenty-four numbers.

summary of the diagonal winner runner-up
mean — the method used here 4 16
median 4 12
fraction effectively identical 8 16

By how much of a diagonal is a literal repeat, the AABA plan’s period is eight — 37.5 per cent of lag eight’s cells against 25 per cent of lag four’s — which is the number the letters name and the number this rung expected. By the mean, and by the median, it is four.

Neither summary is wrong and the mean is the right default, because a period is a claim about the whole distance and not about its best moments. But the earlier sentence that the measurement is “reporting a genuine period that is shorter than the one the letters name” needs its second half: it is doing so because the mean is what was taken, and the statistic that asks how much of the return is exact recovers the letters exactly. The refusal is real and it is a refusal by the mean.

A piece does not have a period. It has periods, and this measurement returns all of them at once, ranked. Under the AABA plan the ranking runs 4, 16, 12, 8 — the turnaround, the distance from the first A to the bridge’s midpoint, and so on down. Under the blues, twelve towers over everything and the next candidate is a fifteenth of its size. That difference between a scheme with one clock and a scheme with several is itself a measurement, and it is the spread printed beside each row.

The scheme with the sharpest clock is not the one with the most repetition

The spreads printed down the right-hand side of the six-scheme figure rank the schemes on something the earlier rungs of this ladder had no way of measuring, and the ranking is not the obvious one.

The verse-and-chorus plan has the largest spread of any scheme here — 0.483, above the blues’ 0.427 — and it is not the most repetitive of them by any count made so far. The redundancy measurement puts it at 0.47 new phrases a bar against the blues’ 0.42, which is to say slightly less repetitive. What the spread reports is something else: how concentrated the repetition is at one distance.

Drawn as a matrix the verse-and-chorus plan is sixteen bars stated twice and nothing else, so a single diagonal at an offset of sixteen carries almost every resemblance in the picture and the rest is nearly empty. That concentration is what a spread of 0.483 measures, and it is not the same quantity as how much of the piece is a repeat.

Sharpness and quantity are different properties and this figure measures the first. A scheme whose material comes back at one distance, always, scores high; a scheme that repeats just as much but at four different distances scores low, because the excess is divided among four diagonals. The AABA plan is the clearest case of the second kind, which is why its tallest bar is short even though its A sections are literal repeats.

That is worth stating because the reflex on seeing a curve like this is to read the height as how repetitive, and it is not. It is how periodic.

What the control has to look like

How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 1 case. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.000 and the smallest 0.000. 0 of 0 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.
Fig. 4 Eight repetitions of a four-bar ostinato. Every diagonal has the same mean as every other, so the excess is zero at every lag and the spread is 0.000 — not a low answer but no answer. The strongest lag is reported as 1 only because something has to be reported when every candidate ties, and the figure says the profile is flat rather than naming a winner.

This is the case the whole ladder is checked against, and it behaves the way it did for the matrix and for the boundary operator: thirty-two bars of one chord produce a completely flat profile.

Worth being exact about what flat means here. It does not mean the ostinato is unrepetitive — it is the most repetitive object in the collection, and the bit count says so by putting it at 0.25 new phrases a bar against everything else’s 0.4 to 0.6. It means the ostinato has no preferred distance. It resembles itself equally at every lag, so there is no length at which it comes round, and a measurement that named one would be inventing it.

The spread, printed on every row, is what separates the two cases: 0.427 for the blues and 0.000 for the ostinato, on the same axis.

The key weight decides which question is being answered

There is a parameter in the encoding that the first essay of this ladder made an argument out of: how much weight the key’s own notes get beside the chord’s. Near zero the matrix is about chords; near one it is about keys, and the sections appear as blocks.

That essay’s recommendation, for looking at sections, was to turn the weight up. This measurement wants it turned down, and the reason is worth following, because it is the clearest case on the site of two figures built from one array of numbers giving opposite advice.

How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 3 cases. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.247 and the smallest 0.056. 2 of 3 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.
Fig. 5 One rondo, three key weights, one vertical scale. At 0.18 the strongest lag is sixteen bars, which is the distance from one refrain to the next. At 0.7 the strongest lag is two, and the spread has fallen from 0.247 to 0.056. Nothing about the music changed between the rows.
How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 3 cases. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.427 and the smallest 0.130. 3 of 3 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.
Fig. 6 The blues given the identical treatment, and it does not behave the way the rondo does. The strongest lag stays at twelve at all three key weights; what changes is the spread, which falls from 0.427 to 0.210 to 0.130. The statistic is degraded and it is not broken, which is the difference between a scheme whose period is carried by its chords and one whose period is carried by where the music is.

Two schemes, two behaviours, and the second is what stops the first being a rule. The blues’ twelve-bar return is a return of chords — the same I7, IV7 and V7 in the same order — so weighting the key down or up scales the whole profile without moving its peak. The rondo’s sixteen-bar return is a return to a key, and the moment the key dominates the vector, every bar in the tonic resembles every other bar in the tonic and the return has nothing left to stand out against.

With the key weighted heavily, every pair of bars in the same key is nearly identical, because most of each vector is the key rather than the chord. Adjacent bars are in the same key almost by definition, so the short lags fill in, the profile tilts, and the winner is lag one or lag two — the trivially true observation that this bar resembles the one before it.

The measurement has not broken. It is answering the question it was given: at a key weight of 0.7, how much does this piece resemble itself at each distance, where resemble mostly means is in the same key. The answer is that it resembles itself most at very short distances, which is correct and useless.

A parameter that improves one reading of a matrix can destroy another reading of the same matrix, and there is no setting that is best for both. That is a specific instance of something the redundancy essay reached from the other direction: the number is a measurement of the description, and the description was chosen.

Where the local operator and the global statistic disagree

Putting the two readings beside each other is the argument this rung exists to make.

How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 1 case. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.247 and the smallest 0.247. 1 of 1 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.
Fig. 7 The rondo’s own profile, which is the case the two period-finders disagree about. It reports sixteen bars: the refrain against the refrain, with an episode in between. The site’s metre-induction function run on the same forty bars — a bar where it usually has a beat, a section start where it usually has an onset — scores eight bars at 3.00 and sixteen at 2.33, so it prefers the section length where this prefers the refrain’s return. Neither is wrong. A hypermetre asks at what spacing the marked bars line up and the marked bars are five section starts, eight apart; a lag profile asks what correlates with what, and the thing that correlates is sixteen apart.

The boundary operator finds two of the rondo’s four section boundaries at a narrow kernel and all four at a wide one, and its false positives land in the middle of sections. The lag profile has none of those difficulties and cannot answer the question at all: it says the rondo repeats every sixteen bars and it does not know where a section starts. Neither statement contains the other.

That disagreement is worth more than an agreement would have been. A hypermetre is a grid — it asks at what spacing the marked bars line up — and the marked bars here are the five section starts, which are eight apart. A lag profile is a correlation, and the thing that correlates is the refrain against the refrain, which is sixteen apart because there is an episode in between. Both numbers are properties of the rondo and they are answers to different questions, and a reader handed only one of them would take it for the piece’s period.

There is a family resemblance here to something the rhythm field settled two phases ago. A rhythm drawn as a circle is a claim that the thing has a period and that positions inside it are what matter; a metre inferred from onsets is a claim about which period wins among several available at once. The lag profile is the second of those operations, run with a bar where the metre finder has a beat, and it returns the same kind of ranked list. The bar above the bar already made that argument at the four-bar level; this makes it at the forty-bar level, and nothing in the arithmetic notices the change of scale.

What this cannot show

Every limitation of the matrix is inherited whole, because this is that matrix summed. The encodings are chord schemes: a roman numeral per bar and the key it is relative to. No melody, no rhythm, no texture, no instrumentation. A bar is the finest thing there is.

Three limitations belong to this measurement in particular.

The profile is a mean, so it cannot tell a piece that repeats twice exactly from a piece that repeats four times approximately. Both can produce the same excess at the same lag. A stripe in the matrix carries that difference and this curve throws it away.

It has no notion of where in the piece the repetition happened. A piece whose first half is a strict canon and whose second half is free improvisation would produce a lag peak from the first half alone, reported as a property of the whole. Nothing in the figure would show that half the piece contributed nothing to it.

And it rewards evenness. A scheme whose returns are all at one distance beats a scheme whose returns are at several, even if the second repeats more of itself in total. The rondo’s 0.247 against the verse-and-chorus plan’s 0.483 is partly that: both return, and only one returns on a single clock.

Where the ladder goes

Something has been quietly assumed in every figure above, and it is not true.

Each of these profiles was computed on the whole matrix — every bar against every other, including bars that had not happened yet at the moment the earlier ones were heard. The blues reports a period of twelve bars, and it reports it from a position outside time, with all thirty-six bars in hand at once.

The period, as the piece goes by. The strongest lag of twelve-bar blues computed on only the bars heard so far, against how many bars that is. The final answer is 12 bars; it is revised 9 times on the way, and is not reached for the last time until bar 24 of 36, which is 67 per cent of the way through and 53 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have.
Fig. 8 The same strongest-lag computation run on only the bars heard so far, at every point in the same blues. The final answer of twelve bars is not reached for the last time until bar 24 — two full choruses, fifty-three seconds at 108 beats a minute — and it is revised nine times before that. Everything the hero figure reports is true and none of it is available yet.

A listener does not have the whole matrix. A listener has its leading corner, growing one bar at a time, and the question of what can be computed from a corner is the next rung — where the answer for the boundary operator turns out to be nearly everything, three bars late, and the answer for this measurement turns out to be nothing until the piece is nearly over.

After that, two questions this rung has left open. This measurement counts a return only if it comes back in the same key, and that is a setting rather than a fact. And it counts a lag of four bars and a lag of twenty-four as the same kind of event, which is the one thing every listener knows is false.

Part 4 of 9

One essay in the series on repetition. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

OstinatoPeriodicityRefrainRepetitionSegmentationSelf-similarityStrophic