Form and structure

The same thing somewhere else

A measure built on which notes are sounding calls a passage that comes back a fifth higher a stranger. There is a dial that fixes this, and turning it is supposed to be a trade — more sensitivity to a transposed return, less specificity against a coincidental one. It is not that trade. Two different statistics answer opposite ways, and the setting that would compromise between them is the worst one available.

Assumes: A piece is mostly itself again

The single most common thing music does with an idea is state it again somewhere else. A fugue subject enters on the dominant. A sequence walks a two-bar figure down by step. A last verse goes up a semitone. The same seven notes started later is one of the site’s oldest arguments and it is about exactly this: an object that keeps its shape and changes its position.

Every measurement in this ladder is blind to it. The vector for a bar sets the chord’s own pitch classes to 1 and the rest of the key to a small weight, so a tonic triad in C and a tonic triad in G share one note out of three and their cosine is small. Move a passage a fifth up and the matrix reports two strangers.

That is fixable, and the interesting part is what the fix costs.

The dial

The fix is to read each chord relative to its own key rather than absolutely, so that the tonic triad of C and the tonic triad of G become the same vector. Done as a switch it destroys something valuable, because it also makes a modulation invisible — and a modulation is most of what makes a rondo episode an episode.

So it is built here as a dial. Each bar’s vector is a mix,

v=(1κ)vabsolute+κvrelative,v = (1-\kappa)\,v_{\text{absolute}} + \kappa\,v_{\text{relative}},

where κ\kappa runs from 0 to 1. At 0 nothing changes and barVector is exactly what it was. At 1 the encoding has forgotten which key anything is in and remembers only which degree of it. In between, a transposed return is partly a return.

The mix is of the vectors and not of two finished similarity matrices. That is a deliberate choice and it is the more honest one: a listener does not run two representations of a passage and average their verdicts, so the question worth asking is what a single representation is sensitive to.

The corpus does not contain the case

Before the dial can be swept, an awkward fact has to be stated, because it decides how much the sweep is worth.

None of the six schemes encoded on this site repeats a section in a new key. The rondo’s three refrains are all at home. The blues repeats its chorus unchanged. The verse and chorus of the popular plan never leave the tonic at all. The thing the dial exists for does not happen anywhere in the corpus.

Sensitivity and specificity on one dial. Aligned similarity — bar i against bar i+L, which is what a return is — for 2 eight-bar comparisons, as the key-invariance dial turns. One comparison is constructed: a literal repeat in the encoding, moved up a fifth, which is a stated manipulation because no scheme encoded here repeats a section in a new key.
Fig. 1 The constructed case and its control, which is the whole of what the dial is for. None of the six schemes encoded here repeats a section in a new key — the rondo’s three refrains are all at home, the blues repeats its chorus unchanged, the verse-and-chorus plan never leaves the tonic — so the sensitive case has to be built, and the smallest available construction is to take the thirty-two-bar plan, whose second A is a note-for-note repeat of the first, and move it up a fifth. The upper line is that; the lower is the same pair untouched. Real music does this constantly and these encodings do not, which is cheaper to say than to pretend otherwise.

So the sensitive case has to be constructed, and constructing it is a stated manipulation of an encoding rather than a claim about a repertoire. The manipulation is the smallest one available: take the thirty-two-bar plan, whose second A section is a note-for-note repeat of the first, and move that second A up a fifth. The result is a literal repeat in a new key — the object the dial is for — and everything else about the encoding is untouched.

Real music does this constantly. The encodings here do not, and saying so is cheaper than pretending otherwise.

On the aligned test, the dial works

A return is an aligned comparison: bar ii against bar i+Li+L, bar i+1i+1 against bar i+L+1i+L+1, and so on down the stripe. That is what the hero figure measures.

The four comparisons in it are the constructed transposed repeat; the rondo’s refrain returning after sixteen bars, which is literal; the rondo’s refrain against its first episode, which is a passage in the dominant that is not a return of anything; and a verse against its own chorus.

The result is clean and it is monotone. The transposed repeat rises from 0.540 at κ=0\kappa=0 to 1.000 at κ=1\kappa=1 — at full invariance it is indistinguishable from a literal repeat, which is correct, because that is what it is. The coincidental case rises too, from 0.519 to 0.894, because full invariance also makes the episode’s I–V–I look like the refrain’s I–V–I. And the margin between them grows the whole way, from 0.021 to 0.106.

By this test there is no trade-off at all. The best setting is 1, the dial should simply be turned to full invariance, and the essay was slated to report a compromise that does not exist.

On the block test, it does the opposite

The trouble is that an aligned comparison is not the statistic a segmentation actually uses.

A method that finds sections from a matrix looks for blocks: regions in which every bar resembles every other bar, without alignment, because a section is a set of bars that belong together and not a sequence that matches another sequence position for position. The number that decides whether a scheme has blocks is the mean similarity of bar pairs inside one section minus the mean of bar pairs in different sections.

What full key-invariance costs a segmentation. For 5 schemes at a key weight of 0.18, the mean similarity of bar pairs inside one section minus the mean of bar pairs in different sections, as the key-invariance dial turns from absolute pitch to full transposition-invariance. Only rondo, ABACA moves at all, because it is the only scheme here with a key change in it: it starts at 0.083, dips to -0.012 at a setting of 0.50 — where its sections are on average no more like themselves than like each other — and ends at 0.070, which is still below where it began. It crosses zero twice, first at about 0.35. 2 of the 5 sit below zero at every setting, because their sections are built from one another's chords.
Fig. 2 The same dial, the same key weight, and the statistic a segmenter thresholds. Only the rondo moves, because it is the only scheme with a key change in it — and it does not simply improve or simply worsen. It starts at 0.083, crosses zero at about 0.35, bottoms at −0.012 at halfway, and recovers to 0.070 at full invariance, which is still below where it began.

At κ=0.5\kappa=0.5 the rondo’s sections are, on average, no more like themselves than they are like each other. A separation of −0.012 is not a weakened segmentation; it is the absence of one. Handed that matrix, a method looking for blocks would find the rondo’s five sections no more readily than five arbitrary eight-bar windows.

Drawn as matrices at the two ends of the dial, the rondo’s block structure is sharp at zero and blurred at one half — the transposed material has been partly forgiven and the modulation partly forgotten, and neither is fully either.

The worst place to stand is the middle, which is exactly where a compromise would put the setting. At 0 the encoding distinguishes an episode from a refrain by their keys. At 1 it distinguishes them by their chord degrees, which are genuinely different — an episode is I–V7–I–V7 where a refrain is I–V–I–IV. At 0.5 it has half of each cue and enough of neither, and the two blur.

Why the two tests disagree

The disagreement is not a paradox and naming its mechanism makes both results usable.

An aligned test asks: is this passage the same thing as that one, bar for bar? Increasing invariance can only help, because the two passages being compared have already been put in correspondence and the only question left is whether corresponding bars match.

A block test asks: do these bars belong together and not with those? Increasing invariance removes a distinction — which key a bar is in — and every distinction removed makes some pairs of bars more alike. Some of those pairs are inside a section and some are across a boundary, and there is no reason for the two effects to be equal.

So the dial’s setting is not a property of the encoding. It is a property of the question, and this ladder has now met the same shape three times: the key weight decides whether the matrix is about chords or about sections, the kernel width decides which level of the hierarchy gets reported, and the coder’s alphabet decides which of five schemes counts as the most repetitive. Every parameter in the field turns out to be a question in disguise.

The case the dial cannot touch

There is one comparison in the hero figure that does not move at all, and it is the most common musical situation of the four.

Sensitivity and specificity on one dial. Aligned similarity — bar i against bar i+L, which is what a return is — for 3 eight-bar comparisons, as the key-invariance dial turns.
Fig. 3 Three aligned comparisons in material that never leaves its home key, across the whole range of the dial. Every line is flat. Whatever separates a verse from a chorus, and whatever a blues chorus’s return consists of, the dial has no purchase on any of it — because the dial only ever discards information about keys, and there is none to discard.
What full key-invariance costs a segmentation. For 2 schemes at a key weight of 0.18, the mean similarity of bar pairs inside one section minus the mean of bar pairs in different sections, as the key-invariance dial turns from absolute pitch to full transposition-invariance. Only twelve-bar blues moves at all, because it is the only scheme here with a key change in it: it starts at 0.141, dips to 0.141 at a setting of 0.00 and ends at 0.141, which is still below where it began. It stays above zero throughout. 1 of the 2 sit below zero at every setting, because their sections are built from one another's chords.
Fig. 4 Two schemes with no modulation in them at all, swept the same way. Their block scores are flat across the dial — nothing changes, because there is nothing for key-invariance to forgive — which is the control the block test needs and the reason the rondo is the only scheme that moves. A parameter that changes nothing on four of five schemes and hurts the fifth is not a dial with a good setting somewhere on it; it is a dial whose only effect in this corpus is a cost.

The verse and chorus of the encoded popular plan sit at 0.658 aligned similarity at every setting from 0 to 1, and their block separation is −0.012 at every setting too, which puts them below zero for the same reason the rondo touches zero at halfway: the verse is I–V–vi–IV and the chorus is the same four chords rotated, so bars in different sections are frequently the same chord.

No setting of this dial separates a verse from its chorus. That was the slated claim and it is confirmed — but not for the slated reason. It is not that separating them costs something elsewhere. It is that transposition is not what distinguishes them, so a parameter about transposition cannot help, and could not have helped at any price.

And the verse-and-chorus plan written out shows why: its two sections differ in their chords and not in their key, so the thing distinguishing them is exactly the thing the dial preserves.

The dial repairs one blindness and installs another

The rondo has two episodes and they are not equally far from home. The first is in the dominant, one step along the chain of fifths, sharing six of seven notes with the refrain’s key. The second is in the relative minor, which shares all seven — which is why the boundary operator cannot see it arrive and why the matrix cannot see it leave. Sweeping the dial over both at once produces the most useful figure in this essay.

Sensitivity and specificity on one dial. Aligned similarity — bar i against bar i+L, which is what a return is — for 3 eight-bar comparisons, as the key-invariance dial turns.
Fig. 5 The refrain against its own return, against the first episode and against the second, aligned bar for bar, across the dial. The return is 1.000 throughout and is the ceiling. The other two cross at about 0.4 — the dominant episode climbing from 0.519 to 0.894 and the relative-minor episode falling from 0.713 to 0.586 — so the two failures recorded earlier are not both cured at either end.

At the dial’s zero, the relative-minor episode is more like the refrain than the dominant episode is, at 0.713 against 0.519. That is precisely wrong as a description of the music: both are departures, and by any account of tonal practice the relative minor is at least as much of one. The reason is arithmetic and was established two rungs ago — the two keys share all seven notes, so a pitch-class encoding has nothing to report.

At the dial’s one, that is repaired. The relative-minor episode falls to 0.586, well clear of the refrain’s return, because its chords are genuinely different degrees: i, iv, i, V7 against I, V, I, IV. Read relative to their own keys, the two passages are not alike, and the measurement finally says so.

And in the same movement the dominant episode climbs to 0.894, which is a new failure of the same size in the opposite direction. Its degrees really are nearly the refrain’s — I, V7, I, V7 against I, V, I, IV — and what made it an episode was the key it was in, which is the one thing full invariance has thrown away.

One dial, two blindnesses, and it exchanges them rather than removing either. A setting near 0.4 has both at once, at moderate strength, which is the worst of the three outcomes and is where a reader who wanted a compromise would have put it.

What a key change actually is, and why halving it is odd

There is a reason the midpoint behaves so badly, and it is a fact about keys rather than about the parameter.

A modulation to a near key turns on one note the home key does not use, and keys a step apart share six of their seven notes — so halving the key-invariance does not half-forget a modulation; it half-forgets one note. That is why the midpoint behaves worse than either end: at zero the encoding sees a modulation as a wholesale change of vectors, at one it sees none at all, and in between it sees a small change that looks like noise.

A key is not a continuous quantity. A passage is in the dominant or it is not, and key distance measured in shared tones puts the nearest neighbours six notes of seven apart and the relative minor seven of seven. Setting the dial to 0.5 does not produce a representation of a passage that is halfway to the dominant, because there is no such passage. It produces a vector that is a blend of two well-formed descriptions and is neither, and blends of well-formed descriptions are not generally well-formed.

That is the sense in which this dial is not really a dial. It has two settings that mean something and a continuum between them that means nothing in particular, and the sweep is worth drawing precisely because the shape of the curve says so — a U with a floor at the halfway point is what a parameter looks like when its interior is not a compromise but a confusion.

Whose music makes this matter

The dial is worth building because of a repertoire it is not being run on.

Transposed restatement is the organising device of eighteenth-century contrapuntal practice: a fugue subject answered on the dominant, an episode built from a sequence that walks the same two bars down through a chain of fifths, an invention whose whole second half is its first half moved. In that repertoire a measure blind to transposition is blind to most of the form. It is nearly as central to the classical style that followed — a sonata exposition states its material and then states it again a fifth up, which is the key plan doing the work of the form — and it survives into popular practice as the modulating last chorus.

None of that is in this corpus, and the reason is worth admitting rather than hiding. The six schemes here are plans — letter schemes and chord successions of a kind that can be written down as a convention — and a plan is exactly the level of description at which transposed restatement disappears, because a plan says “A” twice and does not say where. The thing this dial is for lives one level below the encoding this site uses, in the notes, and a chord scheme is a thin description of a piece.

What this cannot show

The transposed return in the hero figure was made rather than found, and one constructed case is not evidence about a repertoire. Everything the figure says about sensitivity is a statement about that one manipulation of one encoding, and the honest reading is that it shows the dial doing what it claims to do, not that it shows how often the situation arises.

The dial also handles only transposition. Music restates ideas at the same pitch with different harmony, in the parallel minor, in augmentation, in inversion, and with the melody kept and the chords replaced — and none of those is a rotation of a pitch-class vector. A measure invariant to transposition is invariant to one symmetry out of many, and it is the easiest one.

The U is not the dial’s shape, it is the dial’s shape at 0.18

Every figure above is drawn at a key weight of 0.18, and the obvious guess about the other end of that parameter is that full invariance would be worse still, because at a high key weight the vectors are mostly key and the dial is the thing that throws the key away. The guess is wrong, and drawing it is the cheapest way to find out.

What full key-invariance costs a segmentation. For 5 schemes at a key weight of 0.7, the mean similarity of bar pairs inside one section minus the mean of bar pairs in different sections, as the key-invariance dial turns from absolute pitch to full transposition-invariance. Only rondo, ABACA moves at all, because it is the only scheme here with a key change in it: it starts at 0.084, dips to 0.057 at a setting of 0.30 and ends at 0.197, which is above where it began, so at this key weight full invariance is the best setting on offer. It stays above zero throughout. 2 of the 5 sit below zero at every setting, because their sections are built from one another's chords.
Fig. 6 The same five schemes and the same dial at a key weight of 0.7 instead of 0.18. The rondo no longer crosses zero at all: it starts at 0.084, dips only to 0.057 around a third of the way along, and rises to 0.197 at full invariance — more than twice its starting value, and by a distance the 0.18 figure never approaches. At this key weight the dial’s best setting is 1.

So the U survives, in the sense that the middle is still the worst place to stand, and the conclusion drawn from it does not. At 0.18 full invariance is worse than absolute pitch; at 0.7 it is better by a factor of two and a third, and at 0.5 it is better by nearly two.

The mechanism is not mysterious once the two parameters are held side by side. At a high key weight most of each vector is the seven notes of the key, and at full invariance every bar is read in its own key — so that large shared component becomes identical for every bar in the piece and stops distinguishing anything. What is left to distinguish bars is the chord degrees, and the rondo’s sections differ in their degrees rather more than they differ in their keys: I–V–I–IV against I–V7–I–V7 against i–iv–i–V7, the last of them in the minor. The high key weight was hiding that, and full invariance uncovers it.

Two parameters, and neither one’s best setting is a property of the dial alone. The recommendation this essay would have issued from its first two figures — leave the dial at zero — holds at a chords-first key weight and is exactly backwards at a keys-first one. That is not a reason to distrust either figure. It is the reason a second one had to be drawn, and it is the fourth time in eight rungs that a parameter presented as a setting has turned out to be a question.

Where the ladder goes

The dial’s honest recommendation, on this corpus and at the chords-first key weight every other rung of this ladder uses, is to leave it at zero and say so. That is what the rest of the ladder does, and it now has a reason rather than a default — with the reason’s own jurisdiction attached, which the last section is about.

That leaves the ladder’s oldest unexamined assumption still standing. Every measurement so far treats a return at four bars’ distance and a return at twenty-four bars’ distance as the same event, differing only in a number. They are not the same event, and the difference is not in the music at all.

Beyond that lies the question of what a repeat does when it is not exact. This rung has been asking whether a passage counts as a return; the last rung asks where a return stops agreeing with what it returns to, and the answer turns out to depend on a threshold in a way that is more interesting than the answer.

Part 6 of 9

One essay in the series on repetition. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Key-invarianceModulationRefrainRelative minorRepetitionSegmentationSelf-similarityTransposition