Concept

Segmentation — where it appears

The division of a piece into successive parts, whether asserted by an analyst or derived from a measurement of where its content changes. At the level of chords it depends on the metre, so the parts are not readable off the pitches alone.

Named by 25 essays across 5 fields — each of them below, with the objects they name alongside it.

thirty-two-bar AABA, every bar against every other bar. A self-similarity matrix of 32 bars of thirty-two-bar AABA. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.5. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.

A piece is mostly itself again

Take a piece of music, encode each bar as the notes sounding in it, and compare every bar with every other bar. The picture that comes out has blocks and stripes in it, and those blocks and stripes are the form — arrived at by arithmetic that has never heard of an exposition, a chorus or a refrain.

form · Repetition
Boundaries found by a local operator, at three kernel widths. Foote's checkerboard novelty computed on the self-similarity matrix of thirty-two-bar AABA, at kernel widths of 2, 4, 8 bars. The dashed verticals are where the encoding's sections actually change; nothing about them enters the computation. A peak is a place where the bars before resemble each other, the bars after resemble each other, and the two groups do not resemble each other.

The boundary is where the neighbourhood changes

A section boundary can be found by an operator that never sees a section. It walks the diagonal of a similarity matrix asking one local question — do the bars behind me resemble each other, do the bars ahead resemble each other, and do the two groups resemble each other — and where the answer is yes, yes, no, there is an edge. What it cannot find turns out to say more than what it can.

form · Repetition
How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 1 case. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.427 and the smallest 0.427. 1 of 1 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.

How long until it comes back

A self-similarity matrix has a second reading that nobody looks for. Add up each diagonal instead of walking along one, and out falls repetition as a function of how long ago — a period, in bars, with no segmentation, no kernel width and no bar numbers anywhere in the answer. Five of the six schemes here report the length a listener would have named. The sixth reports something better.

form · Repetition
The period, as the piece goes by. The strongest lag of thirty-two-bar AABA computed on only the bars heard so far, against how many bars that is. The final answer is 4 bars; it is revised 6 times on the way, and is not reached for the last time until bar 29 of 32, which is 91 per cent of the way through and 64 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have.

The form a first hearing cannot have

Every figure so far was computed with the whole piece in hand. Run the same methods over only the bars already heard and one of the two methods survives intact — the boundary operator turns out to be causal at a fixed delay of a few bars — while the other collapses. The period of a piece is not knowable until the piece is nearly over, and in two of the six schemes here not until its last bar.

form · Repetition
Sensitivity and specificity on one dial. Aligned similarity — bar i against bar i+L, which is what a return is — for 4 eight-bar comparisons, as the key-invariance dial turns. One comparison is constructed: a literal repeat in the encoding, moved up a fifth, which is a stated manipulation because no scheme encoded here repeats a section in a new key. The shaded band is the margin between the two named comparisons, and it runs from 0.021 to 0.106.

The same thing somewhere else

A measure built on which notes are sounding calls a passage that comes back a fifth higher a stranger. There is a dial that fixes this, and turning it is supposed to be a trade — more sensitivity to a transposed return, less specificity against a coincidental one. It is not that trade. Two different statistics answer opposite ways, and the setting that would compromise between them is the worst one available.

form · Repetition
thirty-two-bar AABA, as a strip of time. thirty-two-bar AABA laid out one cell per bar, coloured by section, with the roman numeral in each bar. the A section's turnaround is the ii-V every variant keeps; the bridge is a chain of applied dominants. At 108 beats a minute in 4/4 the whole of it lasts 71 seconds. Cut into 4 repeat units of 8 bars, 3 pairs of units agree on more than 50 per cent of their bars. 2 of them are not identical, and 2 of those 2 differ in a run of bars ending at the last bar of the unit; the changed bars are marked in orange.

Where a repeat is changed

Cut every scheme into its own repeat unit, compare each unit with every other, and ask where a repeat stops agreeing with what it repeats. The answer is that it stops at the end, in every case the corpus contains — and the number of cases the corpus contains depends entirely on where the threshold for "a repeat" is put. Moving it by nothing at all takes the count from two to twenty and the finding with it.

form · Repetition
How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is.

How much evidence a modulation needs

Run a key-finder bar by bar over a progression that moves to the dominant at bar six. With four bars of history the answer becomes the new key at bar seven and holds. With three bars it never gets there at all, and reports E minor and B minor on the way. The window decides the lag as much as the music does.

perception · Key-relations
The boundary operator run over only what has been heard. Foote's checkerboard novelty on thirty-two-bar AABA at a kernel width of 4 bars, computed twice: once with the whole piece available, and once using only the bars heard up to and including each bar. The kernel reaches 4 bars forward, so every cell it needs has been heard 3 bars after its centre — the retrospective curve replotted 3 bars to the right lands on the causal one, and the operator turns out to be causal at a fixed delay rather than blind. The dashed verticals are the encoding's real section boundaries and are not an input.

An ending that can be heard coming

Two measurements are both called hearing an ending coming and they point in opposite directions. By the halfway mark of an ordinary form almost nothing new arrives — and the cost of coding each bar has not fallen at all. Neither statistic says anything is about to stop, because no statistic over content can: predicting the next event well is not predicting that there will not be one.

form · Closure
The same eight notes, read three ways. The scale, eight quavers, scored against every triad and seventh at every root. Barred as written the best reading is C major7 at 0.850; with the barline one quaver later it is D minor7 at 0.850. With no metre — every note weighted the same — 4 readings tie at 0.625 and the passage has no best analysis at all. The notes are identical in all three. What changed is where the bar starts, which is not a fact about harmony.

Which notes are the chord

A progression is a list of chords, and before there is a list something has to decide which of the notes sounding are chord tones and which are passing. Take the eight notes of a scale as eight quavers and score every triad and seventh at every root: barred as written the best reading is C major seventh, with the barline moved by one quaver it is D minor seventh, and with no metre at all three readings tie exactly and the passage has no best analysis. Same eight notes in all three. Harmonic analysis is a function of a variable that is not harmony.

harmony · Progression
Where the page ends a phrase, and where the ear does. Twinkle, twinkle with two sets of phrase boundaries on it. The lower curve is a local boundary detector — a peak in how much the interval and the note length change from one to the next, with nothing in it about bar lines or harmony — and the marks above it are where the notation puts the phrase ends. It finds 100 per cent of them and 2 boundaries the page does not have. Where the two agree it is because a long note is sitting at the join; where they disagree the page is marking a grammatical unit and the detector is finding a perceptual one.

Where a phrase ends

Run a boundary detector over the three tunes used throughout and it agrees with the notated phrasing on one of them perfectly and on another almost not at all. The reason is which cue each tune uses: Twinkle's phrases all end on a long note, so a duration-weighted detector finds five of five with no false alarms; Ode to Joy's run on in crotchets and its phrasing is in the intervals, where a duration detector finds one of three and a pitch detector finds all three and eight others. No fixed weighting serves both, and the published one is worse on each tune than the single cue that tune uses.

form · Phrase
What survives a change of encoding: verse and chorus. The same 32 bars of verse and chorus under five encodings, scored on the three things measured here measures. Mean off-diagonal similarity says how alike the piece looks to the arithmetic. Recall and precision are the novelty operator's boundaries against the 3 the section plan has, at a kernel of four bars. The period is the strongest peak of the lag profile, in bars. Under the bag of pitch classes every other figure uses, the piece is 90 per cent self-similar and the operator finds 0 per cent of the boundaries; under how far the root moved it finds 100 per cent. The period is the quantity that does not move.

The repeat that is not in the notes

Eight earlier essays compare bars by writing each one as a bag of pitch classes and taking a cosine. Nothing chose that encoding — the first used it and the other seven inherited it. Encode the same six schemes four other ways and one of the three findings survives untouched, one survives with different numbers, and one turns out to have been a statement about the encoding all along: the boundary operator finds none of the section edges in three schemes as a bag of pitch classes and every one of them as tonic, subdominant and dominant.

form · Repetition
What the joint search changes, and what it never changes. Over 552 constructed passages of eight slots with rests, how often the joint reading differs from the pipeline's. The chord differs in 29 per cent and the barline in 31, with both differing in 20. The key differs in 0 per cent — never — because the key is read from a pitch-class histogram, which does not know where the bar starts or which notes are chord tones. Two of the three decisions are entangled and the third is not.

Three decisions that constrain each other

Every model here decides one thing at a time — the key from the pitch classes, the metre from the onsets, the chords from the metre — and an earlier essay ended by saying a listener does all three at once. Resolving them jointly costs a hundred and fifty-seven times the search and changes the reading of two passages in five. It never once changes the key, and the reason it cannot is the reason the whole account is built the way it is.

harmony · Progression
The boundaries that survive each amount of smoothing. The local boundary strengths of Twinkle, twinkle read at every scale: the curve is smoothed with a Gaussian of the width on the horizontal axis and the peaks that survive are counted. Small scales give 11 boundaries and large ones give one, and the notation marks 5. The level with that many falls at a width of 2, where the model finds 100 per cent of the notated boundaries and 100 per cent of what it finds is notated — a comparison with no threshold in it, which is what the scale parameter buys.

A boundary at a stated level

A boundary detector run over three tunes agreed with the notation on one and barely at all on another, and left two things owing: a version with a scale parameter, and a version run on performance timings. Both are paid here, and they pay differently — the scale removes a free parameter from the comparison and does not rescue the hard case, while two per cent of rubato does.

form · Phrase
Twinkle, twinkle, phrased at the level each tempo selects. The number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 5 boundaries, drawn as the flat line; the model matches it best at 160 beats per minute, where the present holds 8.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.

The level the tempo chooses

The boundary detector has a scale parameter and an earlier essay left it free, ending with the sentence that names this one: the scale is in notes and the psychological present is in seconds. The psychological present is two to eight seconds, a tempo converts one to the other, and the level a listener reads then stops being a parameter at all — which is a prediction with teeth, because the same tune at two tempos should be phrased differently at levels the arithmetic names in advance.

form · Phrase
The one number the ordered key-finder was tuned on. For each rate of alternation between two keys, the cost of changing key at which the model stops hearing two keys and starts hearing borrowed chords in one. The threshold rises with the period — 0.95 at 1 bar, 0.95 at 2 bars, 0.95 at 4 bars, 2.00 at 8 bars, 3.50 at 16 bars — so the parameter and the rate trade off against each other exactly. The value tuned earlier, 2.2, sits above every threshold on this axis, which means its verdict about fast alternation was a consequence of the tuning rather than a finding about the music. Filled means the model names two keys; hollow means it names one and calls the rest borrowings.

A modulation and a borrowing are one number apart

The key-finder that keeps the order has one tuned parameter, and said so. Sweep it and the parameter turns out to be the whole verdict: below a threshold the model hears two keys alternating, above it one key with borrowed chords. The threshold rises with how slowly the keys alternate — and the value it chose sits above every threshold in range, so its finding about fast alternation was a consequence of the tuning.

scales · Key-relations
Two kinds of evidence, each measured against its own chance. Every key as a point: across, how many standard deviations its cadence count is above what the same bars resampled would give; up, the same for its pitch-class profile correlation. The winner on cadences is key C at 6.2 standard deviations and on profile C at 1.1, and they agree. Standard deviations above chance are the same unit whatever produced them, which is the commensuration this figure is for — and it costs something: it assumes the two chances are equally interesting, which is a weighting in disguise. The obvious null does not work at all for one of the two: shuffling the bars leaves a pitch-class histogram exactly as it was, so its spread is zero and every key scores nothing against it.

A count and a correlation

Cadence evidence is a count of ordered pairs and profile evidence is a correlation with a template, and the cadence essay refused to total them because they are not in the same units. Score each against its own chance and they are — standard deviations above chance are the same unit whatever produced them. Then the trouble moves: the obvious null does nothing at all to one of the two, because shuffling the bars leaves a pitch-class histogram exactly as it was.

harmony · Progression
How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.

A detector whose resolution the performance sets

The boundary detector lost its free parameter when the psychological present became a number of notes at a stated tempo, and what that held still was named at the time: a performance slows into a phrase end, so the number of notes inside the present is not the same everywhere in a tune — it falls exactly where a boundary is. Making the width follow the performance recovers half of what removing the parameter cost, and honestly leaves the other half.

form · Phrase
The passage built so the two statistics disagree. A body of chords diatonic to C major, of growing length, ended by a ii–V–I in G. The bag of notes says one key and the ordered pair says the other, which is the case the earlier figures never contained. The line is the cadence evidence for G, and under the axis is what each reading actually names at each length. The cadence reading holds G up to a body of 6 bars and is overturned at 8, so one explicit cadence is worth about that many bars of profile evidence — and it is overturned by the body's OWN incidental root motions rather than by the histogram at all. That is the finding the earlier essay could not have: on real material the two statistics cannot be varied independently, because lengthening the profile evidence adds cadence evidence too, for a third key.

The passage built to make them disagree

A cadence count and a key-profile correlation have been put on one scale and run on a passage where the two agree, which tells nobody anything. Building one where they disagree — the notes of one key and the cadences of another — measures the exchange rate at about six bars per cadence, and finds something the agreement case concealed: the two statistics cannot be varied independently, because lengthening the profile evidence adds cadence evidence too, for a third key.

harmony · Progression
How sure the reading is, bar by bar. Every earlier essay reports one best reading. A dynamic program that finds a best path has, by construction, the best score into every state at every bar — so the gap between the best reading and the best reading in any other key is already computed and has never been printed. Here it is, in bits, for the thirty-two-bar AABA. The mean margin is 1.14 bits and 13 of 32 bars are inside one bit of a rival reading, which is where a listener would be genuinely undecided. The reading itself names C, E, B, D, G; the margin says what that naming is worth, and at the weakest bar — bar 31, C over F — it is worth 0.07.

The margin the dynamic program already had

Nine earlier essays produce a single best reading, and the passages worth arguing about are the ones where two readings are nearly equally good. What is needed for that has been inside the model from early on: a dynamic program that finds a best path has, by construction, the best score into every state at every bar — so the gap between the best reading and the best reading in any other key is already computed, and printing it turns every analysis here into a measurement of ambiguity.

scales · Key-relations
The same tune read at six widths of the psychological present. A later essay made the detector's smoothing width a function of position, which removed its last free parameter but one — and the one it cannot remove is the width of the psychological present, because that is a fact about listeners rather than a choice. So the honest object is not a reading but a family of them, one per width. A short present finds 6 boundaries and a long one finds 2, and the family agrees on 0 of them. The fixed-width control, at its own best width, scores 0.67 against the adaptive readings' 0.67, 0.75, 0.33, 0.33, 0.40, 0.40 — so the adaptation does not win, which is what that essay reported too. What the family adds is the ordering: a boundary in every row is a different claim from one in a single row, and a single reading has no way to say so.

A family of readings

Removing the detector's free parameter, and then its constant tempo, cost persistence both times — the property that made its boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than one, indexed by the width of the psychological present, which is the one parameter that cannot be removed, because it is a fact about listeners.

rhythm · Phrase
A long note and a strong note disagree, and the winner is neither. The same 8 notes scored against every triad and seventh at every root, with the weighting run from the metrical one always used to a durational one never drawn. On the left each note counts for its metrical weight; on the right, for how long it is held. The long notes here are on beats 2, 4, 6, 8, which are the weak ones. The two cues point at different chords — C major7 on the left and D minor7 on the right — turning over at a mixture of 40 per cent. And at the crossing the winner is A minor7, which is neither cue's answer — a chord that shares three notes with each and is not the reading either rule asks for. Nothing about the notes changed. What changed is which of two cues a theorist would call obvious is being believed.

The long note and the strong note

The segmentation that produces every object connected here has carried a free parameter since the day it was written: whether a note counts for its metrical weight or for how long it is held. Only the first has ever been drawn. The two name different chords on sixteen per cent of passages where the cues agree about the notes and forty-three per cent where they do not — and where they disagree most sharply a mixture of them picks a third chord neither one asks for.

harmony · Progression
The reading the joint search was never offered. The best chord at each mixture of the two segmentation cues, and what the same weighting gives the same notes shuffled into a different order. Both fall along the axis, and most of the fall is the ruler rather than the music: a metrical weighting over a bar of eight spans a factor of eight and a three-to-one duration spans three, so the weighted note mass is 2.1 times more concentrated at the left of the figure than at the right, and a concentrated mass is easier for four notes to cover. What is not the ruler is the gap. It is widest at a mixture of 0.75, where the reading is D minor7 at 2.15 standard deviations above its own null, against 1.12 for C major7 at a mixture of nought. The joint search holds this axis at nought, so D minor7 is not among the hypotheses it considers.

A fourth decision, and two that were never made

The joint search resolves key, metre and segmentation together and holds the segmentation's cue mixture at zero. Adding the mixture is one loop, and reading the search in order to add it turns up something worse than a missing axis: on the passages it is drawn on, the key it reads is the same key at all forty-eight of its hypotheses and the metre scores every barline identically. The fourth axis then cannot be ranked at all until each reading is measured against its own null, because a mixture changes the ruler and not only the answer.

harmony · Progression
How often the metre, the chords and their product find the barline, chords at 1. Constructed passages of four bars of eight quavers, 100 at each setting, with the barline at the first slot. Rhythm regularity is how much likelier a note is on a strong slot than a weak one; chord regularity is how much likelier a note is to be a tone of its bar's chord than a random scale tone. rhythm 0: metre finds it 10%, chords find it 41%, product finds it 16%; rhythm 0.25: metre finds it 34%, chords find it 35%, product finds it 56%; rhythm 0.5: metre finds it 49%, chords find it 21%, product finds it 66%; rhythm 0.75: metre finds it 50%, chords find it 17%, product finds it 56%; rhythm 1: metre finds it 50%, chords find it 11%, product finds it 45%.

The chords never move the barline

Every hypothesis the joint search had drawn was one bar long, and on one bar with a note in every slot the metre cannot choose a barline at all. Four bars with rests in them make the barline a decision the metre and the chords both have an opinion about, and the prediction was that the chords would move the barline more often than the barline moves the chords. It is the other way round, completely: whenever the two prefer different barlines the search takes the metre's, on up to 72 per cent of passages, and in fifteen hundred passages the chords never once move it. What the chords decide is the one thing the metre cannot see — whether the bar starts on the downbeat or half a bar later — and they decide it right a little over two times in three at best.

harmony · Progression
The product and the sum of standard scores, finding the barline, chords at 1. Constructed passages of four bars of eight quavers, 100 at each setting, with the barline at the first slot. Rhythm regularity is how much likelier a note is on a strong slot than a weak one; chord regularity is how much likelier a note is to be a tone of its bar's chord than a random scale tone. rhythm 0: product finds it 16%, sum of z finds it 22%, metre finds it 10%, chords, z 35%; rhythm 0.25: product finds it 56%, sum of z finds it 58%, metre finds it 34%, chords, z 38%; rhythm 0.5: product finds it 66%, sum of z finds it 61%, metre finds it 49%, chords, z 19%; rhythm 0.75: product finds it 56%, sum of z finds it 58%, metre finds it 50%, chords, z 14%; rhythm 1: product finds it 45%, sum of z finds it 48%, metre finds it 50%, chords, z 12%.

The chords are a weak witness to the barline

Scaled by its own range, the metre overrules the chords every time the two disagree about where a bar begins. The obvious repair is to score each reading against its own chance — the metre against the same number of notes placed at random, the chords against the passage's notes shuffled across its bars — and add the standard scores. It changes very little: the search finds the barline within six points of where the product found it, and the chords gain the power to move the barline only on passages whose rhythm says nothing, where random notes move it nearly as often. The null's real result is the size of the two witnesses. At the written barline the metre stands up to 5.9 standard deviations above chance, and the chords, with every note a tone of its bar's chord, stand 1.55 above it at best.

harmony · Progression
Four bars read by where the chords change, barline by barline. A constructed passage of four bars of eight quavers, its barline at the first slot and its chords C, F, Em, Dm. Notes: slot 1 C, slot 3 E, slot 4 G, slot 5 E, slot 6 C, slot 9 C, slot 10 C, slot 11 F, slot 13 A, slot 16 C, slot 17 B, slot 19 B, slot 20 E, slot 21 E, slot 24 G, slot 25 F, slot 29 F. For each of the eight places the barline could fall: as written metre, z 3.97, chords, z 2.00, change, z 4.30, metre + change, z 8.27; 1 quaver late metre, z -2.45, chords, z 2.14, change, z -0.87, metre + change, z -3.32; 2 quavers late metre, z -1.38, chords, z 2.15, change, z -0.54, metre + change, z -1.93; 3 quavers late metre, z -0.31, chords, z -0.27, change, z -2.64, metre + change, z -2.95; 4 quavers late metre, z 3.97, chords, z -1.63, change, z -4.30, metre + change, z -0.33; 5 quavers late metre, z -2.45, chords, z 1.27, change, z 0.87, metre + change, z -1.58; 6 quavers late metre, z -1.38, chords, z 1.18, change, z 0.54, metre + change, z -0.84; 7 quavers late metre, z -0.31, chords, z 1.05, change, z 2.64, metre + change, z 2.32. Best metre, z: as written and 4 late. Best chords, z: 2 late. Best change, z: as written. Best metre + change, z: as written.

The chords mark the barline by changing there

Read bar by bar, the chords stood barely above chance at the barline and broke the metre's half-bar tie two times in three at best. Read instead by where they change — how different the chords are across a candidate's barlines against how different they are across the middle of its bars — the same notes break the tie right on 81 to 96 per cent of passages, and added to the metre they find the barline on up to 89 per cent against 61. The weakness was the question the old reading asked, not the harmony.

harmony · Progression

Named alongside it

The objects these essays reach for when they reach for this one.

Harmonic analysisKey-findingProgressionRepetitionSelf-similarityMetreModulationMetrical weightNoveltyPhraseCadenceEvidence

All concepts