A return is shorter than its first hearing
Assumes: A proportion is only as fine as its two durations · How much of this is new
A proportion is only as fine as its two durations took the durations of a form’s sections at face value, as clock readings, and asked how finely a listener could compare them. It found that across spans of minutes the answer is two categories — about equal, clearly unequal — and it ended on the assumption every one of its figures made: that a section is remembered as the time it took.
It is not, and forms with returns in them are where the difference is largest. A rondo’s refrain comes round three times and takes three fifths of the clock. The second and third times, a listener who has heard the first already has it. A verse-and-chorus song is exactly balanced, eight bars against eight, twice over — and its second half is made entirely of material the first half supplied.
So there are two proportions for every such form, one measured by the clock and one by what a listener has to keep, and the question is how far apart they are. The answer turns out to depend less on the forms than on what kind of memory is doing the keeping, and two reasonable kinds disagree about it completely.
Two things a listener could be storing
The forms here are the six bar-level chord schemes the form essays have used throughout: a twelve-bar blues played three times, a thirty-two-bar song in AABA form, a rondo in ABACA, a verse-and-chorus song, a sixteen-bar period, and a four-bar ostinato played eight times. Each bar is a chord in a key. Each form is cut into its own repeat unit — the highest common factor of its section lengths and its chorus, four bars for the blues and the ostinato and eight for the rest — so that two consecutive A sections are two occurrences, and a return is any occurrence identical to one heard before.
The clock measure needs no model: a section’s share is its share of the bars.
The storage measure needs one, and the simplest honest one is a model of recognition. A bar costs a listener something to keep only if it could not have been predicted when it arrived — and whether it could have been predicted depends on whether that bar, preceded by the bars before it, has occurred before in the piece. At a bar is stored only if its chord is new. At it is stored if the pair of chords ending on it is new, which is the lower strip above. Larger asks for a longer run of agreement before a return is accepted as one, and it is swept rather than chosen.
A context of one bar is not a guess about the right size. It is a conservative one. At 108 beats a minute a bar lasts 2.2 seconds, so one bar of context fits inside the span a listener holds as a single present with room to spare, and four bars — 8.9 seconds — is already past the top of it.
By the clock the returns are the form, and by storage they are a margin
The upper strips are the proportions an analysis quotes. In four of the six forms the returns take a quarter or more of the clock, and in two of them — the blues and the ostinato — they take most of it.
The lower strips are nothing like them. The rondo’s two returns of the refrain, eight bars each, cost one bar apiece: the first bar of each return is new, because what precedes it is an episode rather than the refrain, and the seven bars after it are the refrain’s own pairs in the refrain’s own order. The verse-and-chorus song keeps five bars of its first verse, two of its first chorus, one of its second verse and nothing at all of its second chorus. The blues keeps seven bars of its first chorus, one of its second and none of its third.
The shape a form has in storage is front-loaded by construction. Everything a return contains was paid for when it was first heard, and a return costs only its seams — the bars where it meets material it did not meet before.
That is the whole mechanism and it is worth being plain about how little is in it. Nothing here knows what a refrain is or that a return is structurally important. It asks one question of every bar — has this happened, in this context, already? — and the forms answer it the way they were built to.
How much recognition asks for
The obvious objection is that one bar of context is too easy: a listener who hears two chords and concludes they are back in the refrain has jumped to a conclusion. The conclusion is a parameter, and it can be made as demanding as anyone likes.
Demanding a longer context can only make a return cost more, since every run of five bars heard before contains a run of four heard before, and the curves do rise. They rise slowly. At four bars of context — a whole four-bar phrase that has to match, plus the bar arriving — the rondo’s returns take 25 per cent of what is kept against 40 per cent of the clock, the verse-and-chorus song’s 24 against 50, the blues’ 25 against 67, and the thirty-two-bar song’s final A section 15 against 25.
What a longer context changes is the width of each seam, and the strips show where the width goes. The verse-and-chorus song keeps 47 per cent of what it stores in its first verse, 29 in its first chorus, 24 in its second verse, and still nothing in its second chorus, because by the time the second chorus begins the second verse has supplied exactly the four bars of context its first chorus had. The blues keeps a quarter of its storage in each section of its first chorus and a quarter in the first section of its second, and nothing after that. The rondo’s two returns cost four bars each.
So the longer context does not move storage toward the clock evenly. It charges the first return of each section and leaves the later returns free, because a later return arrives in a context an earlier return already built.
The ostinato is the instructive exception. Its eight choruses are one four-bar pattern of a single chord, so at every context up to three bars the whole pattern is recognised from the second chorus on and the returns cost nothing. At four bars the context of the second chorus’s first bar reaches back into the start of the piece, where there was nothing before it, so that one bar is stored — a fifth of everything the ostinato costs, charged to a return. That is an edge effect of a piece with no introduction, and it is the only point on the figure where a longer context moves a form’s returns from nothing to something.
At no context any listener could use do the returns get back to their clock share. Recognition needs its seams, and a seam is one or two bars in eight.
The memory that relearns
A context model is one reading of what is stored. There is another, and it is the one how much of this is new measured the same six forms with: an LZ78 coder, which builds a dictionary of phrases as it goes and pays for each new phrase by the size of the dictionary it has built.
The difference that matters here is how each memory treats a return. The context model recognises a return as soon as its context matches. LZ78 cannot recognise anything longer than a phrase it has already stored, and it stores each phrase by extending an old one by a single symbol. So an eight-bar refrain heard a second time is not one recognition but several new phrases, each a little longer than one the coder knew, each charged at a price that rises as the dictionary grows.
Under the coder the lower strips are very nearly the upper strips. A memory that learns repeats one bar longer each time it meets them charges a return almost exactly what the clock charges it, and on the ostinato — eight identical choruses — it charges each of the last three choruses more bits than the first, because each new phrase is indexed in a larger dictionary.
That is not a defect of the coder. It is what LZ78 is: a universal compressor that makes no assumption about the source, and whose efficiency is guaranteed only in the limit of long sequences. The same essay found that at a musical length its ratio sits above one and concluded that the number is a statement about the description rather than the music. The section shares say the same thing from the other side. The coder is a model of a listener who has to be told a passage several times, at increasing length, before it can repeat it back; the context model is a model of a listener who knows it after once.
Balanced by the clock, and by nothing else
The difference between the two memories becomes a difference in proportion the moment it is asked of the two halves of a form. By the clock, every one of the six forms is exactly balanced — its first half is its second half’s length. By storage, almost none is.
The verse-and-chorus song is the clearest case. Its two halves are identical — verse and chorus, then verse and chorus again — and a listener recognising from one bar of context keeps seven times as much of the first half as of the second. With four bars of context the ratio is 3.2 to 1. Both are further from 1 : 1 than 3 : 1 is, and the resolution arithmetic found 1 : 1 and 3 : 1 separate at every Weber fraction reported for a judgement over minutes. Whichever storage a retrospective judgement follows, the clock’s answer and the context model’s answer are in different categories.
Under the memory that relearns, the verse-and-chorus song is balanced after all, and so is every other form. The two memories do not disagree about the size of an imbalance; they disagree about which of the two categories a listener over minutes has — about equal, clearly unequal — the form belongs in. That turns the question of which memory a retrospective judgement follows into one with a sharp answer rather than a matter of degree. A listener asked afterwards whether the second half of a verse-and-chorus song lasted as long as the first would say yes under one memory and no under the other, and the resolution arithmetic says the difference between those answers is one a listener can report.
The rondo and the period lean slightly the other way under the coder, and for a reason that belongs to the coder rather than to the forms: their second halves hold the minor episode and the changed cadence, and a coder whose dictionary has grown charges every new phrase more than it charged the same phrase earlier.
The thirty-two-bar song goes the other way, and the reason is its bridge. Its first half is A and A again, and the second A is almost free; its second half is the bridge, which is entirely new, and a closing A whose last three bars differ from the others. So by storage the song is heavier at the end, 1 : 1.4 with one bar of context — a ratio inside the blur band, and so, on the evidence of the resolution arithmetic, balanced to a listener after all.
The sixteen-bar period is the case that shows how much the context parameter can do. Its consequent repeats its antecedent’s first six bars and changes the cadence. With one bar of context, all but the seam and the changed cadence are recognised and the halves stand at 4 : 1; with four bars of context, the consequent’s opening bars are preceded by the antecedent’s close — a run of four bars never heard before — so they are stored as well, and the ratio falls to 1.3 : 1. A varied repeat is where the choice of memory decides the proportion, and a literal one is where it barely matters.
What the storage-size account adds, and what it does not
None of the storage measures is a claim about how long a listener thinks a section lasted. The claim that connects them is older than any of them. Robert Ornstein proposed in 1969 that the remembered duration of an interval grows with the amount of information stored from it — his storage-size hypothesis — and later accounts refined amount of information into the number of changes or segments a listener registered. The pattern that survives in the reviews, among them Richard Block and Dan Zakay’s meta-analysis, is a split by paradigm. When a listener knows in advance that duration will be asked about, judgements follow attention to time. When the question comes as a surprise, afterwards, they follow what memory holds, and a more eventful interval is judged longer.
On that account a verse-and-chorus song is a different length depending on when the question is asked. While it plays it is balanced, and after it ends it is front-loaded, by a factor the figures above put between three and seven — as long as recognition works the way the context model says.
What the account does not do is choose between the memories. If a retrospective judgement follows something like the LZ78 bits, the song is balanced afterwards too, and nothing about the storage-size hypothesis prefers one reading. The evidence that would decide it is about listeners rather than coders: whether a second chorus, once recognised, adds to the remembered length of a song, or whether only its seams do.
What the model assumes
That a bar is a chord. Every measure here reads a bar as its harmony and its key, which is what the schemes encode. A real second verse has different words, and often a different melody over the same chords, and every one of those differences is something new to store. The storage shares here are a floor for a literal return and an underestimate for a varied one.
That recognition is exact. A bar is recognised if its context was heard before, identically. A listener who half-recognises a passage stores some of it, and the model has no half. The context parameter is the only dial between recognising too easily and not recognising at all.
That nothing is forgotten. A refrain heard a minute ago and a refrain heard ten seconds ago are recognised alike. A return has to be remembered put a decay on the comparison between two bars and found the ranking of these same forms changed, and a return whose first hearing has faded costs more than one that has not. That would raise the returns’ share in the long forms and leave the short ones where they are.
That the listener arrives knowing nothing. Every piece here starts from an empty memory. A listener who knows the song, or the style, has most of the first verse already, and the first half becomes cheaper too. What happens to the ratio then depends on which half the prior knowledge covers, and the model does not say.
What a strip of stored bars cannot establish
That a listener compares halves at all. The resolution essay’s warning holds here with full force: the evidence that listeners attend to the large proportions of a form is thin, and a storage ratio of seven to one is only a remembered proportion if somebody remembers a proportion.
What the returns are for. A refrain costs a margin of storage and it is the part a listener sings on the way out. A form that is front-loaded in storage is not thereby worse, and the arithmetic here is silent about value. Rondo and verse-and-chorus forms are built for their returns, and a return that costs nothing to keep is plausibly the reason they work — which is a different claim and not one a strip can make.
Anything about a first hearing inside a first hearing. Recognition here is decided with the whole of the preceding piece available, in order. Nothing about when the listener recognises a return — after one bar, after four — enters the proportions, only whether they do. The form a first hearing cannot have is about that other question, and its answer is that some periods are not knowable until the piece is nearly over.
Whose forms, and the ones where this bites
The forms where the clock and the storage measures part furthest are the ones built from literal repetition: strophic songs, twelve-bar blues played chorus after chorus, dance forms whose sections repeat by sign, and ostinato music of every kind. The forms where they part least are through-composed, where nothing returns and every bar is new under every memory — though that case is not in these six schemes, and the claim for it is arithmetic rather than measurement.
The classical repeat sign is the sharpest instance, because it is a composer writing the second hearing into the clock on purpose. An exposition played twice doubles its share of the clock and, on the context model, adds one bar of storage for the seam where the repeat begins. A repeat sign adds time and almost no content, which is the reason performers argue about whether to take it and a question the storage strips can price but not settle.
The cyclic traditions stand at the far end. A cycle that outruns the memory found a gong cycle of forty seconds long enough to exceed what a listener holds, and the ostinato here is its limit case at the level of chords: a piece whose every chorus after the first is free to keep. What such music gives a listener to store is not in its repetitions at all, and a cycle cannot cadence found the density curve standing where a cadence would be.
Still open: whether counting restores the clock
Both of the proportions here, by the clock and by storage, are quantities a listener estimates. There is a third way to know how long a section was, and it is not an estimate: counting it. A listener who has induced a four-bar hypermetre knows that a verse was eight bars long, and knows it exactly, whatever the verse contained and however familiar it was. The resolution essay found a one-bar phrase extension invisible to timing and plain to a count. What is owed now is the cost of the count itself — each counted unit a chance to slip — and whether, set against the blur of timing and the bias of storage, counting in beats, in bars or in phrases gives a listener a proportion close to the one on the page.
Part 2 of 6
One essay in the series on proportion. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Description lengthDurationMemory decayMusical formProportionRedundancyRefrainRepetition
- The part of the tune that is kept description length, memory decay, redundancy
- A piece is mostly itself again refrain, repetition
- An ending that can be heard coming description length, redundancy
- An expectation cannot rescue a cycle too slow to time duration, memory decay
- How long until it comes back refrain, repetition
- Repetition buys least where it is needed most memory decay, repetition