The part of the tune that is kept
Assumes: The shape that survives everything else · A return has to be remembered
The third rung of this ladder found that a melody’s contour — the bare sequence of ups and downs, with no interval sizes in it — survives transposition, a change of tuning, a doubling of every interval and a change of instrument. It survives them because it is a coarser description than any of those transformations is fine enough to disturb.
The reason usually given for caring about it is that contour is what a listener keeps. That is a claim about memory — and what a listener forgets between two hearings is measurable — it is repeated everywhere, and it was left as an assertion when the rung closed. It is not an assertion that has to be left alone: a code has a size, memory has a capacity, and both are numbers.
What a shape costs
Take a melody of n notes over eight scale degrees. There are eight to the n of them and each therefore carries three bits a note. Its contour is a sequence of n − 1 signs, each up, down or level, and if all three were equally likely that would cost log₂3 = 1.585 bits a note.
They are not equally likely, and the difference is the first thing worth measuring.
Ninety-six against two hundred and forty-three is a substantial discount and it comes from a fact about walks rather than about music. A sequence of pitches drawn at random from eight degrees produces alternating contours far more often than monotonic ones: the two strictly alternating six-note shapes occur 15,106 times each in the enumeration and the all-level shape occurs eight times. The distribution is a binomial in disguise, and an entropy is what accounts for it.
So the shape of a tune costs less to store than the count of shapes implies, and the amount less is 1.33 bits at six notes. That is the discount a code gets for free when it is a code over something with a shape.
The length at which a shape identifies
If a contour costs 1.28 bits a note, the number of tunes it can distinguish grows by a factor of 2.4 per note. That is a slow growth and it is the interesting quantity, because it says how much of a melody a listener has to have heard before its shape is a name for it rather than a description of it.
At six notes the effective alphabet is ninety-six. At seven it is 234; at eight, 568. Extrapolating the same 1.28 bits a note, at nine notes it is about 1,380 and at ten about 3,350.
Nine notes of contour is enough to pick one tune out of a thousand.
That number wants two checks before it is worth anything, and both are available.
The first check is against a listener’s capacity. Short-term memory for arbitrary material is conventionally put at seven or so chunks, and for material of this kind the usual estimate of the capacity in bits is somewhere between ten and twenty. Nine notes of contour is 11.6 bits. That is inside the range, and near the bottom of it — which is the right place for it to be, because a listener has to do something with the information after storing it.
The second check is against the phrase, and it is the one this collection can make from its own measurements.
How long a phrase lasts across a range of tempi and metres is two to eight seconds, which at ordinary speeds is four to eight notes — so the length at which a contour begins to identify a tune and the length a listener can hold as one present event are the same length, arrived at from two directions.
Nine notes at a hundred and twenty beats a minute is four and a half seconds. Ten at a hundred is six. Both are inside the window, and both are near the middle of it.
So the length at which a contour becomes identifying and the length of a phrase are the same length, and the two were arrived at from entirely different places: one from an enumeration over sequences with no listener in it, the other from tempo and metre data with no combinatorics in it. That is the result of this rung, and it says what a phrase is for: a phrase is the unit over which a shape becomes a name.
The two codes are worth holding apart, because this ladder has used both and called them both contour. The class is nine values and answers “what shape is this”; the sign sequence is 2.4 to the n and answers “which tune is this”. Only the second is a candidate for what a listener stores, and only the first is what the census results were about.
Why a shape and not the intervals
The obvious objection is that a listener could store the intervals instead, which would be more informative at three bits a note rather than 1.28, and would identify a tune in four notes rather than nine.
They could, and there is a good deal of evidence that they partly do — trained listeners retain interval information that untrained ones do not, and everybody retains more of it for a familiar tune than an unfamiliar one. The question is what a first hearing keeps, and the case for contour is that it is what survives the things a first hearing has to survive.
Put in terms of the count: a contour is a lossy compression whose losses are exactly the parameters that vary between performances. Transposition is a global offset and contour discards offsets. A retuning of up to twenty-two cents is a perturbation of each interval and contour discards sizes. Doubling every interval is a scaling and contour discards scale. The 63 per cent thrown away is not thrown away because it is unimportant; it is thrown away because it is the part that will be different next time.
There is an experimental result here that this collection can describe and cannot check. Listeners asked whether two short melodies are the same, when the second is transposed, make a characteristic error: they accept a transposition that preserves the contour but alters the intervals, and reject one that preserves the intervals and alters the contour — and the error is much stronger for unfamiliar tunes and for short delays. That is the finding the whole “contour is what is retained” claim rests on, it is quoted here, and nothing on this site measures anything of the kind.
What this site can add is the counting above, and the counting says the claim is at least possible: a contour is small enough to be held, specific enough at the length of a phrase to be a name, and invariant under the transformations a second hearing applies. Any of those three failing would have made the experimental result hard to believe. None of them fails.
A random walk with a pull toward the middle produces contours with the same statistics as real tunes over four or five notes, which is the caution the enumeration supplies for free: a shape that a tune shares with a random walk is not evidence that the tune was remembered by its shape.
What the uniform assumption costs, measured
That figure also supplies the test of this essay’s largest assumption, and running it reverses a claim made below.
Every count above is uniform over sequences of scale degrees: each note drawn independently and with equal probability from the eight. Real melodies are not drawn that way, and the caveat at the end of this essay reasons that a code fitted to their real distribution would be shorter than 1.28 bits a note, so that the figures here are an upper bound on what a contour costs.
The site has a better melodic distribution than uniform and it is the one drawn above: a centred walk with the pull the leap rung fitted to reproduce post-skip reversal. Sampling contours from it instead of from the uniform enumeration gives, at the same four lengths:
| length | uniform | centred walk |
|---|---|---|
| 4 notes | 4.02 bits (16 shapes) | 4.32 bits (20) |
| 5 | 5.31 (40) | 5.75 (54) |
| 6 | 6.59 (96) | 7.17 (144) |
| 7 | 7.87 (234) | 8.58 (382) |
It is higher at every length, not lower — 1.420 bits a note against 1.282. The uniform enumeration is a lower bound on what a contour carries, and the caveat has the direction wrong.
The reason is in this essay’s own third paragraph. Independent uniform draws over eight degrees produce alternating contours far more often than monotonic ones, because a note drawn high is likely to be followed by a lower one whatever came before; that is a negative correlation between consecutive signs, and a correlation of any kind lowers entropy. A melodic walk takes small local steps, so a run of them is not improbable, the signs are closer to independent, and the distribution over shapes is flatter. Making the melodies more realistic makes their contours more varied rather than less.
And it can be enumerated one length further than the uniform census reaches. At eight notes the walk’s effective alphabet is 988 shapes — measured, not extrapolated — against the 568 the uniform slope predicts. Eight notes of contour, over a melodic distribution with real melodic constraints in it, distinguishes almost exactly one tune in a thousand.
That moves the headline in the reassuring direction, which is worth being suspicious of and does not survive being suspicious of, because the mechanism is stated above and does not depend on the answer. The identifying length falls from 8.6 notes to 8.0 for a repertoire of a thousand, and from 11.2 to 10.3 for one of ten thousand. The coincidence this rung rests on — that the length at which a shape becomes a name is the length of a phrase — is unchanged and slightly tightened, since eight notes at a hundred and twenty is four seconds and sits nearer the middle of the psychological present than nine did.
What does change is the epistemic status of the claim. It was resting on an enumeration whose distribution was known to be wrong and assumed to be wrong in a safe direction; it is now resting on one that is still wrong, in the opposite direction from the one recorded, and bracketed on both sides. Two distributions a long way apart give 8.0 and 8.6, which is a much better reason to believe “eight to eleven notes” than either of them alone.
The other memory, which is longer
Everything above is about the seconds. A tune also has to be recognised across days, and that is a different store with a different characteristic.
The long-term store is not a contour, and the evidence is that people who cannot sing a tune in tune can nevertheless recognise it from a recording in the right key — absolute pitch information survives in long-term memory for familiar recordings, which is a result the pitch-standard ladder already carries. So the coarse code is a first-hearing code, and what replaces it with repetition is finer.
That gives the two rungs a division of labour worth stating. Contour is what makes a tune recognisable on a second hearing; something much more specific is what makes a recording recognisable on the hundredth.
At the scale of a whole piece the currency changes: repetition rather than shape, measured as how much a coder saves as the sequence lengthens, and a boundary detector run over a movement finds section edges rather than phrase shapes. Both are memories and neither is this one.
The two scales are worth putting side by side because the arithmetic is the same and the conclusion is opposite. At the scale of a phrase the coarse code is enough: nine signs pick a tune out of a thousand. At the scale of a movement it is not, because the number of things to be distinguished does not grow but the amount of material does, so the listener’s problem stops being identification and becomes segmentation. That is the boundary between this ladder and the repetition ladder, and it falls at about the length of the psychological present.
Which computation produced the numbers
Every count is an exhaustive enumeration, not a sample. At six notes over eight degrees that is 262,144 sequences; at eight notes it is 16,777,216, which takes about three seconds and is why the figure stops at seven and the eight-note figure is quoted in the text.
The entropy is the Shannon entropy of the contour distribution the enumeration produces — sum of −p log₂ p over the distinct contours, with p the fraction of sequences having each one. The “effective alphabet” is two to that entropy, which is the number of equally likely shapes that would cost the same.
The 1.28 bits a note is the slope of entropy against length over the four lengths computed, and it is very nearly constant: 1.28, 1.29, 1.28 between consecutive pairs. The extrapolation to nine and ten notes uses that slope and is an extrapolation rather than a computation; the eight-note value it predicts, 9.15 bits, matches the enumerated one to two decimal places, which is the check that makes the extrapolation worth quoting.
The walk comparison is a sample rather than an enumeration, because the space it draws from is not finite in the same way. Four thousand centred walks of eighty notes each, cut into non-overlapping windows of the length being measured, give between fifty and eighty thousand contours per length — enough that the entropy is stable in the third decimal across seeds, which is the only accuracy the comparison needs. Non-overlapping windows matter: sliding them would count each note in several contours and correlate the samples, which inflates the apparent variety.
The pull is 0.6, the value the leap rung fitted so that the walk reproduces the observed rate of post-skip reversal. It was fitted for a different purpose and is used here unchanged, which is what makes it a second distribution rather than a parameter tuned until the answer came out well.
The memory-capacity figures are quoted from the literature and are not computed here. So is the claim that trained listeners retain more interval information. Nothing in this essay measures a listener.
The phrase durations come from this collection’s own phrase ladder and are computed from tempo and metre, not from a corpus of performances.
What the picture cannot show
It assumes every melody is equally likely, which is the largest weakness. The enumeration is uniform over sequences of degrees and real melodies are not. This was previously recorded here as a safe assumption — a code fitted to the real distribution being shorter than 1.28 bits a note, making these figures an upper bound. The section above computes it against the site’s own melodic walk and finds the opposite: 1.42 bits a note, so the uniform figure is a lower bound. The centred walk is not a real corpus either, and what the two together give is a bracket rather than an answer.
It has no rhythm in it. A tune is recognised from its rhythm at least as readily as from its shape — “Happy Birthday” is identifiable from a rhythm alone — and a joint code over contour and duration would identify a tune in far fewer notes than nine. The duration debt was paid one rung ago and this rung does not use the result.
It is a code, not a mechanism. Nothing here says a listener stores bits, and the coincidence between the phrase length and the identifying length is a coincidence between two numbers rather than evidence about a process. The honest version of the result is that the two quantities are the same order of magnitude and were arrived at independently.
It has no harmony. A tune heard over chords is a different object from a tune heard alone, and the chords supply a great deal of the identifying information — how much of a passage is new is measurable over harmony as readily as over pitch, and a listener uses both.
And nine notes is a repertoire of a thousand. A listener who knows ten thousand tunes needs eleven, and one who knows a hundred needs seven. The result scales as the logarithm of the repertoire, which is slow enough that the answer is “eight to eleven notes” for any repertoire anybody has.
The ladder from here
This rung paid the second of the ladder’s three debts. The last one is the width: the rung about a tune’s register set out to explain why a tune fits inside about a twelfth, found that the singer’s register break does not explain it, and left the question open. The answer turns out to be in the same parameter the leap rung fitted and did not use again — a walk with no walls whose central tendency reproduces post-skip reversal has a range that grows logarithmically, and over the whole plausible length of a tune that range sits between an octave and a fifteenth.
Part 7 of 8
One essay in the series on melody. The essays either side of this one:
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
ContourDescription lengthEnumerationMemory decayPerceptual presentPitch memoryRedundancyRelative pitch
- A return is shorter than its first hearing description length, memory decay, redundancy
- An ending that can be heard coming description length, redundancy
- How long a limping bar can be enumeration, perceptual present
- Knowing every metre is slower than knowing none enumeration, perceptual present