A melody is a walk, not a set
Assumes: A voice is a stream, and the ear decides which · Seven of the twelve, chosen unevenly
One hundred and sixty-six essays about why music sounds the way it does, and the object a listener would say music is — a tune, the thing that gets whistled — has not appeared in any of them as a subject in its own right.
The reason is visible in what the collection is good at. A scale is a set and sets can be counted: all 349 shapes in twelve, every universe from four to thirty, the sizes at which a chain of fifths closes. A chord is a simultaneity and simultaneities can be scored. An interval vector is a multiset. A rhythmic necklace is defined up to rotation. Every one of those objects is the same read forwards and backwards, and that symmetry is exactly what made them enumerable.
A tune is not symmetric in time, and it is the first object here that is not.
The set and the path
Start with the smallest possible demonstration, which needs no theory at all.
Drawn as a necklace the tune is five degrees and nothing else — every property the rest of this collection computes returns the same answer for it as for any other arrangement of those five. A walk has boundaries and a set does not, and that is the cheapest demonstration of the difference.
The sorted version is the set. It has the same pitch classes with the same frequencies of occurrence, the same range, the same interval vector, the same everything this collection has ever measured about a scale — and it is not recognisable, because none of what makes the theme a theme is in any of those quantities.
The difference between the two is entirely in the ordering, and the ordering has a measurable property. Sorted, the mean step between adjacent notes is 0.24 semitones, because a sorted list mostly repeats. As written, it is 1.24. What is worth noticing is how small the second number still is — smaller than the difference between the two step sizes of the diatonic set, and smaller than most of the intervals the scale it walks on actually contains: a tune whose notes span a fifth moves, on average, by rather less than a whole tone at a time.
Almost all of it is steps
That turns out not to be a peculiarity of one theme.
Beethoven’s theme is the extreme case and the extremity is exact: its largest melodic interval anywhere in eight bars is two semitones. Thirty notes, twenty-nine moves, and not one of them is bigger than a whole tone. That is a well-known fact about the tune and it is worth stating in the form the figure gives it — the theme does not merely favour steps, it contains nothing else.
The nursery tunes are less austere. Twinkle’s opening is a fifth, and Frère Jacques drops a fifth at its cadence and rises a fourth in its third bar. Even so, 91 per cent of Twinkle’s moves are two semitones or fewer, and the leaps are few, large and structurally placed — at the start of a phrase, at the end of one.
This is the shape everyone who has looked at melodic corpora reports, in every repertoire that has been counted. The question this essay is about is why, and the answer usually given — that steps are easier to sing, or more pleasing — is not a bad answer but it is not the binding one.
There is a second reason to look at the differences rather than the agreement. Twinkle’s distribution has a gap: 37 per cent of its moves are two semitones, 2 per cent are five, 7 per cent are seven, and nothing at all lies at three, four or six. That is not a smooth falling-off; it is a stepwise regime with a few structural leaps and an empty band between them. Frère Jacques, which has more leaps, fills the band — 16 per cent of its moves are four semitones, which is the third that the triad ladder is entirely about, and its melody is a broken chord for most of two bars.
So “melodies are stepwise” is really two claims that ought to be separated. One is that the mean move is small. The other is that the distribution is bimodal — a dense cluster of steps and a sparse tail of leaps, with a thin region between — and the second is the more interesting, because a single smooth distribution would be what a preference for small intervals produces and a bimodal one is what two different mechanisms produce.
The second claim can be given a test even on a corpus this small, and it survives one. Pool all three tunes — 101 moves — and the shape is not a decay:
| interval | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| share of moves | 31% | 12% | 43% | 1% | 5% | 5% | 0% | 4% |
Eighty-six per cent of everything sits at nought, one or two semitones; then the count falls off a cliff, and the bottom of the cliff is at three semitones, where there is a single move in a hundred and one. A minor third is the rarest interval below an octave in this corpus — rarer than the fourth above it, rarer than the fifth above that. A smoothly decaying preference for small intervals cannot produce a minimum with larger values on both sides of it.
Twinkle’s empty band is the weaker half of the evidence and the arithmetic says so. Draw 41 moves from the pooled distribution and the chance that none of them lands at three, four or six is 8.2 per cent — unremarkable, which means one tune’s gap is not independent evidence of anything. Draw the same 41 from a smooth geometric decay with the same mean of 1.72 semitones and the chance is 0.04 per cent.
That is the right shape for the claim. The gap in one tune proves nothing on its own and refutes a smooth model decisively, and what carries the bimodality is the pooled trough at the minor third rather than any single melody’s silence.
Three tunes and a hundred moves remain an illustration. What has changed is that it is now an illustration with a falsifiable statement attached to it, and the statement a larger corpus should be asked is where the trough sits rather than whether the mean is small.
The ear has a rule about it, and the rule has numbers
There is a constraint on melodic leaps that has nothing to do with taste, singers or style, and it is one this collection has already used for something else.
A sequence of tones that alternates between two pitch regions splits into two streams when the regions are far enough apart and the sequence is fast enough. Van Noorden measured the two boundaries of this in the 1970s and they are not the same boundary: below the fission limit a sequence cannot be heard as two however hard a listener tries, and above the coherence limit it cannot be held as one.
Van Noorden’s two boundaries against the time between tones enclose a region where the percept is genuinely ambiguous — a listener can hear one stream or two and can switch at will. Everything in this essay sits inside or below that region, which is why the walk is a walk: at ordinary melodic rates the notes are one line, and the set-like description only becomes available when the rate pushes the percept apart.
Read that curve as a statement about melody rather than about laboratory sequences and it says something sharp. The largest interval a melodic line can contain and still be one line depends on how fast the line is going:
| Notes a second | Largest interval that stays one stream |
|---|---|
| up to about 6 | past the top of the measured range |
| 8 | 9.1 semitones — a major sixth |
| 10 | 6.6 — a tritone |
| 12 | 5.5 — a fourth |
| 16 | 4.5 — a major third |
At a slow rate the constraint does not bite at all: a singer taking two notes a second can leap two octaves and the listener will follow. At the speed of a fast violin passage the ceiling has come down to a fourth.
The rate at which the choice disappears
The two boundaries converge, and where they meet is the hardest number in this essay.
At sixteen notes a second the coherence limit is 4.5 semitones and the fission limit is also 4.5. The bistable region — the whole area in which a listener may decide whether to hear one line or two — has closed to nothing. Above a major third the sequence splits and there is no choosing otherwise; below it, it fuses and there is no choosing otherwise either.
The consequence for melodic writing is a hard one and it is not a stylistic observation. A fast passage that leaps is not a fast passage that leaps; it is two slower lines. Composers discovered this and gave it a name — compound melody, or the implied polyphony of an unaccompanied Bach line — and the technique consists of doing it on purpose, writing one line of notes that the ear separates into two voices because the leaps are large and the rate is high.
That is the same phenomenon, used rather than avoided. What the boundary explains is the default: a line that is meant to be one line, at speed, has to move in steps, and the size of step available shrinks as the tempo rises.
There is a corollary worth stating because it inverts the usual advice. A leap is not made safe by being approached slowly; it is made safe by the notes either side of it being slow. The boundary is a function of the time between successive tones, so a long note before a large leap and a long note after it put that leap far to the right of the curve, where the ceiling is high. This is what a melody that leaps actually does — Twinkle’s opening fifth falls between two crotchets at a walking pace, and Frère Jacques’ descending fifth is the last two notes of a phrase, at the slowest point in the bar.
That is why the constraint is easy to miss. It is invisible in any tune played at the speed tunes are usually played at, and it becomes the dominant fact about melodic writing exactly where music gets fast — which is also where composers stopped writing leaps and started writing scales, arpeggios and turns.
What the constraint does not explain
It is worth being careful about how much weight this carries, because the arithmetic is clean enough to invite overclaiming.
It does not explain slow melodies. Below about six notes a second the coherence limit is past the top of van Noorden’s measured range, and a slow melody can leap as far as it likes without any risk of splitting. Yet slow melodies are also mostly stepwise. Something else is doing the work there, and the two candidates — that a singer’s larynx moves more accurately over short distances, and that a stepwise line is easier to remember — are both plausible and neither is measured here.
It is a boundary measured on repeating three-tone sequences, not on melodies. Van Noorden’s stimuli were ABAB patterns of identical tones at a fixed rate. A melody has notes of different lengths, a shape, a metre, a harmonic context and words; every one of those is a reason the boundary might sit somewhere else in music than it does in the laboratory. Applying it to a tune is a stronger claim than the experiment supports, and this essay is making it deliberately and saying so.
And it is a limit on one line, not on music. The constraint says a fast leaping sequence becomes two streams. It does not say that is bad.
The same argument, seen from voice leading
The interesting corroboration is that this collection has already run into this boundary from the opposite direction, and got an answer that fits.
The essay on streaming and part-writing asked whether a listener follows four written voices, and found the fission limit to be the binding one there: two parts closer than about five semitones cannot be heard as two at any speed. That is the same pair of curves read for the opposite failure. Counterpoint’s problem is parts that are too close together and fuse; melody’s problem is a line whose intervals are too wide and splits.
The two problems have the same solution and it is the same number. A texture works when its parts are far enough apart to stay separate and each part’s own steps are small enough to stay together, and the window between those two conditions is exactly the shaded region in the figure above. Four-part writing and stepwise melody are the two edges of one constraint, and neither was derived from the other.
One cue is enough for some boundaries and not for others, which is the practical form of the streaming argument: the rate axis decides whether a listener is given a line at all, and once they are, the cues decide where it divides.
Which computation produced the numbers
The interval distributions are the melodies themselves, differenced. Each tune is stored as a list of pitch-and-duration pairs, the intervals are the differences between successive pitches, and a repeated note counts as a move of zero rather than being dropped — because whether a tune repeats notes is part of what its distribution says, and dropping the zeros would report Beethoven’s theme as having no unison motion when nearly a third of its moves are one.
The boundaries are van Noorden’s, interpolated. His two curves are stored as measured points against tone repetition time and read linearly between them; the table above is that function evaluated at the note rates named, with the tone repetition time taken as one over the rate. Above 160 milliseconds the coherence curve is outside the measured range and the function returns the last measured value, which is why the table says “past the top of the measured range” rather than a number.
Whose music, and when
The three tunes are European and two of them are nursery songs. That is a small and unrepresentative corpus, and it is used here only to show a shape.
The claim about the boundary is not a claim about a repertoire at all. It is a claim about the auditory system, measured on synthetic tones with listeners who were not asked about music, and it applies to any melodic tradition whatever — which is the reason for preferring it to a claim about what sounds nice. Where stepwise motion has been counted outside the European tradition, in Japanese, Chinese, Native American and West African repertoires, the same weighting toward small intervals appears; the sources for those counts are ethnomusicological surveys rather than anything this site computes.
The one place a tradition can be seen deliberately working against the constraint is instrumental. A fast passage on a plucked or keyed instrument can leap freely because it is meant to sound like several lines — a lute divisions passage, a Baroque violin partita, a bebop line at three hundred to the minute. Every one of those is a case where the boundary was crossed on purpose.
What the picture cannot show
It cannot show a melody a listener already knows. Van Noorden’s boundaries describe an unfamiliar sequence. A familiar tune is held together by memory in ways a novel one is not, and it is entirely possible that a known melody survives leaps that would split an unknown one. Nothing here measures that.
It cannot show timbre. Two notes an octave apart played by the same instrument and by two different ones do not stream alike; a change of timbre segregates a sequence almost as strongly as a change of pitch does, and that is a separate mechanism with its own essay. Every number here assumes one instrument.
And it has no harmony in it. A leap that lands on a chord tone and a leap that lands off one are the same leap in this model, and are not the same event in any music that has chords. The same blindness runs the other way round: the harmonic-rhythm essay measures how fast the chords change and has nothing to say about what the melody over them is doing, and the two rates are not independent in any repertoire.
Nor is there a metre. Every note in these figures is one column wide whatever its length, and the beat a listener infers is what decides which notes of a line are structural. A leap from a weak position to a strong one and the same leap the other way round are different melodic events, and this model cannot tell them apart.
The ladder from here
This is the first rung of the first anchor in this collection about a tune rather than about the material a tune is made of, and the rungs it obviously needs are the ones that ask what else about melodic shape is a consequence of something other than a rule.
The next asks about the most-taught melodic rule there is — that a leap is followed by a step in the opposite direction — and finds that a walk with two walls and no rule at all does most of it. After that: what survives when everything except the shape is thrown away, why a tune fits inside a twelfth, and what happens to a leading note when the argument for its tuning is horizontal rather than vertical.
What none of them will do is produce a rule for writing one. The whole point of a constraint is that it says what is impossible, and everything possible is still available.
Part 1 of 8
One essay in the series on melody. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 14.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Auditory scene analysisContourFission boundaryMelodic intervalScale degreeStreamingTemporal coherenceTessitura
- An interval is two posteriors subtracted melodic interval, scale degree
- The cue that settles it auditory scene analysis, streaming
- The exchange rate nobody has auditory scene analysis, streaming
- The shape that survives everything else contour, melodic interval
- Which end the mistuning is on melodic interval, scale degree