Form and structure

The twelfth, and where it comes from

Melodies occupy about an octave and a fifth, and an earlier essay set out to explain that by the singer's register break and found that it does not: the chest mechanism alone spans two octaves and a semitone. The answer is in a parameter the essay on leaps fitted and then put down. A walk with no walls whose central tendency reproduces the post-skip reversal rate has a range that grows logarithmically — six semitones at eight notes, twelve at thirty, eighteen at a hundred and twenty — so across every length a tune plausibly has, the span is between an octave and a fifteenth.

Assumes: The leap that pays itself back · A melody is a walk, not a set

Melodies are narrow. The usual figure is that a tune occupies about an octave and a fifth — a twelfth — and it is repeated in enough places, about enough repertoires, that it is worth asking where it comes from.

This ladder set out to answer that once already and did not. Its fourth rung was to explain the twelfth by the singer’s register break, on the ground that a tune staying inside one vocal mechanism stays inside one colour. The machinery refused: the chest mechanism alone runs from 82 to 349 hertz, which is two octaves and a semitone, and no arrangement of the numbers produces a twelfth. The rung became a different and better essay about the clarino register, and the original question was recorded as owed.

This essay pays it, and the answer was already in the collection — sitting in a parameter that the leap rung fitted, used once, and put down.

First, the obvious model, and it fails in the informative direction

If a melody is a walk on scale degrees with the step distribution these three tunes have, then how wide it gets is a property of the walk and the number of notes. That is a completely determinate quantity, and it can be computed exactly rather than simulated.

A walk of a tune's length is wider than a tune. The exact mean range of a random walk with the step distribution measured off the three tunes carried here — computed by enumeration over the states a range depends on, not sampled. At eight notes it is 5.04 semitones; at thirty, 18.12. The three tunes have thirty to forty-two notes and spans of 7, 9, 14 semitones, so every one of them is narrower than a walk with its own step distribution. The range is imposed; it is not what the steps produce.
Fig. 1 The exact mean range of a walk with the measured step distribution, computed by enumerating the states a range depends on rather than by sampling. At eight notes it is 8.07 semitones; at thirty, 18.12; at forty, 21.36.

Now put the tunes beside it.

Three tunes’ spans against their lengths are seven semitones over thirty notes, nine over forty-two and fourteen over thirty-two — a spread among the three that is larger than the discrepancy with any model, which is the honest limit of a corpus of three.

So the range is imposed and not produced. A free walk with the right marginal statistics wanders about twice as far as a melody does, and it keeps going: at a hundred and twenty notes it covers thirty-nine semitones, which is more than three octaves.

That is a negative result and it is the useful kind, because it says exactly what has to be added. Something pulls a melody back, the something is not a wall at the edge of a vocal range — a wall three octaves out would never be reached — and it has to act throughout.

It is worth naming the three candidates that were available before any of this, because the elimination is most of the argument.

The instrument. A voice, a violin or a flute has a compass — and on a wind instrument the register where the series becomes a scale is narrower than the compass is, and a tune that fits inside it is confined by it. This fails on arithmetic: every one of those compasses is over two octaves, and a walk that never reaches the wall is not affected by it.

The register break. A tune staying in one vocal mechanism stays in one colour, so the mechanism’s width would set the tune’s. This is what the fourth rung tried and it fails on the same arithmetic, one octave down.

The walk. A tune of n notes with small steps only gets so far. This is what the figure above tests, and it fails in the opposite direction: the walk gets further than a tune does.

Two of the three are too wide and one is too wide as well. What is left is that something acts continuously rather than at a boundary, and the leap rung already knew what it was.

The parameter that was already fitted

The leap rung ran into the same shortfall from the other direction. A bounded walk reproduces 61.3 per cent post-skip reversal against the 70-odd the corpus studies report, and closing the gap needs a central tendency: a probability of stepping back toward the middle that rises with distance from it.

How much of the rule a walk with no rule reproduces. Post-skip reversal in 20,000-note random walks with no melodic knowledge of any kind. An unbounded walk reverses after 50.0 per cent of leaps, which is the chance rate and is the check that the measurement is right. Confining it to 12 semitones raises that to 61.3 per cent. Reaching the 70 per cent that corpus studies report needs a central tendency of 0.95 — an almost deterministic pull back toward the middle at the edges of the range. A wall is not enough; there has to be a spring.
Fig. 2 That result, as it was drawn earlier. Reversal against how strongly the walk is pulled toward the middle of its range. Two hard walls and nothing else give 61.3 per cent; reaching the reported 70 needs a pull of 0.95 out of 1.

The trouble with that figure for the present question is that its walk is told how wide to be. It has walls at nought and twelve, and asking it how wide a melody is would return twelve.

So the model has to be rebuilt without walls, and this is the whole of the change: the probability of stepping back toward a fixed centre rises with how far out the walk already is, over a stated scale, and there is nothing to stop it going further. The range becomes an output.

One parameter, κ. Fit it to the reversal rate — a statistic that says nothing whatever about how wide anything is — and then read the range off.

Where a twelfth comes from. The range of a walk with no walls, against how many notes it runs for, at three settings of the one parameter it has. The parameter is fitted to the post-skip reversal rate and to nothing else; the range is then read off. With no central tendency at all the walk passes two octaves by 60 notes and keeps going. At the setting that reproduces 70 per cent reversal — κ = 0.78 — the range is 11.9 semitones at thirty notes and 17.9 at a hundred and twenty. It grows logarithmically, so over the whole plausible length of a tune it sits between an octave and a fifteenth, and a twelfth is the middle of that. The three tunes carried here are marked and all three fall below the curve.
Fig. 3 The range of the wall-free walk against how many notes it runs for, at three settings of κ, each fitted to a reversal rate and to nothing else. With no central tendency at all the walk passes two octaves before thirty notes and three before a hundred and twenty. At κ = 0.78, which reproduces seventy per cent reversal, the range is 11.9 semitones at thirty notes and 17.8 at a hundred and twenty. The three tunes are marked and all three sit below the curve.

At the fitted κ the walk’s range passes a twelfth at about a hundred and sixty-five notes and reaches 20.4 semitones at two hundred and forty.

Why the answer is a twelfth rather than something else

The number itself matters less than the shape of the curve, and the shape is the result.

The range grows logarithmically. Doubling the length of the tune adds about two and a half semitones, so from thirty notes to two hundred and forty — a factor of eight, which is more variation in length than tunes actually show — the range moves only from 11.9 to 20.4.

That is the answer to why the figure quoted is always about the same. A quantity that grows logarithmically has no characteristic value in the ordinary sense and every characteristic value in the practical one: over the whole range of lengths a melody plausibly has, it sits between an octave and a fifteenth, and a twelfth is the middle of that. Nothing has to be tuned to produce it. A tune long enough to be a tune, walked with the reversal statistic melodies have, is about a twelfth wide.

The sensitivity check runs the same way. The published reversal figures range from about sixty-five to about seventy-five per cent; fitting κ to each end gives spans at a hundred and twenty notes of 19.7 and 16.1 semitones. So over the joint uncertainty — which figure from the literature, and how long the tune is — the prediction is eleven to twenty-three semitones, centred near a twelfth.

How much of it is simply how wide the range is. Post-skip reversal in a walk with two walls and no rule, against how far apart the walls are. It falls from 88.1 per cent at 5 semitones to 53.2 at 30, because a wide range puts most leaps nowhere near a wall. The 70 per cent that corpus studies report lands at about 9 semitones — which is narrower than any melody's total compass and about the width of a phrase. That is a prediction, and a sharp one: the bound that produces the effect is local, not the tune's whole range.
Fig. 4 The other earlier reading, and the consistency check between the two models. A walk with hard walls reproduces the reported reversal rate at a width of about nine semitones — and a nine-semitone box entered by an eight-note walk is visited to a mean depth of 5.4 semitones, against the wall-free walk’s 6.15 at the same length. The two models agree about the local behaviour and differ about what happens over a whole tune, which is exactly where one of them has no answer.

What the tunes still say against it

The model overshoots the three tunes and that has to be reported rather than smoothed.

At thirty notes it predicts 11.9 semitones. Ode to Joy has thirty notes and spans seven. Frère Jacques has thirty-two and spans fourteen, which is above the prediction; Twinkle has forty-two and spans nine, which is well below. Mean observed is ten against a predicted twelve to thirteen.

Three tunes is not a corpus and the spread among them — seven to fourteen — is larger than the discrepancy with the model, so the honest statement is that these tunes are consistent with the prediction and cannot test it. What they can do is show the structure the model does not have.

Where the reversals actually are. Post-skip reversal split by where the leap ended: heading for the near edge of the range, or back toward the middle of it. In both walks the leaps that end near an edge reverse far above chance and the leaps that end in the middle reverse BELOW it. A melodic rule predicts neither asymmetry, so this is the measurement that separates the two accounts — and it is a prediction the null model makes rather than a fit it achieves.
Fig. 5 Where the reversals actually are, which is what makes the null model a real explanation rather than a coincidence of averages. Split by where the leap ended: leaps that finish near an edge of the range reverse far above chance and leaps that finish in the middle reverse below it. A rule that said “large leaps are followed by a step in the opposite direction” would apply everywhere; a wall applies only near the wall. That is a testable difference and it is the one the corpus figure of 0.7 does not distinguish, because an average over all positions hides it.

That decomposition holds on all three, and it is worth writing down as an identity:

A tune’s span is its phrase’s span plus the drift of its phrase centres. Ode to Joy: 4.5 plus 2.5 gives 7, and the observed span is 7. Twinkle: 6.3 plus 2.0 gives 8.3 against an observed 9. Frère Jacques: 4.9 plus 9.5 gives 14.4 against an observed 14.

The measured phrase-centre step is 1.87 semitones across all fifteen phrase transitions in the collection. So the same logarithmic argument applies one level up: the centres do a walk of their own, with small steps, and the total span is a phrase’s width plus a slow drift.

And the phrase’s own width is set by its length in time rather than by anything melodic: a phrase is two to eight seconds and notes come at one or two a second, so a phrase is eight to sixteen notes, which the exact walk puts at eight to twelve semitones of free range.

The bigger the leap, the more a wall reverses it. Post-skip reversal by the size of the leap, in three walks with no melodic rule in any of them. The unbounded walk is flat at chance, as it must be. Both bounded walks rise with leap size, because a larger leap is likelier to have ended near an edge — which is a prediction a melodic rule does not make: a rule that says reverse after a leap has no reason to care how big it was.
Fig. 6 The second prediction the wall makes and a rule does not. Post-skip reversal by the size of the leap: the unbounded walk is flat at chance, as it must be, and both bounded walks rise with leap size — a larger leap is likelier to have ended near an edge, so it is likelier to be turned round. A melodic rule stated as “a leap is followed by a step back” predicts a constant; the wall predicts a slope. The corpus figure everybody quotes is one number and cannot tell them apart, and the slope is what a corpus study would have to report.

That points at where the second term of the decomposition comes from, and it is not melodic at all. Phrase centres drift when the music is going somewhere — up through a sequence, out to a new key, toward a climax — and they do not when it is repeating, which is to say that the key plan is the form and the tune’s width is partly a reading of it. So the width of a tune is partly a fact about its form, and the widest of the three tunes here is the one whose eight phrases are a through-composed climb rather than four statements of two.

The two arrivals that were not arranged

Two other things in this collection land on the same number and neither was set to.

A four-part realisation of an ordinary progression puts each voice inside a range of about a twelfth without anybody choosing that, because the parts have to stay apart and stay inside a compass — which is the same arrival from the harmony side. And a contour of six notes over seven degrees collapses to far fewer distinct shapes than the degrees allow, which is the same arrival from the memory side.

That is not a derivation, because a choral range is partly a convention about what is comfortable and partly a physical fact. It is a coincidence worth flagging: the width a tune has and the width a part is written within are the same width, and the second was measured for a completely different purpose.

The other arrival is the clef, and it is the debt the notation ladder left one rung ago. A staff of five lines with one ledger line above and below holds fifteen positions, which is two octaves and a note; the reason there are several clefs is to slide that window so that a given voice fits inside it. A notation designed to keep a part inside the lines is evidence about how wide a part is, and it points at the same order of magnitude.

And the two laryngeal mechanisms overlap by about eight semitones, which is the same quantity arriving from the instrument: a singer’s comfortable range within one mechanism is roughly a twelfth, and the seam sits inside it.

The claim, stated so it can be refused

Put together, the account is three sentences and each is separately checkable.

A melody is locally a walk with small steps, which is what the first rung of this ladder measured and which the streaming boundary explains: at the note rate a tune uses, a large interval stops being one line.

It is pulled toward a centre with a strength that reproduces post-skip reversal, which is what the second rung measured and which was left as a fitted parameter with nothing else asked of it.

Its range is therefore logarithmic in its length, and lands between an octave and a fifteenth for every tune anybody writes. That is this rung, and it is a prediction rather than a fit, because the parameter was set by a statistic that contains no information about width.

The refutation is available and it is cheap: measure the ambitus and the post-skip reversal rate of the same corpus, tune by tune, and check that the two are related the way the curve says. If a repertoire with a high reversal rate is not narrower than one with a low rate, the account is wrong. Nothing in the literature reports the pair together, which is why this is still a prediction.

Run on the three tunes here, it comes out backwards

Three tunes is not a corpus and it is the corpus this collection has, so the test can at least be performed rather than only described.

tune notes span leaps reversals rate
Ode to Joy 30 narrowest 8 1 0.13
Twinkle 42 middle 5 2 0.40
Frère Jacques 32 widest 18 10 0.56

The ordering is exactly inverted. The model says a strongly reverting walk is narrow, so the narrowest tune should reverse most; here the narrowest reverses least and the widest reverses most, and the three are in perfect reverse order.

Two things stop that being a refutation and neither makes it comfortable. The leap counts are eight, five and eighteen, which is far too small to rank three rates. And the ordering is not robust to the definition: raising the leap threshold from three semitones to five leaves only Frère Jacques with any leaps at all.

What is robust is the level, and it is the more awkward finding. None of the three reverses at anything near the seventy per cent the whole account is fitted to. Pooled over all three tunes the rate is 13 of 31, which is 0.42, and it stays between 0.33 and 0.44 at every leap threshold from two semitones to four. If the true rate were 0.70, seeing 13 of 31 has a probability of about one in a thousand, and seeing Ode to Joy’s 1 of 8 has about the same.

So the parameter this rung reads its answer off is a published figure that the collection’s own melodies do not reproduce, and refitting κ to 0.42 would give a much weaker central tendency and a much wider walk. The essay’s caveat that “the prediction is conditional” is therefore doing more work than it looked like: the condition is not merely unverified here, it is contradicted by the only three tunes available to check it against.

The likeliest explanation is the repertoire. The published rate comes from large corpora of folk and art song; three well-known tunes chosen for other purposes — two of them built almost entirely of steps — are not a sample of that. But the direction of the discrepancy is worth carrying: these three tunes are narrower than a walk fitted to their own reversal rate would be, which is the same shortfall the free walk showed, one level down.

Which computation produced the numbers

The step distribution is measured off the three tunes: the fraction of intervals of each absolute size, with 30.7 per cent repeated notes, 42.6 per cent whole tones and a thin tail to a fifth.

The exact span is a dynamic programme over the pair (offset above the running minimum, span so far), which is everything the range of a walk depends on. It is not a simulation and its output is a distribution rather than a mean; the figure prints the mean and the median. Probability escaping the state bound is reported, and it is under two parts in ten thousand at forty notes.

The wall-free walk is simulated, because its state is unbounded and the exact method does not apply. Each figure averages several thousand runs with deterministic seeds, so the same figure is drawn every time; the standard error on a mean span over three thousand runs is under a twentieth of a semitone.

κ is found by bisection on a forty-thousand-note walk, twenty-two iterations, against a target reversal rate. Post-skip reversal is measured by the same function the previous rung used, with the same definitions: a leap is three semitones or more, and repeated notes leave the denominator.

The phrase boundaries used in the decomposition are the notated ones, entered from the tunes’ own structure. The centre of a phrase is the midpoint of its extremes, not its mean pitch; using the mean instead changes the centre ranges by under a semitone and changes no conclusion.

What the picture cannot show

The model has one parameter and it was fitted to a quoted number. Seventy per cent post-skip reversal is a figure from the corpus literature, and a section above measures it on the three tunes this collection carries and gets 0.42. The prediction is therefore conditional: if melodies reverse after leaps at the rate that is reported, then a wall-free walk reproducing that rate is about a twelfth wide over a tune’s length. Both halves are the model’s.

It is a walk and melodies are not walks. There is no tonic in it, no phrase structure, no cadence, no harmony and no metre. Its resemblance to music was one number deep in the previous rung and is two numbers deep now, which is not the same as being right.

The logarithmic growth is a property of mean-reverting walks in general and not something discovered about music. What is specific to music is the value of κ, and κ came from one statistic.

The decomposition is arithmetic, not a model. Span equals phrase span plus centre drift is close to a tautology for tunes whose phrases do not overlap in register, and it is wrong for tunes whose phrases do. It is offered as a description of where the width sits, not as a second derivation.

And the twelfth itself is quoted. No corpus of ambitus measurements is computed here; the figure is the one that appears in the literature and in pedagogy, and this collection’s own three tunes span seven, nine and fourteen. The essay explains a number it did not measure, using a parameter it did not measure, and the honest description of what it has shown is that the two quoted numbers are consistent with each other under a model with nothing else in it.

The ladder from here

This closes the three debts the melody ladder recorded against itself — duration, memory and the width — and leaves the ladder at eight rungs with a different debt in place of them: every quantity in the last three essays is conditional on a corpus statistic that this collection quotes and cannot check, and the measurement that would settle all three is the same one the leap rung named, which is post-skip reversal conditioned on where the leap ended, counted over a real corpus with each melody’s own local range as the boundary.

Part 8 of 8

One essay in the series on melody. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Melodic intervalNull modelPhrasePost-skip reversalRandom walkRegisterRegression to the meanTessitura