Scales and modes

A reader does not read notes

Eleven earlier essays count notes, and the page cannot tell one line of eight quavers from another. A reader can: a scale of eight is one object where eight leaps are eight. Measured against the melodic interval distribution, the same eight notes are four times as much to read — and the eye–hand span, the best-measured quantity in the reading literature, is four notes of a tune and one of a leaping line.

Assumes: The page is read by an eye · A melody is a walk, not a set

The page is read by an eye put a reader into this ladder for the first time: the eye–hand span, the fixation, the acuity bound that cuts the rastral series exactly where practice does. It ended by naming what every quantity in the anchor still assumes:

Every quantity above counts notes, and a reader does not read notes — they read figures, and a scale fragment of eight notes is one object where eight unrelated pitches are eight.

The page cannot tell those apart. Eight quavers occupy eight quavers of horizontal space under every spacing rule this ladder has drawn, whichever line they spell.

A reader can, and this collection has the means to say by how much.

The same eight notes are four times as much to read. How many bits each note of a line carries, taken as minus the log of the probability of the interval that reached it, under the distribution of melodic steps measured over the tunes used throughout. A scale costs 1.76 bits a note and a wide leaps costs 7.02 — a factor of 4.0 at the same number of notes on the page. Every quantity computed until now counts notes, and the page cannot tell these apart: eight quavers are eight quavers of horizontal space whichever line they spell.
Fig. 1 How many bits each note of a line carries, taken as minus the log of the probability of the interval that reached it, under the distribution of melodic steps measured over the tunes used throughout.

A note’s cost is the interval that reached it

A melody is a walk, not a set is where the ingredient comes from. That ladder measures the distribution of melodic interval sizes over the tunes this collection holds, and it is very far from flat: a repeated note or a step of a tone accounts for nearly three quarters of everything, a third for five per cent, a fifth for four.

So the information a note carries is a number — minus the log of the probability of the step that reached it — and it runs from under a bit for a step to seven for a leap nobody makes.

Six kinds of line, at eight notes each:

bits a note eight notes
a scale 1.76 12.3
a tune 1.89 13.3
a chromatic run 3.07 21.5
an arpeggio 5.00 35.0
an unpatterned line 6.09 42.6
a wide-leaping line 7.02 49.1

A factor of four, at the same number of notes and the same width on the page.

It is worth noticing where the tune falls. A real melody costs 1.89 bits a note, which is barely above a scale and a third of a leaping line — because a real melody is mostly steps. That is the melody ladder’s own finding arriving here from a different direction: a melody is a walk, and a walk is cheap to read for the same reason it is easy to sing.

Which is what a chunk is

The word the reading literature uses is chunk: a reader takes in a group as one object when the group is one they recognise, and the eye–hand span measured in notes is really a span measured in chunks that happens to be reported in notes because notes are what the page has.

This is that idea with a number under it. A scale is cheap because its steps are the commonest thing in the distribution; a wide-leaping line is expensive because each of its steps is individually improbable; and the gap between them is the whole content of “a reader reads figures”.

The idea is old and the number is not. Chunking has been the standard account of expert sight-reading for fifty years, and it is usually demonstrated the other way round — by showing that experts reproduce a real passage far better than a scrambled one and that the advantage vanishes when the material is scrambled. That is a measurement of the effect with no scale on it, and what an information measure adds is a scale: not that a scrambled passage is harder but that it is four times as much.

It also says why an arpeggio is not free. An arpeggio is as recognisable a figure as a scale to a musician, and on this measure it costs 5.00 bits a note — nearly three times a scale — because its steps are thirds and fourths, which the melodic distribution says are rare. The model has no term for a chord, and an arpeggio is a chord read horizontally.

That is the clearest limitation of this rung and it is worth naming before the results are used: the interval distribution is a first-order melodic one, and a reader has more than that.

The span in notes is a span in something else

Four notes of a tune, and one of a leaping line. The eye–hand span is measured in notes, and it is 4 of them for an ordinary sight-reader. Holding the load fixed instead — the number of bits the reader is carrying — gives a span that depends on what the music is. The same load is 4.3 notes of a scale and 1.1 of a wide leaps. So a sight-reader looking a bar ahead in a scale is looking a beat ahead in a wide-leaping line, on the same page at the same tempo.
Fig. 2 The eye–hand span held at a constant load rather than at a constant number of notes.

The eye–hand span is about four notes for an ordinary sight-reader and up to seven for a very good one, and it is the best-measured quantity in the whole music-reading literature.

Hold the load fixed instead — the number of bits the reader is carrying — and the span becomes 4.3 notes of a scale, 4.0 of a tune, 2.5 of a chromatic run, 1.5 of an arpeggio and 1.1 of a wide-leaping line.

A sight-reader looking a bar ahead in a scale is looking a beat ahead in a leaping line, on the same page, at the same tempo, in the same notation. That is a large effect and it is one every teacher describes without a number.

What it does to the eleventh rung’s ceiling

The eleventh rung ended with a bound: the layout runs out of eye movements at 428 beats a minute, and the acuity bound cuts the nine standard rastrals exactly where practice does — the five used for parts clear it and the four used for study scores do not.

Both of those are computed at a fixed number of notes per beat. If the load rather than the count is what the reader is limited by, then the ceiling is a function of the music and the 428 is the ceiling for a line of average predictability.

For a scale it is higher by the ratio of the loads, and for a leaping line lower by a factor of nearly four. A page of scales can be read at a tempo a page of leaps cannot, at the same rastral, on the same stand.

That is not a correction to the eleventh rung’s arithmetic. It is a second axis it did not have, and the two multiply: a bound set by the eye and a bound set by the music.

The band a rastral size has to sit in, and where it closes. The upper line is the largest notehead the eye can still saccade past fast enough; the lower is the smallest it can identify. The band between them is where a rastral size has to be, and it narrows with tempo because the saccade bound falls and the acuity bound does not move at all. At 40 beats a minute the band is a factor of 11.5 wide; at 300 it is 1.54. It does not close inside any tempo a player meets, which is the answer to whether a page can be laid out too densely to sight-read: not by the eye's speed. The dots on the lower line are the standard rastrals, and 5 of the nine sit inside the band at every tempo here.
Fig. 3 The band a rastral size has to sit in, computed earlier at a fixed number of notes a beat. Everything on it moves with the predictability of the line.

The chromatic run is the interesting middle case

Of the six lines, the chromatic run is the one whose number is least obvious and most informative.

It costs 3.07 bits a note — above a scale and a tune, below an arpeggio, well below a leaping line. On the melodic distribution a semitone is uncommon: it accounts for about twelve per cent of the intervals in these tunes, against forty-three for a tone. So a chromatic run is nearly two bits a note more expensive than a diatonic one.

Every sight-reader would call that the wrong way round. A chromatic run is easy to read — it is the most regular figure there is, one step per note in one direction — and the accidentals that spell it are a nuisance rather than a difficulty.

The model is measuring the wrong regularity. A first-order distribution over interval sizes knows that a semitone is rarer than a tone and does not know that a run of seven identical semitones is a pattern. The information in a sequence is not the sum of the information in its steps unless the steps are independent, and in a scale or a chromatic run they are as far from independent as they can be.

That is a real defect and it points at the fix. A second-order model — the probability of a step given the step before — would make any run of identical intervals nearly free and would leave the leaping line where it is, which is the ordering a reader would give. The machinery for it is a table this collection does not have, and it is one more thing the corpus would supply.

So the numbers here should be read as an upper bound on the load, tightest for lines whose steps really are independent and loosest for runs. The factor of four between a tune and a leaping line is safe; the position of the chromatic run in the list is not.

And it is the first quantity in this anchor that depends on the music

That is the debt’s own phrase and it is worth taking seriously, because it is a change in what the anchor is about.

Eleven rungs measure a notation: how many staff positions a clef reaches, what a time signature claims, how much horizontal space a duration is given, how many systems a page holds. Every one of them is a property of the printing and none of them knows what is printed.

This one is not. Two pages identical in every measurable respect — same rastral, same spacing rule, same note count, same system count — differ by a factor of four in what they ask of a reader, and the difference is entirely in which notes.

So the anchor’s own quantities are an upper bound on legibility rather than a measure of it. A page that clears every layout rule can still be unreadable, and this rung says by how much and why.

That reframes several earlier results without contradicting any of them. How much music a page holds is a capacity in notes and is exactly right as a capacity in notes; what it is not is a statement about how much a reader can take off that page. What a tablature keeps compares two notations by what each records, and the comparison would look different in bits, because a tablature spells hand positions whose distribution is not the melodic one. Every rung of this anchor is a measurement of a page and this is the first of a reading.

How much music each spacing rule fits on a page. The two axes multiplied. Vertically, a system is as tall as the staves the range needs; horizontally, a system holds as many notes as fit once the shortest is wide enough to read. proportional: 3 staves to the system, 4 systems and 27 notes to a system, 108 notes to the page — 32 seconds at 100 beats a minute, so a page turn every 32 seconds; Ross, 1970: 3 staves to the system, 4 systems and 38 notes to a system, 152 notes to the page — 46 seconds at 100 beats a minute, so a page turn every 46 seconds; Gould, 2011: 3 staves to the system, 4 systems and 40 notes to a system, 160 notes to the page — 48 seconds at 100 beats a minute, so a page turn every 48 seconds; one column a note: 3 staves to the system, 4 systems and 50 notes to a system, 200 notes to the page — 60 seconds at 100 beats a minute, so a page turn every 60 seconds. one column a note holds 1.85 times what proportional does, which is a difference of 0.85 page turns a minute — and a page turn is a thing a player with two hands occupied cannot do.
Fig. 4 How much music each spacing rule fits on a page, computed earlier. Every one of those numbers is in notes, and this essay is the argument that notes are the wrong unit.
The eye-hand span does not fit inside one fixation. A sight-reader's eye sits a roughly fixed number of notes ahead of the sounding one, and a fixation takes in a roughly fixed number of millimetres, because that is what the fovea is. Those are two facts in different units and the spacing rule converts between them. At this layout a note occupies 5.4 millimetres, so a span of 4 notes is 22 millimetres against a useful field of 19 — the span does not fit, and reading is therefore a sequence of fixations rather than one. Each fixation covers 3.6 notes, and at 100 beats a minute with 2 to the beat that is 0.9 saccades a second against a ceiling of about four. There is a great deal of headroom: this layout runs out of eye movements at 428 beats a minute, which nobody plays.
Fig. 5 The earlier reading distance: how far ahead of the sounding note the eye sits, in millimetres on the page. It is computed from a span in notes, which is the assumption this essay replaces.

What this would predict, if anybody measured it

The rung makes one clean prediction and it is measurable with equipment the reading literature already owns.

Eye-tracking studies of sight-readers report the eye–hand span in notes and report it as roughly constant for a given reader. If the span is really a load rather than a count, then the same reader’s span in notes should shrink on unpredictable material by the ratio of the loads — from four notes on a tune to about one on a leaping line.

That is a factor of four in a quantity that is routinely measured, on material that is trivial to construct, with a prediction stated in advance. It is also the sort of thing that may well already be in the literature under another description, because “the span is shorter in difficult music” is exactly what a reader would report.

What would refute it is a span that stays at four notes whatever the material, which would say the reader is buffering objects rather than information — and that is the other reasonable model and the one this rung is arguing against.

The staff holds 11 positions and nothing fits in it. Each clef's eleven staff positions — five lines, four spaces and the space either side — as a bar on an axis that counts letters, with eight ranges laid underneath. The clefs step through the axis in thirds and cover fifteen positions of offset between them. Every range drawn is wider than eleven positions: the four voices span 13, 12, 13, 13 and the four instruments 24, 24, 25, 23, so the best clef for each still leaves 1 to 7 positions off the staff. A clef is a choice of which end sticks out.
Fig. 6 What the staff holds, computed earlier. Nothing on this page changes it: the argument is that two pages identical in every respect measured so far can differ by a factor of four in what they ask.
Every standard rastral size against the two bounds a reader imposes. Print the notes larger and the eye-hand span stops fitting inside one fixation, so the reader has to saccade ahead faster than the eye can move. Print them smaller and a notehead stops subtending enough angle to be identified. Both bounds come from the reader and neither from the music. At 100 beats a minute with 2 notes to the beat, the acuity bound sits at 1.63 millimetres and the saccade bound at 7.5 — so the saccade rate is nowhere near binding and acuity is doing all the work, which is the opposite of what the eye-hand span suggests. 5 of the 9 standard rastrals clear the acuity bound: rastral 4 and larger. Those are exactly the sizes used for parts, and the ones below are used for study scores — which are read at a desk rather than played from at a stand, and a shorter viewing distance moves the bound with them.
Fig. 7 The layout constraint built on earlier, for scale. This essay’s claim is that a second constraint sits beside it, is invisible on the page, and moves by more.

Which computation produced the numbers

The interval distribution is melodySteps, unchanged from the melody ladder: the sizes of every melodic interval in the tunes this collection holds, as a probability over absolute size in semitones, with a floor for sizes that do not occur.

A note’s information is minus the log to base two of the probability of the interval that reached it. A line’s load is the sum over its notes and its cost per note is the mean.

The six lines are constructed rather than found: a diatonic scale, a triadic arpeggio, a chromatic run, a wide-leaping line, an unpatterned one, and the opening of a tune from the collection’s own set.

The span at constant load takes the tune’s cost per note as the reference, because that is the material the four-note span was measured on, and divides the same budget by each other line’s cost.

Where the model stops

A first-order interval distribution is not a reader. It has no chords, no scale degrees, no key, no rhythm and no memory of the bar before. An arpeggio comes out expensive on it and is not, which is the model failing rather than the arpeggio being hard.

The distribution is three tunes. The melody ladder’s own sample is small and Western and simple, so the probabilities are indicative rather than measured over a repertoire — which is another use for the corpus three ladders on this site have recorded wanting.

Bits are not effort. Converting information to reading difficulty assumes a reader whose capacity is a channel, and nothing here measures a reader at all. What the number is good for is ratios between lines, which is how it is used above.

The floor is doing work. Intervals that never occur in the sample are given a small probability rather than zero, and the leaping line’s cost depends on what that floor is: a tenth of the value used here would add about two bits a note to it. The ordering does not move and the factor does.

And the span is not a buffer. The eye–hand span is a measured distance between where the eye is and where the hands are, and treating it as a fixed quantity of information rather than a fixed number of objects is exactly the substitution this rung is proposing. It is a proposal.

What the picture cannot show

It cannot show rhythm. Every line here is eight equal notes, and the reading literature is clear that rhythmic complexity costs a reader at least as much as pitch complexity does.

It cannot show the key signature. A line in five sharps and the same line in C are identical in interval sizes and are not identical to read, which is a fact about what a notation makes explicit rather than about melody.

Nor can it show the fingers. A pianist’s difficulty is partly a hand’s, and a wide-leaping line is hard to play as well as hard to read.

It cannot show a second reading. Sight-reading is the first time; the tenth time through, the whole page is one chunk and none of this applies.

It cannot show the performer’s memory. A sight-reader is not holding the next four notes in a vacuum; they are holding them against everything the piece has done so far, which is a memory that fades and would make a familiar idiom cheaper than its interval distribution says.

And it cannot show the layout interacting with the music. A leaping line placed across a system break is worse than either problem alone, and this ladder has the layout and now the music and has not put them together.

Whose notation, and when

The staff, the rastrals and the spacing rules are the European engraving tradition’s. The interval distribution is from simple Western tunes.

The historical observation available is a narrow one about repertoire and difficulty. The nineteenth and twentieth centuries produce a great deal of music whose melodic intervals are far from the distribution measured here — and every complaint about the unreadability of that music is a complaint about exactly the quantity on this page, made without it. What the model adds is that the notation is not at fault: the staff is doing the same job it always did, and the material has moved.

Where this ladder goes next

Twelve rungs. The stave is not a ruler; two names for one key; the time signature is a claim; a mark that is not a level; what a tablature keeps; three notations for one progression; the clef is an integer; the notations invented for the overflow; the axis that is not a time axis; the two axes multiplied; the eye that has to cross them; and now the fact that what is being crossed is not uniform.

What the ladder owes now is the vertical. Everything here is one line of one voice, and a reader of a score is taking in several at once — which multiplies the load and does not multiply it by the number of staves, because the parts are related. Two voices in parallel thirds are very nearly one line’s worth of information and two independent voices are two, and this collection has a voice-leading ladder that measures exactly how related two parts are. What comes out is why a four-part chorale is easier to read than a two-part invention, which every keyboard player knows and no layout rule contains.

Part 12 of 18

One essay in the series on notation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

ExpectationInformationMelodyMemory decayNotationSight-readingStaff