A reader does not read notes
Assumes: The page is read by an eye · A melody is a walk, not a set
The page is read by an eye put a reader into this ladder for the first time: the eye–hand span, the fixation, the acuity bound that cuts the rastral series exactly where practice does. It ended by naming what every quantity in the anchor still assumes:
Every quantity above counts notes, and a reader does not read notes — they read figures, and a scale fragment of eight notes is one object where eight unrelated pitches are eight.
The page cannot tell those apart. Eight quavers occupy eight quavers of horizontal space under every spacing rule this ladder has drawn, whichever line they spell.
A reader can, and this collection has the means to say by how much.
A note’s cost is the interval that reached it
A melody is a walk, not a set is where the ingredient comes from. That ladder measures the distribution of melodic interval sizes over the tunes this collection holds, and it is very far from flat: a repeated note or a step of a tone accounts for nearly three quarters of everything, a third for five per cent, a fifth for four.
So the information a note carries is a number — minus the log of the probability of the step that reached it — and it runs from under a bit for a step to seven for a leap nobody makes.
Six kinds of line, at eight notes each:
| bits a note | eight notes | |
|---|---|---|
| a scale | 1.76 | 12.3 |
| a tune | 1.89 | 13.3 |
| a chromatic run | 3.07 | 21.5 |
| an arpeggio | 5.00 | 35.0 |
| an unpatterned line | 6.09 | 42.6 |
| a wide-leaping line | 7.02 | 49.1 |
A factor of four, at the same number of notes and the same width on the page.
It is worth noticing where the tune falls. A real melody costs 1.89 bits a note, which is barely above a scale and a third of a leaping line — because a real melody is mostly steps. That is the melody ladder’s own finding arriving here from a different direction: a melody is a walk, and a walk is cheap to read for the same reason it is easy to sing.
Which is what a chunk is
The word the reading literature uses is chunk: a reader takes in a group as one object when the group is one they recognise, and the eye–hand span measured in notes is really a span measured in chunks that happens to be reported in notes because notes are what the page has.
This is that idea with a number under it. A scale is cheap because its steps are the commonest thing in the distribution; a wide-leaping line is expensive because each of its steps is individually improbable; and the gap between them is the whole content of “a reader reads figures”.
The idea is old and the number is not. Chunking has been the standard account of expert sight-reading for fifty years, and it is usually demonstrated the other way round — by showing that experts reproduce a real passage far better than a scrambled one and that the advantage vanishes when the material is scrambled. That is a measurement of the effect with no scale on it, and what an information measure adds is a scale: not that a scrambled passage is harder but that it is four times as much.
It also says why an arpeggio is not free. An arpeggio is as recognisable a figure as a scale to a musician, and on this measure it costs 5.00 bits a note — nearly three times a scale — because its steps are thirds and fourths, which the melodic distribution says are rare. The model has no term for a chord, and an arpeggio is a chord read horizontally.
That is the clearest limitation of this rung and it is worth naming before the results are used: the interval distribution is a first-order melodic one, and a reader has more than that.
The span in notes is a span in something else
The eye–hand span is about four notes for an ordinary sight-reader and up to seven for a very good one, and it is the best-measured quantity in the whole music-reading literature.
Hold the load fixed instead — the number of bits the reader is carrying — and the span becomes 4.3 notes of a scale, 4.0 of a tune, 2.5 of a chromatic run, 1.5 of an arpeggio and 1.1 of a wide-leaping line.
A sight-reader looking a bar ahead in a scale is looking a beat ahead in a leaping line, on the same page, at the same tempo, in the same notation. That is a large effect and it is one every teacher describes without a number.
What it does to the eleventh rung’s ceiling
The eleventh rung ended with a bound: the layout runs out of eye movements at 428 beats a minute, and the acuity bound cuts the nine standard rastrals exactly where practice does — the five used for parts clear it and the four used for study scores do not.
Both of those are computed at a fixed number of notes per beat. If the load rather than the count is what the reader is limited by, then the ceiling is a function of the music and the 428 is the ceiling for a line of average predictability.
For a scale it is higher by the ratio of the loads, and for a leaping line lower by a factor of nearly four. A page of scales can be read at a tempo a page of leaps cannot, at the same rastral, on the same stand.
That is not a correction to the eleventh rung’s arithmetic. It is a second axis it did not have, and the two multiply: a bound set by the eye and a bound set by the music.
The chromatic run is the interesting middle case
Of the six lines, the chromatic run is the one whose number is least obvious and most informative.
It costs 3.07 bits a note — above a scale and a tune, below an arpeggio, well below a leaping line. On the melodic distribution a semitone is uncommon: it accounts for about twelve per cent of the intervals in these tunes, against forty-three for a tone. So a chromatic run is nearly two bits a note more expensive than a diatonic one.
Every sight-reader would call that the wrong way round. A chromatic run is easy to read — it is the most regular figure there is, one step per note in one direction — and the accidentals that spell it are a nuisance rather than a difficulty.
The model is measuring the wrong regularity. A first-order distribution over interval sizes knows that a semitone is rarer than a tone and does not know that a run of seven identical semitones is a pattern. The information in a sequence is not the sum of the information in its steps unless the steps are independent, and in a scale or a chromatic run they are as far from independent as they can be.
That is a real defect and it points at the fix. A second-order model — the probability of a step given the step before — would make any run of identical intervals nearly free and would leave the leaping line where it is, which is the ordering a reader would give. The machinery for it is a table this collection does not have, and it is one more thing the corpus would supply.
So the numbers here should be read as an upper bound on the load, tightest for lines whose steps really are independent and loosest for runs. The factor of four between a tune and a leaping line is safe; the position of the chromatic run in the list is not.
And it is the first quantity in this anchor that depends on the music
That is the debt’s own phrase and it is worth taking seriously, because it is a change in what the anchor is about.
Eleven rungs measure a notation: how many staff positions a clef reaches, what a time signature claims, how much horizontal space a duration is given, how many systems a page holds. Every one of them is a property of the printing and none of them knows what is printed.
This one is not. Two pages identical in every measurable respect — same rastral, same spacing rule, same note count, same system count — differ by a factor of four in what they ask of a reader, and the difference is entirely in which notes.
So the anchor’s own quantities are an upper bound on legibility rather than a measure of it. A page that clears every layout rule can still be unreadable, and this rung says by how much and why.
That reframes several earlier results without contradicting any of them. How much music a page holds is a capacity in notes and is exactly right as a capacity in notes; what it is not is a statement about how much a reader can take off that page. What a tablature keeps compares two notations by what each records, and the comparison would look different in bits, because a tablature spells hand positions whose distribution is not the melodic one. Every rung of this anchor is a measurement of a page and this is the first of a reading.
What this would predict, if anybody measured it
The rung makes one clean prediction and it is measurable with equipment the reading literature already owns.
Eye-tracking studies of sight-readers report the eye–hand span in notes and report it as roughly constant for a given reader. If the span is really a load rather than a count, then the same reader’s span in notes should shrink on unpredictable material by the ratio of the loads — from four notes on a tune to about one on a leaping line.
That is a factor of four in a quantity that is routinely measured, on material that is trivial to construct, with a prediction stated in advance. It is also the sort of thing that may well already be in the literature under another description, because “the span is shorter in difficult music” is exactly what a reader would report.
What would refute it is a span that stays at four notes whatever the material, which would say the reader is buffering objects rather than information — and that is the other reasonable model and the one this rung is arguing against.
Which computation produced the numbers
The interval distribution is melodySteps, unchanged from the melody ladder: the sizes of every melodic interval in the tunes this collection holds, as a probability over absolute size in semitones, with a floor for sizes that do not occur.
A note’s information is minus the log to base two of the probability of the interval that reached it. A line’s load is the sum over its notes and its cost per note is the mean.
The six lines are constructed rather than found: a diatonic scale, a triadic arpeggio, a chromatic run, a wide-leaping line, an unpatterned one, and the opening of a tune from the collection’s own set.
The span at constant load takes the tune’s cost per note as the reference, because that is the material the four-note span was measured on, and divides the same budget by each other line’s cost.
Where the model stops
A first-order interval distribution is not a reader. It has no chords, no scale degrees, no key, no rhythm and no memory of the bar before. An arpeggio comes out expensive on it and is not, which is the model failing rather than the arpeggio being hard.
The distribution is three tunes. The melody ladder’s own sample is small and Western and simple, so the probabilities are indicative rather than measured over a repertoire — which is another use for the corpus three ladders on this site have recorded wanting.
Bits are not effort. Converting information to reading difficulty assumes a reader whose capacity is a channel, and nothing here measures a reader at all. What the number is good for is ratios between lines, which is how it is used above.
The floor is doing work. Intervals that never occur in the sample are given a small probability rather than zero, and the leaping line’s cost depends on what that floor is: a tenth of the value used here would add about two bits a note to it. The ordering does not move and the factor does.
And the span is not a buffer. The eye–hand span is a measured distance between where the eye is and where the hands are, and treating it as a fixed quantity of information rather than a fixed number of objects is exactly the substitution this rung is proposing. It is a proposal.
What the picture cannot show
It cannot show rhythm. Every line here is eight equal notes, and the reading literature is clear that rhythmic complexity costs a reader at least as much as pitch complexity does.
It cannot show the key signature. A line in five sharps and the same line in C are identical in interval sizes and are not identical to read, which is a fact about what a notation makes explicit rather than about melody.
Nor can it show the fingers. A pianist’s difficulty is partly a hand’s, and a wide-leaping line is hard to play as well as hard to read.
It cannot show a second reading. Sight-reading is the first time; the tenth time through, the whole page is one chunk and none of this applies.
It cannot show the performer’s memory. A sight-reader is not holding the next four notes in a vacuum; they are holding them against everything the piece has done so far, which is a memory that fades and would make a familiar idiom cheaper than its interval distribution says.
And it cannot show the layout interacting with the music. A leaping line placed across a system break is worse than either problem alone, and this ladder has the layout and now the music and has not put them together.
Whose notation, and when
The staff, the rastrals and the spacing rules are the European engraving tradition’s. The interval distribution is from simple Western tunes.
The historical observation available is a narrow one about repertoire and difficulty. The nineteenth and twentieth centuries produce a great deal of music whose melodic intervals are far from the distribution measured here — and every complaint about the unreadability of that music is a complaint about exactly the quantity on this page, made without it. What the model adds is that the notation is not at fault: the staff is doing the same job it always did, and the material has moved.
Where this ladder goes next
Twelve rungs. The stave is not a ruler; two names for one key; the time signature is a claim; a mark that is not a level; what a tablature keeps; three notations for one progression; the clef is an integer; the notations invented for the overflow; the axis that is not a time axis; the two axes multiplied; the eye that has to cross them; and now the fact that what is being crossed is not uniform.
What the ladder owes now is the vertical. Everything here is one line of one voice, and a reader of a score is taking in several at once — which multiplies the load and does not multiply it by the number of staves, because the parts are related. Two voices in parallel thirds are very nearly one line’s worth of information and two independent voices are two, and this collection has a voice-leading ladder that measures exactly how related two parts are. What comes out is why a four-part chorale is easier to read than a two-part invention, which every keyboard player knows and no layout rule contains.
Part 12 of 18
One essay in the series on notation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
ExpectationInformationMelodyMemory decayNotationSight-readingStaff
- Leaps do not fall where offbeats do information, notation, sight-reading
- The notes in between expectation, melody, memory decay
- A chord, given a key and a predecessor expectation, information
- A return has to be remembered expectation, memory decay
- Against a pulse the bell pattern is the easiest to place information, memory decay
- An expectation cannot rescue a cycle too slow to time information, memory decay