Scales and modes

A note lasts until the next one starts

Pricing where a note is against which note it is left duration as the term it had not, with a prediction that it would be small. Measured on the three tunes these readings are built on, it is exactly zero — and it is zero by construction, because those tunes are stored as pitches and lengths with no rests in them, so every duration is its own inter-onset interval. The prediction cannot be tested on the corpus that produced it. Priced directly, a rest costs 0.67 bits a note where a tenth of the notes have one, which is not well under half a bit.

Assumes: Where the note is costs more than which note it is · A reader does not read notes

The fourteenth rung prices a page along two axes. Which note it is costs a reader 1.89 bits; where the note is costs 2.23. It ends by naming the term it does not have — duration — with a prediction attached: the term will be small, well under half a bit, because a note lasts until the next one starts and the information is only in the exceptions.

The measurement is the same measurement all fourteen rungs use, on the same three tunes. It takes about a line to write, and its answer is not a number that could be compared with half a bit.

The term that was owed, and the corpus cannot hold it. For each of the three tunes everything here is measured on, how many of its notes have a duration that differs from the gap to the next onset. The answer is none, in 101 notes: these tunes are stored as a list of pitches and lengths with no rests in them, so a note's duration IS its inter-onset interval and conditioning one on the other leaves exactly zero bits. That is a fact about the representation rather than about music. The prediction was that the term would be small, and it could not have been known that the corpus would make it identically zero — which means the prediction cannot be tested here and the exceptions have to be priced directly.
Fig. 1 For each of the three tunes measured throughout, how many of its notes have a duration that differs from the gap to the next onset.

None of them, in a hundred and one notes. The duration term is exactly zero.

Zero because of the representation, not because of the music

The three tunes are stored as a list of pairs — a scale degree and a length — and nothing else. There is no rest in the data structure, because there is no way to write one: a note’s length is the time until the next note begins, by construction of the format.

So conditioning the duration on the onsets leaves nothing to condition, and the measurement is not a measurement at all. It is a statement about a data format that thirteen previous rungs have been reading numbers out of.

That is worth naming precisely, because it is a defect of a particular kind. Nothing here is wrong: every number the ladder has published is correct for what it computed. What is wrong is that the corpus cannot express the phenomenon the fourteenth rung asked about, and no gate on this site could have said so, because the format is doing exactly what it was built to do.

The fourteenth rung ended by calling its prediction “the last term available without a corpus or a reader”. It was neither: it needed a corpus, and the corpus it named was the wrong one.

Which leaves the exceptions to be priced directly

The prediction’s reasoning is still right, and it is the whole of what can be salvaged. A reader who knows where every note starts knows every duration except where a note stops early — so the term is a decision about the exceptions, and the exceptions are rests.

Written out, the decision is one bit’s worth of question per note and then, where the answer is yes, a length:

bits = H(r) + r · log₂(k)

with r the share of notes followed by a rest and k the number of rest lengths a reader has to tell apart.

What a rest costs a reader, against how often there is one. The duration term as a function of the share of notes that stop before the next one starts. A reader who knows the onsets has one binary decision per note — is there a rest here? — worth its own binary entropy, plus, where there is, which of 4 lengths it has. At a tenth of the notes the total is 0.67 bits, of which 0.47 is the decision and 0.20 is the length. The prediction was well under half a bit. That holds only for a repertoire whose notes stop early less than about one time in twenty, which is sparser than most music and a great deal sparser than any of the three tunes measured here — none of which has a rest at all.
Fig. 2 The duration term against how often a note stops early, with four distinguishable rest lengths. The dashed line is the earlier prediction.

At a tenth of the notes the total is 0.67 bits, of which 0.47 is the decision and 0.20 is which length. The fourteenth rung’s “well under half a bit” holds only below about one note in twenty.

The exponent matters more than the constant here, and it is worth reading the curve rather than a point on it. The term rises steeply from nothing and flattens: it is 0.18 bits at one note in fifty, 0.39 at one in twenty, 0.67 at one in ten and 1.12 at one in five. So the first few per cent of rests are the expensive ones, and a repertoire that has decided to have rests at all has already paid most of what it will pay.

One in twenty is sparse. It is sparser than a Bach chorale, very much sparser than an orchestral wind part, and it is infinitely sparser than the three tunes, which have no rests whatever. So the prediction is wrong for most music and right for the corpus that produced it, which is the same defect the first half of this rung found, arriving from the other side. A prediction made from a corpus that cannot contain the phenomenon will be a prediction that the phenomenon is small, every time, and it will be right about that corpus and about nothing else.

Why a decision costs more than the thing it decides

The shape of the curve is worth reading rather than the number, because it says something a reader would not guess.

The term is dominated by the binary entropy, not by the length. At a tenth, the decision is 0.47 bits and the length is 0.20 — the question “is there a rest?” costs more than twice the answer “how long is it?”, and it costs most at a rest density of one half, where a reader has no idea which way it will go.

That is the general shape of any decision priced this way and it has a consequence for notation. A repertoire in which rests are rare pays little; a repertoire in which they are rare-but-not-negligible pays most of the maximum. The cost is not proportional to how much silence there is; it is a function of how unpredictable it is.

Three terms, and the smallest of them is the one that was missing. What a reader of one line of music is charged for, per note, in bits: which note it is, where it is in the bar, and how long it lasts at a rest density of 10 per cent. The first two are the earlier numbers, 1.89 and 2.23; the third is 0.67, which is 14 per cent of the total of 4.79. So a single line's reading load is fully accounted for by two axes and a correction, which is what was hoped — and the correction is a seventh of the load rather than the tenth it expected, because a rest is a decision and a decision costs a whole bit at the density where it is most uncertain.
Fig. 3 The three terms on one scale: which note it is, where the note is, and how long it lasts at a tenth of the notes carrying a rest.

Set beside the other two, the term is 0.67 against 1.89 and 2.23 — the smallest of the three, and 14 per cent of a total of 4.79 bits a note. So the fourteenth rung’s hope is met in its substance even though its number is not: a single line’s reading load is two axes and a correction, and the correction is a seventh of the load rather than a tenth of it.

What the corpus could have held and does not

It is worth being specific about what a corpus with rests in it would need, because the fix is small and the reason it was never made is instructive.

A note on the downbeat costs 1.89 bits and one on the offbeat 5.70. What each position in a bar of 4/4 asks of a reader, by two routes. The solid bar counts where the notes of this collection's own three tunes actually fall — 28, 2, 26, 4, 28, 0, 16, 0 notes at the 8 positions — and takes minus the log of the frequency. The rule across each bar is the same quantity from the stated metrical weights, 1, 0.15, 0.5, 0.15, 0.85, 0.15, 0.5, 0.15, normalised and logged the same way. Nothing makes the two agree. They put the eight positions in the same order, and they price the tunes' own rhythm a fifth of a bit apart — while differing by more than a whole bit about the quaver after the downbeat, which two notes in a hundred and four ever use. The two positions these tunes never touch at all are drawn at the floor, which is the same floor the melodic measure gives an interval nobody plays. A weight was always a probability waiting to be read as one.
Fig. 4 The position axis measured on the same three tunes: what each place in the bar costs, by the preference-rule weights and by how often these tunes actually put a note there. The measured route needs only the onsets, which is why the corpus was never asked for anything else.

Every rung from the ninth onward needs onsets and pitches and nothing else. The position axis is a histogram over where notes start; the pitch axis is a histogram over intervals; the eye-hand span counts noteheads. A format holding pairs of degree and length supplies all of it, and there was never an occasion to notice that the length was doing double duty as a duration and as a gap.

Adding a third number per note — the sounded length, beside the inter-onset interval — would cost three lines and would let every rung above be recomputed. What it would not supply is a repertoire: three tunes chosen for being universally known are three tunes with almost no rests in them whatever the format allows, and measuring a rest density needs music chosen for something else.

That is the honest state of the debt. The format is fixable in minutes and the corpus is not, and this rung has priced the term parametrically because the parameter is the thing that is missing.

What is left over is not a term at all

The fourteenth rung named one more thing beyond duration — the interaction between its two axes, which it could not reach — and the accounting above still treats the three terms as adding.

They do not, and the reason is easy to state and is measured two rungs from here. A leap does not land where an offbeat lands: tunes put their large moves on strong beats far more often than between them, which is a thing the metrical weights already assert and the corpus can check, so a reader who has seen where a note falls already knows something about how far it moved. The three terms above therefore over-count, and the amount they over-count by is a measurable quantity on the same three tunes.

That is the seventeenth rung, and it is measurable on the corpus this one has just complained about — because the interaction needs only onsets and pitches, which are exactly what the format holds. What matters for this one is only that the total of 4.79 is an upper bound rather than a value, and every reading load on this anchor is high by the same amount.

A rest is not the only way a note stops early

The decision the term prices is binary — does this note fill its gap or not — and there are two ways for the answer to be no, which the accounting treats alike and a reader does not.

A rest is written. The staff puts a symbol where the silence is, and the reader’s decision is made by seeing it: the information is on the page and the cost is the cost of reading it.

A staccato is not. The mark says the note is short and does not say how short, and the loudness ladder found that a staccato is a dynamic mark as much as a rhythmic one — its effect on what a listener receives is a matter of duty cycle rather than of duration. A reader meeting one has not been told a length; they have been told a manner.

Those are two different objects and only the first is in the arithmetic. A staccato costs the reader almost nothing on this accounting — one symbol, no length — and it removes the same amount of sound a rest would. So a repertoire that marks its short notes rather than writing rests after them pays a fraction of the term, and gets a performance that is less precisely specified.

That is a real trade and notation makes it constantly. A Classical wind part is full of staccato dots and comparatively short of rests; a Baroque keyboard part is the reverse. The accounting says the first is cheaper to read and the second is more exactly determined, and both traditions presumably knew which they wanted.

What the pictures cannot show

Neither r nor k is measured here. The rest density is a property of a repertoire and this collection has no corpus with rests in it; the four rest lengths are a guess at how many a reader distinguishes, and a reader who tells six apart pays 0.26 rather than 0.20 for the length. The shape is not sensitive to either — the binary term dominates for any plausible k — and the headline number is sensitive to both.

The model also assumes a rest is a property of the note before it, which is how the arithmetic is set up and is not how a staff writes one. A staff writes a rest as its own symbol occupying its own horizontal space, so a reader meets it as an object rather than as an attribute — and an object on the page has a position to be read, which is a cost on the other axis that this accounting has just failed to charge.

There is a subtler problem with the binary framing and it is worth stating because it would change the number. The decision priced here is per note — does this one stop early? — and a reader does not scan noteheads asking that question. They read a bar, see whether there is a rest in it, and only then attend to which note it follows. That is a decision per bar rather than per note, and at four notes to a bar it would divide the binary term by four while leaving the length term alone: 0.32 bits rather than 0.67. Which of the two is the right accounting is a claim about how a page is scanned, and this ladder’s own eye-movement rung is the only evidence available and does not settle it.

And the whole measure charges a reader for information they may not extract. A sight-reader plays approximately and corrects; nothing in an entropy is about accuracy, and the eleventh rung’s eye-hand span is the only place this ladder has anything resembling a performance measure in it.

What the term does to the eleventh rung’s span

There is a consequence downstream and it is the one a sight-reader would feel.

Four notes of a tune, and one of a leaping line. The eye–hand span is measured in notes, and it is 4 of them for an ordinary sight-reader. Holding the load fixed instead — the number of bits the reader is carrying — gives a span that depends on what the music is. The same load is 4.3 notes of a scale and 1.1 of a wide leaps. So a sight-reader looking a bar ahead in a scale is looking a beat ahead in a wide-leaping line, on the same page at the same tempo.
Fig. 5 The eye-hand span read as a load rather than as a note count: how many notes of each kind of line a reader can hold at one bit rate. A line whose notes cost more is a line a reader holds fewer of.

The eleventh rung found that a sight-reader’s eye sits a fixed number of notes ahead of the sounding one, and the twelfth replaced the note count with a bit rate: a reader holds a constant amount of information rather than a constant number of noteheads, so a hard line is read with a shorter span.

Adding 0.67 bits to every note shortens that span by the same proportion — about 14 per cent on the numbers above. On a span of four notes that is a little over half a note, which is not nothing and is not much either.

What would be much is a repertoire with a high rest density. At a third of the notes the term is 1.48 bits, the load is 5.6, and the span shortens by a quarter. A part full of rests is read with a shorter span than a part full of notes, which reverses the obvious expectation — a part with fewer things in it is not an easier part — and is a prediction an orchestral player would be able to judge immediately.

Whose notation, and what a rest is for

The staff notation this prices is the one common to European music since about 1600, in which a rest is a distinct symbol with its own hierarchy of values and its own placement rules. Other notations handle the same information differently and the differences are legible in the accounting.

A tablature writes what the hand does and often leaves duration to be inferred entirely, which sets r out of reach and charges nothing — and produces exactly the complaint every lutenist makes about it. A proportional notation writes time as horizontal distance, so a rest is a gap rather than a symbol and the decision is made by the eye without a mark to read: the term goes to zero and the position axis becomes continuous, which is a trade the twentieth century made deliberately in several notations and abandoned in most of them.

Neither of those is priced here, and both are the natural test of whether this accounting is measuring anything. A notation that removes a term should be easier to read by exactly that term, and that is a sight-reading experiment rather than an argument.

The historical evidence runs the other way and is worth setting against it. Proportional notation is easier to read on this accounting and has been proposed, in one form or another, as often as any other reform, and it has failed every time. Every explanation offered for that is about inertia, and inertia is a real force; but a reader who has learned the durational symbols is not paying the term the accounting charges, because the symbols have become chunks. A cost measured on a naive reader is not the cost a trained one pays, and this whole anchor’s measure is a measure over a distribution rather than over a person.

That does not make the arithmetic useless. It makes it a statement about learning: the term is what a reader has to acquire, rather than what they spend once they have.

Where this ladder goes next

Fifteen rungs, and this one paid a debt by finding that it could not be paid where it was owed. The term is zero on the corpus and 0.67 bits on a repertoire with ordinary rests in it, and the difference between those two numbers is a fact about a data structure.

What is owed now is the tie, and the section above has already named why. A rest is a note stopping early and the accounting handles it as an attribute; a tie is something else entirely — it is the one mark on the staff that makes a single note out of two noteheads, so it puts an object on the page that a reader must read completely and then not play. Every quantity on this anchor is charged per notehead, and a tied continuation is a notehead carrying no event. What that costs is not the decision that identifies it; it is the decision plus the whole reading of a notehead that turns out to have been unnecessary, and the second is the larger of the two.

Part 15 of 18

One essay in the series on notation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DurationEntropyInformationInter-onset intervalNotationSight-reading