The axis that is not a time axis
Assumes: The notations invented for the overflow · The stave is not a ruler
Every rung of this ladder has the same shape: a notation fixes some coordinates of a musical object and leaves the rest to be supplied, and which ones it fixes is a claim about what the object is. The staff fixes the letter and leaves the pitch; the signature fixes the bar and leaves the grouping; the dynamic mark fixes an order and leaves the spectrum; the tablature fixes the action and leaves the sound.
All eight are about the vertical axis, or about symbols attached to it. Not one has asked about the other coordinate, and the other coordinate is the one every page carries and no essay here has counted.
What the measurement is, and why it is a measurement
Every stave on this site is laid out by VexFlow, which is a typesetting engine with the same job as a human engraver and the same conventions in it. So the spacing rule is not something to be looked up; it is available to be measured, by formatting a system and reading back where each notehead ended up.
Doing that on a run of notes from semiquavers to a semibreve gives spans of 69 points for an eighth, 80 for a quarter, 106 for a half and 156 for a whole. Fitted against duration those are
with d in crotchets. The constant is nearly the whole of the eighth-note’s width, and a note four times as long as another is about two and a half times as wide.
That is not an idiosyncrasy of one program. The published rules agree on the shape and differ on the exponent: Ross’s spacing has a note twice as long taking about √2 the space, which is a power of 0.5; Gould’s is flatter still. A pure power law with an exponent near a half and a constant-plus-proportional rule are hard to tell apart over the range a bar actually uses, and both say the same thing about the page.
Why the constant exists
The 58 points are not slack. They are what a notehead, its stem, its flags, its accidental and the white space a reader needs to see it as a separate event occupy — and none of those has anything to do with how long the note lasts.
So a page’s horizontal axis is a sum of two things that answer different questions. The proportional part answers when; the constant part answers is this legible. And the constant is much the larger over the range of durations most music uses, which means the page is mostly a legibility layout with a weak duration signal riding on it.
That has a consequence that is easy to check and easy to miss: a bar full of short notes is much wider than a bar containing one long one, even though the two occupy the same amount of time. Every engraver knows this as a fact about how bars are widened to fit a system; this arithmetic says it is not an adjustment, it is the rule.
How far the page is from the time it claims to represent
The useful summary is not the exponent; it is the error, and it is bounded in a way that makes it worse rather than better.
The error is zero at both ends by construction — the system starts when the phrase starts and ends when it ends — so all of the departure is in the middle. On a thirteen-note phrase with a mixture of durations it reaches 5.9 per cent of the system’s width under Ross’s rule.
Six per cent of a system is, on a normal page, about a centimetre — and a section below shows that six is a low number for this rule rather than a typical one. A reader following the page as a time axis is a centimetre out at the worst point, and the worst point is systematic: it is immediately after the longest note, because that is where the page has most under-spent relative to time and has the most catching up to do.
Which is the opposite of where a time axis would be least useful. A reader is most likely to be tracking position against time when the music is slow and sparse, and that is exactly where the representation is least faithful.
What this says about the two ladders that have been counting milliseconds
There is a consequence for the parts of this collection that are about timing, and it sharpens a complaint two other ladders have been making.
The microtiming ladder is about deviations of tens of milliseconds from a notated grid, and the perceptual-centre ladder is about tens of milliseconds between when a note starts and when it is heard. Both of them describe the notation as discarding those quantities, as does the essay about what the onsets left out. That is true, and it is not the whole of it.
What the page actually does is worse than discarding. It represents time on an axis that is not monotone in any useful sense against duration — a semiquaver and a minim are within a factor of two and a half of each other in width and a factor of eight in time — so a reader who forms any intuition at all about elapsed time from horizontal position is forming a systematically wrong one.
And the errors run in a particular direction. Short notes get proportionally more page than they deserve and long notes proportionally less, which means the page systematically exaggerates activity and compresses stasis. A held chord looks brief; a run of ornaments looks long. Whether that shapes how music is composed by people who write it down is not a question this arithmetic can answer, and it is a question worth having.
The one notation where the axis is honest
Tablature is the exception and it is instructive, because it is honest by giving up.
A lute or guitar tablature puts a number on a line for each string, and the horizontal axis is the action rather than the sound — where a finger goes, in the order it goes there. Some tablatures carry rhythm signs above the staff and some do not, and the ones that do not have an axis that makes no claim about time at all. It is a list, and it is honest about being a list.
The staff is the awkward middle: it nearly claims a time axis, by being roughly monotone in duration, and it does not deliver one. That is arguably worse than either extreme, and it is the reason the piano-roll representation — where the axis really is time and a note’s length really is its duration — looks so immediately legible to anyone who has never read music and so strange to anyone who has.
The horizontal axis counts notes and a reader does not. Eight quavers of a scale and eight quavers of a leaping line occupy the same width on the page and are four times apart in what they ask of the person reading them — which is the same failure as the timing one, in the other coordinate.
What a bar line does to the arithmetic
There is one feature of conventional spacing that pulls the other way and is worth naming, because it is the reason the six per cent is not larger.
Bars are spaced to be roughly equal in width where their contents allow. A system with four bars of common time tends to put them at four roughly equal widths, whatever is inside them — so the error accumulated inside a bar is largely reset at each bar line, and the whole page is a sequence of short, individually-wrong stretches rather than one long drift.
That is a real constraint and it improves the representation considerably. The thirteen-note phrase above is a single bar’s worth of unevenness; a page of four-bar systems has that error four times and never lets it compound. The bar line is, among other things, an error-correcting device on the horizontal axis, which is not one of the things the time signature essay lists it as doing.
Which computation produced the numbers
The spacing measurement is taken by formatting a system with the site’s own VexFlow, calling for the absolute x-coordinate of every notehead, and differencing. It is the same engine and the same call the stave figures on this site are drawn with, so the number is a property of the notation a reader is actually looking at.
The fit is least squares of span against duration in crotchets over the distinct durations present, giving an intercept and a slope. It is reported as intercept-plus-slope rather than as a power law because that is the shape the data has: a power law fitted to the same points gives an exponent of about 0.29, which is flatter than either published rule and is an artefact of the constant term rather than a description of a rule.
The published rules are implemented as pure powers of the duration: 1 for proportional, 0.5 for Ross, 0.42 for Gould, 0 for one-column-per-note. Each note’s page position is the running sum of its predecessors’ widths over the total; its time position is the running sum of its predecessors’ durations over the total; the error is the difference.
The thirteen-note phrase is invented, chosen to have a mixture of durations with the longest notes in the middle. A phrase with uniform durations has zero error under every rule, and the section below asks what a phrase that was not chosen would give.
Proportional notation is the one rule that is a time axis, and it fits the least on a page — which is the trade the whole convention rests on. The page is not a time axis because being one is expensive, and the cost is measured in how often the player has to turn it.
The invented phrase is not a hard case, it is an easy one
Six per cent is one number from one phrase somebody picked, and the honest question is where that phrase sits in the range of phrases. Generating twenty thousand thirteen-note phrases at random from an ordinary pool of durations and measuring each one’s worst departure:
| rule | median | 90th percentile | 99th | worst seen |
|---|---|---|---|---|
| proportional | 0.0% | 0.0% | 0.0% | 0.0% |
| Ross | 7.6% | 11.9% | 16.2% | 22.2% |
| Gould | 8.8 | 13.7 | 18.5 | 27.1 |
| one column a note | 14.9 | 23.1 | 31.3 | 43.9 |
The essay’s phrase is below the median. A typical mixed phrase departs by seven and a half per cent under Ross’s rule rather than six, one in ten by twelve, and one in a hundred by sixteen.
The ordering matters more than the durations do, which is easier to see by keeping the same thirteen durations and rearranging them:
| ordering | departure |
|---|---|
| the best arrangement there is | 2.9% |
| as drawn above | 5.9 |
| long notes in the middle | 7.2 |
| long notes at the ends | 9.6 |
| all long notes first, or all last | 14.3 |
A factor of five, from one multiset of durations. And the worst case is the one that says what the rule is doing: a phrase that spends all its long notes at one end puts the page maximally out of step with the clock, because the whole of the under-spending happens before any of the catching up. That also corrects a guess in the caveats — long notes at the ends is not better than long notes in the middle, it is worse, because it is two runs of the bad case rather than one.
So the summary to carry is not six per cent. It is that a conventional page is out of step with the clock by five to fifteen per cent of a system, depending almost entirely on how the long notes are distributed, and that a chord chart’s one-column-a-note rule is twice as bad again at every percentile.
Where the model stops
One engine is not every engraver. VexFlow’s rule is one implementation and a human engraver’s is another; Ross and Gould are two more, and a nineteenth-century plate engraver’s is a fourth. What they share is the shape — a large constant and a weak duration term — and the specific 58 and 15 belong to one program.
Ties, beams, accidentals and lyrics are not in it. All four change spacing substantially and none of them is a duration. A bar with a beamed run and a bar with the same durations unbeamed are spaced differently, and a syllable of text under a note can widen it past anything the duration would ask for — so a vocal line’s spacing is partly a property of its words, which is a coordinate the clef rung never had to think about. The constant term is therefore not one constant; it is a distribution whose mean is 58.
The random phrases are not phrases. Twenty thousand draws from a pool of durations is a distribution over multisets and orderings, and real music is neither uniform over durations nor uniform over their arrangements — long notes cluster at phrase ends, short ones come in beamed runs, and a bar’s contents sum to the bar. A sample of real engraved music would give a different median, almost certainly a lower one, because the arrangements music uses are not drawn at random. What the sweep establishes is the range the rule can produce and where one chosen phrase sits in it, not what a page of Brahms departs by.
And the system justification is not modelled. A real engraver spaces a bar and then stretches or compresses the whole system to fill the page, which is a global operation on top of the local rule. That is what makes the within-system error the right thing to measure and the between-system comparison meaningless, and this essay has only computed the first.
Whose notation, and when
Proportional spacing — where horizontal distance really is proportional to time — exists, and where it is used says something.
It is the notation of the twentieth-century avant-garde: music with no metre to speak of, in which the page’s job is to say when rather than what, and in which a bar line would be a lie. It is also the notation of every piece of software that shows music as a piano roll, which is the representation this whole essay’s error is measured against and which is now much the commonest way music is looked at.
Conventional spacing is a compromise made for a reader who is playing, and a player does not read horizontal position as elapsed time. They read note values, which are symbols, and convert them to time using a tempo they hold. The horizontal axis exists to put the symbols in order and to keep them apart — which is exactly the job the 58-point constant does and the 15-point slope is nearly incidental to.
So the page is not a failed time axis. It is a successful list, and the list has a weak duration signal in it because a small amount of proportionality helps a reader group things. The time signature is a claim about how the symbols are grouped and it is on the same axis for the same reason.
What the picture cannot show
Whether any reader is misled. A trained musician reads durations as symbols, not as distances, and may form no spatial intuition about time at all. The error computed here is an error in a representation nobody may be using.
And it cannot show what a page is for, which is not the same as what it represents. A page has to fit on a stand, turn at a place where a hand is free, and be readable at three feet in bad light — and every one of those is a constraint on the horizontal axis that has nothing whatever to do with music.
Where this ladder goes next
Nine rungs, and the coordinate that had never been looked at is measured: the staff’s horizontal axis is a legibility layout with a weak duration term riding on it, and the departure from time reaches six per cent of a system on an ordinary phrase.
What the ladder owes now is the one thing both axes have in common and neither rung has asked. Every notation this ladder has measured — the staff, the tablature, the chromatic staves of the eighth rung — is a page, which is a two-dimensional object read in a fixed order by a reader who has to turn it. The count that would say something is how much music a system holds under each notation, because that is the quantity that decides how often a page turns and therefore what can be written for a player with two hands occupied. The horizontal rule measured here is one of the two terms in it, the vertical capacity from the eighth rung is the other, and neither has been multiplied by the other.
Part 9 of 18
One essay in the series on notation. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
DurationEngravingLegibilityNotationProportional notationRhythmStaffTempo
- A staccato is a dynamic mark duration, notation, tempo
- A phrase is a number of seconds notation, tempo
- A proportion is only as fine as its two durations duration, notation
- A reader does not read notes notation, staff
- A silence long enough to be an ending notation, tempo
- An ending is a deceleration duration, tempo