A form is sharp at the bottom and vague at the top
Assumes: A count is not an estimate · A proportion is only as fine as its two durations
Every proportion the account here has priced is a single ratio: one section against another, at one level of one piece. A form is not one ratio. A movement divides into sections, a section into phrases, a phrase into bars, and each of those divisions is a proportion in its own right, made of durations an order of magnitude shorter than the one above.
That matters because the blur on a duration is not a constant. A stretch of minutes judged afterwards is blurred by about thirty-seven per cent; a few seconds judged without counting by fifteen; a single second by seven and a half. So the hierarchy is not uniformly vague. It should be sharp at the bottom and vague at the top, and how much sharper is a number.
The timed hierarchy is almost flat
The expected answer is a smooth improvement downward, and it is not what the arithmetic gives.
Timed, the seven levels return 2.14, 2.14, 2.14, 2.14, 2.14, 5.21 and 5.21 distinguishable proportions. Five of the seven are identical and the last two jump together. The spread from top to bottom is a factor of 2.4, and it is a single step rather than a slope.
The reason is that a Weber fraction is not a continuous function of duration. The bands used throughout are three: an interval near a second, several seconds uncounted, and a stretch of minutes judged afterwards. Every level of a piece longer than half a minute falls in the same band, and an eight-minute movement halved five times is still thirty seconds at the fifth level. The hierarchy does not reach the finer bands until it is down among the phrases.
So a listener timing a form has one resolution for nearly the whole of it. The proportion between two halves of a movement, between two sections of a half, between two phrase-groups of a section: all three are judgements at thirty-seven per cent, and all three admit about two distinguishable answers.
The counted hierarchy sharpens
The count behaves completely differently, and the difference is the essay.
Counted in units of two seconds, the same seven levels return 2.2, 2.5, 3.1, 4.1, 5.4, 12.2 and 10.3 — a spread of 5.5 from top to bottom, climbing at nearly every step. The top of the form is no better counted than timed, because a count of 240 bars has almost certainly lapsed; the bottom is more than twice as good.
That climb is what an induced hypermetre buys a listener, priced one level at a time. A count improves as the thing being counted gets shorter and a clock does not. That is the previous essay’s arithmetic seen from a different angle: the lapse rate is what ruins a long count, and the levels of a hierarchy differ in exactly the quantity lapse punishes. Sixteen bars can be counted reliably; two hundred and forty cannot.
The one place the counted column falls rather than rises is the bottom, from 12.2 to 10.3, and that is the other failure arriving. At four units a slip is a quarter of a unit’s worth of error and the slip term has stopped falling fast enough to pay for itself. The counted hierarchy is best in the middle of its own range, which is what the essay that priced counting against timing found for a single section and is here a property of a whole form.
Neither hierarchy separates the golden section
There is one result that holds at every level of both columns, and it is worth stating on its own because it is the strongest negative the account here has produced.
At no level does a timed judgement separate 3 : 2 from the golden section. Not at the top, not at the bottom, not at the phrase level where the durations are a few seconds. The two ratios are 1.500 and 1.618, which is a log distance of 0.0757, and one criterion’s worth of ratio spread is 0.51 in the minutes band and 0.21 in the seconds band. Even the single-second band, at 0.106, is larger.
the first essay computed the Weber fraction at which those two merge and put it at 5.4 per cent, finer than any band the essays here carry. This essay says the same thing about a whole form at once: there is no level of any piece at which a timed listener can hear the difference between a section divided in three-to-two and one divided at the golden section. the next essay asks what that does to the claims analyses make.
Counting changes that, and only at the bottom. At twelve or more distinguishable proportions between 1 : 1 and 3 : 1, the spacing is fine enough for the two to be separate objects — which puts the only level at which a golden-section claim is audible at the phrase, over spans of ten or fifteen seconds, in a passage whose metre is stable enough to count.
What the flatness costs a composer
The flat timed column has a consequence for writing that is worth separating from the consequence for analysis, because they point opposite ways.
A composer choosing proportions at the top of a form is choosing among two audible categories. That is not a reason to stop choosing — a balanced movement and a lopsided one are different objects and the difference is large — but it does mean the precision of the choice is spent on nothing. A movement in 3 : 2 and a movement at the golden section are, to a timing listener, the same movement.
Where precision is not spent on nothing is further down, and the arithmetic says how far down: below about thirty seconds a section enters the seconds band and the resolution more than doubles, and below about ten seconds with a countable pulse it more than doubles again. A composer who wants a proportion heard has to put it where the durations are short, which is the phrase level and the phrase-group level, and which is where the return that is shorter than its first hearing found the storage measure also does its work.
That is a claim about what proportion is for, and it is at odds with the way proportional analysis is usually done — which starts at the top, with the whole movement, and works down. The resolution runs the other way.
Which level a proportional claim can be made at
Putting the two columns together gives a short answer to a question analyses ask constantly and rarely state.
A proportional claim about a whole movement is not a claim about anything a listener hears. Two distinguishable answers over the range from 1 : 1 to 3 : 1 means a listener can tell a balanced movement from a lopsided one and nothing finer. Every published ratio at that level — 3 : 2, 5 : 3, the golden section — is the same judgement.
A proportional claim about a phrase-group may be. At five to twelve distinguishable answers there is real structure to hear, and the claim becomes checkable in principle: a listener asked whether this phrase was longer than that one has a basis for answering.
And the level in between is where the count decides. At sixteen to sixty units — sections of half a minute to two minutes — timing gives two answers and counting gives four or five. Whether a listener hears the proportion of a sonata’s exposition to its development depends entirely on whether the metre held well enough to count it, which is a property of the piece rather than of the listener — and which an eight-bar phrase that accelerates shows a composer can take away deliberately.
Dividing in threes rather than halves does not change the shape of either column. That is worth checking because the branching factor is the one thing about a form that an analyst chooses, and if the finding depended on it the finding would be about the analysis.
The one number that is worse at the bottom
Two of the three routes improve downward and one of them does not, and it is the one the second essay built.
Storage is a count of bars a listener could not have predicted, and it is a proportion of two such counts. At the top of a form the counts are large — a section holds dozens of unpredictable bars — and their relative sampling error is small. At the bottom they are tiny: a four-bar phrase holds one or two new bars on a second hearing, and a ratio of one to two is not a proportion, it is a coin.
So the storage measure runs the opposite way to the other two. It is informative about the proportions of a whole form and meaningless about the proportions of a phrase, which is exactly the reverse of both timing and counting, and it means the three measures do not merely differ in precision — they differ in where their precision is.
That reframes the disagreement the previous essay described. It is not that a listener has three noisy instruments and should use the best. It is that the three are instruments for different levels of one object: storage at the top, counting in the middle, and the perceptual present at the bottom. A form read by all three at once is read at three resolutions, each sharpest where the others are blunt — which is a more interesting claim about listening than any of the three makes alone, and it is not one any essay here set out to make.
Which computation produced the numbers
A form is modelled as a piece of a stated length divided repeatedly by a stated branching factor, with each level’s characteristic duration the total divided by the branch to the power of the level. Each level’s timing band is the one the essays here carry for that duration, and its Weber fraction is the band’s midpoint.
The timed spread at a level is one criterion’s worth of the ratio spread at that fraction — the same measure the first essay uses, with the factor of root two for comparing two independently timed durations. The counted spread is the previous essay’s mixture, with the number of units at each level taken as the level’s duration divided by a stated unit, and with the same slip and lapse rates.
The number of distinguishable proportions is the log of three divided by the spread, so every figure in both columns is on the same scale as the first essay’s own headline numbers and can be read against them directly.
The unit is two seconds, which at a moderate tempo is a bar of four. Changing it moves the counted column bodily and does not change its shape, because it multiplies every level’s unit count by the same factor.
A level is not a thing a piece has
One assumption runs under every figure here and it is the kind that is easy to import without noticing: that a form has levels, that they are discrete, and that a listener judges a proportion at one of them.
A binary tree is a convenient way to generate seven characteristic durations from one piece. It is not a description of any movement. A sonata exposition has a first group, a transition, a second group and a closing group — four parts of quite unequal length, at one level — and the transition is not a sibling of the first group in any sense a listener would recognise. Below that the phrases are four bars, six bars and three, and below that the bars are equal only if nothing is written in five.
What the model actually needs is much weaker than a tree, and saying so is what keeps the finding from resting on the tree. It needs only that a form contains durations at several scales, and that a listener comparing two of them is comparing durations of roughly that scale. The resolution then follows from the scale alone, which is why dividing in threes gives the same shape as dividing in halves and why an irregular division would too.
So the columns should be read as a function of duration rather than of depth. At eight minutes a proportion admits two answers, at thirty seconds two, at fifteen seconds five, at seven seconds five and countable twelve — and where those durations sit in a particular piece’s structure is that piece’s business.
Where the model stops
The bands are steps and a listener is not. The three timing bands are a summary of a literature, and the underlying quantity surely varies continuously with duration. Modelling it as three steps is what makes the timed column flat, so the flatness is partly an artefact — a continuous function would give a gentle slope rather than a plateau and a jump. What survives is the magnitude: the whole range from a second to ten minutes spans a factor of five in Weber fraction, which is small next to the factor of five hundred in duration.
Every level is treated as independent. A listener judging a section’s proportion has already judged the phrases inside it, and those judgements are evidence about the section. A model that let information flow up the hierarchy would give better answers at the top than this one does, and the amount is not small: four phrase-level counts add up to a section-level one.
And the hierarchy is regular. Real forms are not binary trees. A sonata exposition is not two equal halves, and a rondo’s refrain is not the length its clock says, so the levels above are characteristic durations rather than actual ones.
What the picture cannot show
It cannot show the perceptual present. Below about three seconds a listener is not comparing two remembered durations at all; they are hearing one thing, and the window inside which a series of events is heard as one is a different mechanism from either route here. The bottom level of an eight-minute form in sevenths is inside it, and the figures there should not be trusted.
Nor a listener who has heard the piece before. Everything here is a first hearing. A listener who knows the form is not estimating its proportions; they are recalling them, and how much of this is new counts what a second hearing costs, and the second essay’s own findings about what a return costs suggest recall is a different measure again.
And it cannot show what a composer was doing. A form built on a proportion the composer computed is a fact about the score whether or not anybody can hear it. This essay is about what is audible, and the two questions have been run together often enough in the literature on proportion that keeping them apart is most of the work.
Still open: whether the levels pool
The strongest simplification above is that each level is judged on its own, and it is the one that would change the answers most.
A listener who has counted four phrases of four bars has, in principle, counted the section: sixteen bars, by a route that never required a sixteen-unit count to survive. If counts pool that way, the lapse rate that ruins the top of the counted column is the wrong rate to use there, and the column would climb toward the bottom instead of turning.
Whether they do is an empirical question with a clean shape. Pooling requires that a lapse at one level not destroy the level above it, and the obvious reason it might is that both are maintained by the same pulse: a listener who loses the bar has lost the phrase as well. If lapses are perfectly correlated across levels, a hierarchy is exactly one count and this essay’s figures stand. If they are independent — because the phrase level is tracked by cadences and the bar level by beats, which are different cues — then a four-level count of a movement is far more robust than any single count of it, and the vague top of a form is an artefact of assuming a listener counts in one unit.
That is a question about cues rather than about arithmetic, and it is the first thing the account here has needed that a stopwatch could not settle.
Part 4 of 6
One essay in the series on proportion. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
DurationHypermetreMusical formPerceptual presentPhraseProportionWeber fraction
- The ceiling everybody names is the loose one duration, musical form, perceptual present, phrase
- A detector whose resolution the performance sets perceptual present, phrase
- An ending is a deceleration duration, perceptual present
- An ending that exists so a bigger one can hypermetre, phrase
- An expectation cannot rescue a cycle too slow to time duration, weber fraction
- The level the tempo chooses perceptual present, phrase