Form and structure

A form is sharp at the bottom and vague at the top

A movement is divided into sections, each into phrases, each into bars, and every level is a ratio of two estimates. Timed, the hierarchy is almost uniformly blunt — 2.1 distinguishable proportions at the top and 5.2 at the bottom, because a Weber fraction is a step function of duration and six of a piece's seven levels fall in one step of it. Counted, the same hierarchy runs from 2.2 to 12.2 and sharpens monotonically downward. At no level of either does timing separate 3 : 2 from the golden section.

Assumes: A count is not an estimate · A proportion is only as fine as its two durations

Every proportion the account here has priced is a single ratio: one section against another, at one level of one piece. A form is not one ratio. A movement divides into sections, a section into phrases, a phrase into bars, and each of those divisions is a proportion in its own right, made of durations an order of magnitude shorter than the one above.

That matters because the blur on a duration is not a constant. A stretch of minutes judged afterwards is blurred by about thirty-seven per cent; a few seconds judged without counting by fifteen; a single second by seven and a half. So the hierarchy is not uniformly vague. It should be sharp at the bottom and vague at the top, and how much sharper is a number.

Timing blurs a whole form evenly; counting sharpens it downward. A piece of 480 seconds divided 7 times, each level half the length of the one above, with how many proportions between 1 : 1 and 3 : 1 a listener can tell apart at each. Timed, the answer is 2.1 at the top and 5.2 at the bottom, a spread of 2.4 — because a timing judgement's Weber fraction is a step function of duration and almost every level of a piece falls in one step of it. Counted, in units of 2 seconds, the answer runs 2.2 to 10.3, a spread of 5.5. At no level does timing separate 3 : 2 from the golden section.
Fig. 1 A piece of eight minutes divided seven times, each level half the length of the one above, with how many proportions between 1 : 1 and 3 : 1 a listener can tell apart at each. The upper bar of each pair is timed and the lower is counted.

The timed hierarchy is almost flat

The expected answer is a smooth improvement downward, and it is not what the arithmetic gives.

Timed, the seven levels return 2.14, 2.14, 2.14, 2.14, 2.14, 5.21 and 5.21 distinguishable proportions. Five of the seven are identical and the last two jump together. The spread from top to bottom is a factor of 2.4, and it is a single step rather than a slope.

The reason is that a Weber fraction is not a continuous function of duration. The bands used throughout are three: an interval near a second, several seconds uncounted, and a stretch of minutes judged afterwards. Every level of a piece longer than half a minute falls in the same band, and an eight-minute movement halved five times is still thirty seconds at the fifth level. The hierarchy does not reach the finer bands until it is down among the phrases.

So a listener timing a form has one resolution for nearly the whole of it. The proportion between two halves of a movement, between two sections of a half, between two phrase-groups of a section: all three are judgements at thirty-seven per cent, and all three admit about two distinguishable answers.

Six named proportions, as blurred as the durations that make them. Six proportions between two parts of a piece — 1 : 1, 4 : 3, 3 : 2, golden section, 2 : 1, 3 : 1 — placed on one axis by the logarithm of the ratio of the longer part to the shorter, and drawn as bars one criterion wide (d′ = 1) for a listener timing both parts with a Weber fraction of 7.5%, 15%, 38%. Bars that overlap are proportions that listener cannot tell apart. At 7.5%, 4 of 5 neighbouring pairs stay apart; at 15%, 3 of 5 neighbouring pairs stay apart; at 38%, 0 of 5 neighbouring pairs stay apart.
Fig. 2 The first essay’s figure, which is what “two distinguishable answers” looks like: the named proportions on one ratio axis, each drawn a criterion wide, at three Weber fractions. In the minutes band the bars for 1 : 1, 4 : 3, 3 : 2 and the golden section overlap into one.

The counted hierarchy sharpens

The count behaves completely differently, and the difference is the essay.

Counted in units of two seconds, the same seven levels return 2.2, 2.5, 3.1, 4.1, 5.4, 12.2 and 10.3 — a spread of 5.5 from top to bottom, climbing at nearly every step. The top of the form is no better counted than timed, because a count of 240 bars has almost certainly lapsed; the bottom is more than twice as good.

That climb is what an induced hypermetre buys a listener, priced one level at a time. A count improves as the thing being counted gets shorter and a clock does not. That is the previous essay’s arithmetic seen from a different angle: the lapse rate is what ruins a long count, and the levels of a hierarchy differ in exactly the quantity lapse punishes. Sixteen bars can be counted reliably; two hundred and forty cannot.

The one place the counted column falls rather than rises is the bottom, from 12.2 to 10.3, and that is the other failure arriving. At four units a slip is a quarter of a unit’s worth of error and the slip term has stopped falling fast enough to pay for itself. The counted hierarchy is best in the middle of its own range, which is what the essay that priced counting against timing found for a single section and is here a property of a whole form.

A count is least reliable at both ends and best in the middle. How precisely a listener knows the length of a section they are counting, against how many units long it is, at three kinds of timing judgement. A slip — one unit miscounted, at 2% a unit — accumulates as a random walk, so its relative cost FALLS as the section lengthens. A lapse — the count lost altogether, at 1% a unit — compounds, so the chance of still having the count falls geometrically and a long enough section is certain to lose it. A listener who has lost the count is back to timing, so the two failures mix into a floor. Against a Weber fraction of 38% the count is worth most at 4 units, where it is 3.7 times finer than timing, and falls back under a quarter better by 102 units. The length at which it stops being worth much is nearly the same in all three, because it is set by the lapse rate alone.
Fig. 3 The count’s own floor, drawn over the range a form’s levels cover. Reading the levels of the figure above off this curve is the whole of why the counted column climbs and then turns.

Neither hierarchy separates the golden section

There is one result that holds at every level of both columns, and it is worth stating on its own because it is the strongest negative the account here has produced.

At no level does a timed judgement separate 3 : 2 from the golden section. Not at the top, not at the bottom, not at the phrase level where the durations are a few seconds. The two ratios are 1.500 and 1.618, which is a log distance of 0.0757, and one criterion’s worth of ratio spread is 0.51 in the minutes band and 0.21 in the seconds band. Even the single-second band, at 0.106, is larger.

the first essay computed the Weber fraction at which those two merge and put it at 5.4 per cent, finer than any band the essays here carry. This essay says the same thing about a whole form at once: there is no level of any piece at which a timed listener can hear the difference between a section divided in three-to-two and one divided at the golden section. the next essay asks what that does to the claims analyses make.

Where each pair of neighbouring proportions merges. For each pair of neighbouring named proportions, the Weber fraction below which a listener can still tell them apart at d′ = 1 (the solid bar) and at d′ = 2 (the tick). 1 : 1 and 4 : 3: 21%; 4 : 3 and 3 : 2: 8.3%; 3 : 2 and the golden section: 5.4%; the golden section and 2 : 1: 15%; 2 : 1 and 3 : 1: 29%. The shaded ranges are the fractions reported for an interval near a second, several seconds, not counted, minutes, judged afterwards.
Fig. 4 The Weber fraction above which each adjacent pair of named proportions merges. The 3 : 2 and golden-section pair merges at a finer fraction than any duration judgement anybody makes, at either criterion.

Counting changes that, and only at the bottom. At twelve or more distinguishable proportions between 1 : 1 and 3 : 1, the spacing is fine enough for the two to be separate objects — which puts the only level at which a golden-section claim is audible at the phrase, over spans of ten or fifteen seconds, in a passage whose metre is stable enough to count.

What the flatness costs a composer

The flat timed column has a consequence for writing that is worth separating from the consequence for analysis, because they point opposite ways.

A composer choosing proportions at the top of a form is choosing among two audible categories. That is not a reason to stop choosing — a balanced movement and a lopsided one are different objects and the difference is large — but it does mean the precision of the choice is spent on nothing. A movement in 3 : 2 and a movement at the golden section are, to a timing listener, the same movement.

Where precision is not spent on nothing is further down, and the arithmetic says how far down: below about thirty seconds a section enters the seconds band and the resolution more than doubles, and below about ten seconds with a countable pulse it more than doubles again. A composer who wants a proportion heard has to put it where the durations are short, which is the phrase level and the phrase-group level, and which is where the return that is shorter than its first hearing found the storage measure also does its work.

That is a claim about what proportion is for, and it is at odds with the way proportional analysis is usually done — which starts at the top, with the whole movement, and works down. The resolution runs the other way.

Which level a proportional claim can be made at

Putting the two columns together gives a short answer to a question analyses ask constantly and rarely state.

A proportional claim about a whole movement is not a claim about anything a listener hears. Two distinguishable answers over the range from 1 : 1 to 3 : 1 means a listener can tell a balanced movement from a lopsided one and nothing finer. Every published ratio at that level — 3 : 2, 5 : 3, the golden section — is the same judgement.

A proportional claim about a phrase-group may be. At five to twelve distinguishable answers there is real structure to hear, and the claim becomes checkable in principle: a listener asked whether this phrase was longer than that one has a basis for answering.

And the level in between is where the count decides. At sixteen to sixty units — sections of half a minute to two minutes — timing gives two answers and counting gives four or five. Whether a listener hears the proportion of a sonata’s exposition to its development depends entirely on whether the metre held well enough to count it, which is a property of the piece rather than of the listener — and which an eight-bar phrase that accelerates shows a composer can take away deliberately.

Timing blurs a whole form evenly; counting sharpens it downward. A piece of 480 seconds divided 7 times, each level half the length of the one above, with how many proportions between 1 : 1 and 3 : 1 a listener can tell apart at each. Timed, the answer is 2.1 at the top and 10.4 at the bottom, a spread of 4.8 — because a timing judgement's Weber fraction is a step function of duration and almost every level of a piece falls in one step of it. Counted, in units of 2 seconds, the answer runs 2.2 to 5.5, a spread of 5.5. At no level does timing separate 3 : 2 from the golden section.
Fig. 5 The same piece divided in threes rather than halves, which reaches the finer timing bands two levels sooner and has fewer levels to do it in. The shape of both columns survives the change: flat then stepped when timed, climbing then turning when counted.

Dividing in threes rather than halves does not change the shape of either column. That is worth checking because the branching factor is the one thing about a form that an analyst chooses, and if the finding depended on it the finding would be about the analysis.

The one number that is worse at the bottom

Two of the three routes improve downward and one of them does not, and it is the one the second essay built.

Storage is a count of bars a listener could not have predicted, and it is a proportion of two such counts. At the top of a form the counts are large — a section holds dozens of unpredictable bars — and their relative sampling error is small. At the bottom they are tiny: a four-bar phrase holds one or two new bars on a second hearing, and a ratio of one to two is not a proportion, it is a coin.

So the storage measure runs the opposite way to the other two. It is informative about the proportions of a whole form and meaningless about the proportions of a phrase, which is exactly the reverse of both timing and counting, and it means the three measures do not merely differ in precision — they differ in where their precision is.

That reframes the disagreement the previous essay described. It is not that a listener has three noisy instruments and should use the best. It is that the three are instruments for different levels of one object: storage at the top, counting in the middle, and the perceptual present at the bottom. A form read by all three at once is read at three resolutions, each sharpest where the others are blunt — which is a more interesting claim about listening than any of the three makes alone, and it is not one any essay here set out to make.

Which computation produced the numbers

A form is modelled as a piece of a stated length divided repeatedly by a stated branching factor, with each level’s characteristic duration the total divided by the branch to the power of the level. Each level’s timing band is the one the essays here carry for that duration, and its Weber fraction is the band’s midpoint.

The timed spread at a level is one criterion’s worth of the ratio spread at that fraction — the same measure the first essay uses, with the factor of root two for comparing two independently timed durations. The counted spread is the previous essay’s mixture, with the number of units at each level taken as the level’s duration divided by a stated unit, and with the same slip and lapse rates.

The number of distinguishable proportions is the log of three divided by the spread, so every figure in both columns is on the same scale as the first essay’s own headline numbers and can be read against them directly.

The unit is two seconds, which at a moderate tempo is a bar of four. Changing it moves the counted column bodily and does not change its shape, because it multiplies every level’s unit count by the same factor.

A level is not a thing a piece has

One assumption runs under every figure here and it is the kind that is easy to import without noticing: that a form has levels, that they are discrete, and that a listener judges a proportion at one of them.

A binary tree is a convenient way to generate seven characteristic durations from one piece. It is not a description of any movement. A sonata exposition has a first group, a transition, a second group and a closing group — four parts of quite unequal length, at one level — and the transition is not a sibling of the first group in any sense a listener would recognise. Below that the phrases are four bars, six bars and three, and below that the bars are equal only if nothing is written in five.

What the model actually needs is much weaker than a tree, and saying so is what keeps the finding from resting on the tree. It needs only that a form contains durations at several scales, and that a listener comparing two of them is comparing durations of roughly that scale. The resolution then follows from the scale alone, which is why dividing in threes gives the same shape as dividing in halves and why an irregular division would too.

So the columns should be read as a function of duration rather than of depth. At eight minutes a proportion admits two answers, at thirty seconds two, at fifteen seconds five, at seven seconds five and countable twelve — and where those durations sit in a particular piece’s structure is that piece’s business.

Where the model stops

The bands are steps and a listener is not. The three timing bands are a summary of a literature, and the underlying quantity surely varies continuously with duration. Modelling it as three steps is what makes the timed column flat, so the flatness is partly an artefact — a continuous function would give a gentle slope rather than a plateau and a jump. What survives is the magnitude: the whole range from a second to ten minutes spans a factor of five in Weber fraction, which is small next to the factor of five hundred in duration.

Every level is treated as independent. A listener judging a section’s proportion has already judged the phrases inside it, and those judgements are evidence about the section. A model that let information flow up the hierarchy would give better answers at the top than this one does, and the amount is not small: four phrase-level counts add up to a section-level one.

And the hierarchy is regular. Real forms are not binary trees. A sonata exposition is not two equal halves, and a rondo’s refrain is not the length its clock says, so the levels above are characteristic durations rather than actual ones.

What the picture cannot show

It cannot show the perceptual present. Below about three seconds a listener is not comparing two remembered durations at all; they are hearing one thing, and the window inside which a series of events is heard as one is a different mechanism from either route here. The bottom level of an eight-minute form in sevenths is inside it, and the figures there should not be trusted.

Nor a listener who has heard the piece before. Everything here is a first hearing. A listener who knows the form is not estimating its proportions; they are recalling them, and how much of this is new counts what a second hearing costs, and the second essay’s own findings about what a return costs suggest recall is a different measure again.

And it cannot show what a composer was doing. A form built on a proportion the composer computed is a fact about the score whether or not anybody can hear it. This essay is about what is audible, and the two questions have been run together often enough in the literature on proportion that keeping them apart is most of the work.

Still open: whether the levels pool

The strongest simplification above is that each level is judged on its own, and it is the one that would change the answers most.

A listener who has counted four phrases of four bars has, in principle, counted the section: sixteen bars, by a route that never required a sixteen-unit count to survive. If counts pool that way, the lapse rate that ruins the top of the counted column is the wrong rate to use there, and the column would climb toward the bottom instead of turning.

Whether they do is an empirical question with a clean shape. Pooling requires that a lapse at one level not destroy the level above it, and the obvious reason it might is that both are maintained by the same pulse: a listener who loses the bar has lost the phrase as well. If lapses are perfectly correlated across levels, a hierarchy is exactly one count and this essay’s figures stand. If they are independent — because the phrase level is tracked by cadences and the bar level by beats, which are different cues — then a four-level count of a movement is far more robust than any single count of it, and the vague top of a form is an artefact of assuming a listener counts in one unit.

That is a question about cues rather than about arithmetic, and it is the first thing the account here has needed that a stopwatch could not settle.

Part 4 of 6

One essay in the series on proportion. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DurationHypermetreMusical formPerceptual presentPhraseProportionWeber fraction