Form and structure

A proportion is only as fine as its two durations

Analyses of form measure proportions in bars and report them to three figures — a climax at 0.618, a section in the ratio 3 : 2. A listener has each part only as an estimate of how long it lasted, and a ratio of two estimates is blurred by both. Timed as well as anyone times a single second, eleven proportions fit between 1 : 1 and 3 : 1; timed from memory over minutes, two do. The golden section is told from 3 : 2 only below a Weber fraction of 5.4 per cent.

Assumes: A phrase is a number of seconds · An interval is two errors

A great deal of writing about musical form is about proportion. A binary movement is balanced when its halves are equal. A sentence spends half its length presenting an idea and half developing it. A development section is two thirds the length of the exposition it follows. And an influential school of analysis has found climaxes placed at the golden section of a movement — about 0.618 of the way through — in Bartók, in Debussy, and in a long list of composers who were not asked.

Every one of those statements is made in bars, and every one is quoted with more precision than a listener could possibly use. That is not a complaint about the analyses. A bar count is exact and it is what a score offers. But a proportion that is to be heard is a ratio of two durations, and a listener has each of those durations only as an estimate of how long something lasted. A ratio of two estimates is blurred by both of them.

A phrase is a number of seconds found that the unit a listener can hold is set by a clock rather than by a bar line. This essay asks the next question about the same clock: how finely it can compare two of the things it holds. The answer is a number of proportions, and it depends on the scale so steeply that the proportions a movement’s form is described in and the proportions a listener could possibly distinguish are different lists.

Six named proportions, as blurred as the durations that make them. Six proportions between two parts of a piece — 1 : 1, 4 : 3, 3 : 2, golden section, 2 : 1, 3 : 1 — placed on one axis by the logarithm of the ratio of the longer part to the shorter, and drawn as bars one criterion wide (d′ = 1) for a listener timing both parts with a Weber fraction of 7%, 15%, 35%. Bars that overlap are proportions that listener cannot tell apart. At 7%, 4 of 5 neighbouring pairs stay apart; at 15%, 3 of 5 neighbouring pairs stay apart; at 35%, 0 of 5 neighbouring pairs stay apart.
Fig. 1 Six named proportions on one axis, placed by the logarithm of the ratio of the longer part to the shorter. Each is drawn as a bar exactly one criterion wide, so two bars that overlap belong to proportions a listener with that timing cannot tell apart at that criterion. At a Weber fraction of 7 per cent four of the five neighbouring pairs stay apart; at 15 per cent three do; at 35 per cent none does, and the whole axis from 1 : 1 to 3 : 1 is one smear.

A ratio has two errors in it

The model is the oldest regularity in the study of time perception and among the best supported. When a listener estimates a duration, the spread of the estimate grows in proportion to the duration: a second is judged to within some fraction of a second, ten seconds to within the same fraction of ten. That fraction is the Weber fraction for duration, ww, and the regularity is called the scalar property of timing.

An estimate whose spread is a fixed fraction of its size is naturally a quantity on a logarithmic scale. Its spread in natural-log units is

σ=ln(1+w2),\sigma = \sqrt{\ln\left(1 + w^2\right)},

which for any fraction a listener could have is almost exactly ww itself: 0.050 at five per cent, 0.34 at thirty-five.

A proportion between two parts is the ratio of two such estimates, and on a logarithmic scale a ratio is a difference. Two independent spreads add in quadrature, so the heard ratio is blurred by 2σ\sqrt{2}\,\sigma. The square root of two is the whole cost of comparing rather than timing, and it is paid before anything about music enters.

The shape of that argument is not new here. An interval is two errors made exactly this move for pitch: two notes in succession are two pitch estimates, what a listener judges is their difference, so a melodic interval is heard less finely than either note in it. A proportion is the same object in the time domain. A section is a note whose pitch is its length, and two sections in succession are an interval between lengths — heard, like every melodic interval, a factor of 2\sqrt{2} less finely than either term.

The analogy holds one level down as well. How late is a different note found the edges between duration categories — where a long note becomes a different written value — at the midpoints between simple ratios. A proportion between sections is that same question about ratios, asked of spans a hundred times longer and answered with a clock several times worse.

Eleven proportions, then five, then two

Take the stretch of ratios from 1 : 1 to 3 : 1, which holds every proportion a treatise on form ordinarily names, and ask how many steps of one criterion’s width fit inside it. The count is the logarithm of three divided by the blur, and it is an exact function of the Weber fraction and nothing else.

How many proportions a listener can hold apart. The number of proportions that fit between 1 : 1 and 3 : 1 when each is one criterion's worth of ratio apart, for a listener who times each part with the Weber fraction on the horizontal axis. The curves are for d′ = 1 and d′ = 2. The shaded ranges are the fractions reported for an interval near a second (5% to 10%), several seconds, not counted (10% to 20%), minutes, judged afterwards (25% to 50%). At 7% there are 11.1; at 15% there are 5.2; at 35% there are 2.3 at d′ = 1.
Fig. 2 The number of proportions between 1 : 1 and 3 : 1 when each is one criterion’s width apart, against the Weber fraction for each part. The solid curve is the criterion used throughout, d=1d' = 1; the dashed curve is twice as strict. The shaded ranges are the fractions reported for three kinds of judgement. At 7 per cent there are 11.1 proportions, at 15 per cent 5.2, and at 35 per cent 2.3.

What decides which count applies is which Weber fraction applies, and that is not one number. The shaded bands in the figure are three kinds of duration judgement, and each is drawn as a range because the published values are ranges.

An interval near a second, compared with another at once, is the best case and the one most studied — the span where a series of events can be taken as a beat at all, and where timing has the most to work with. Reviews of interval timing put the Weber fraction there at around five to ten per cent, lower for trained listeners and filled intervals, higher otherwise. Eleven proportions between 1 : 1 and 3 : 1 is what that timing buys.

An interval of several seconds, judged without counting, is where the fraction is reported to rise. Studies reviewed by Simon Grondin find it growing once an interval outlasts a second or so and the listener is prevented from subdividing it, which is exactly the condition a section of music imposes on anyone not counting its bars. Ten to twenty per cent leaves four to eight proportions.

A stretch of minutes, judged afterwards, is the case a movement’s large proportions actually present. Retrospective duration judgements — made by a listener who was not told in advance that time would be asked about — are known to be both shorter and far more variable than prospective ones; individual estimates routinely scatter by a third of the duration or more. At twenty-five per cent there are 3.2 proportions between 1 : 1 and 3 : 1, and at thirty-five there are 2.3.

Two proportions is a verdict, not a scale. It is about equal and clearly unequal, with a region between them where a listener could say either. Every finer statement about the large proportions of a movement is a statement about the score.

The tuning essays rest on a number of the same kind: about five cents in the middle of the range sorts the tuning arguments into those about music and those about arithmetic. A Weber fraction for duration does that sorting for form, and it sorts far more of the arguments onto the arithmetic side.

The criterion is part of the number

A count of proportions means nothing without saying what “told apart” means, and the choice above is deliberately generous. At d=1d' = 1 a listener asked which of two passages divides more unequally is right about three times in four. That is detection, and it is well short of anything a listener would stake a description on.

Doubling the criterion, to about ninety-two per cent correct, halves every count. The dashed curve in the figure above is that criterion, and on it the stretch from 1 : 1 to 3 : 1 holds 5.6 proportions at the timing of a single second and 1.1 at the timing of minutes.

Six named proportions, as blurred as the durations that make them. Six proportions between two parts of a piece — 1 : 1, 4 : 3, 3 : 2, golden section, 2 : 1, 3 : 1 — placed on one axis by the logarithm of the ratio of the longer part to the shorter, and drawn as bars one criterion wide (d′ = 2) for a listener timing both parts with a Weber fraction of 5%, 10%, 20%. Bars that overlap are proportions that listener cannot tell apart. At 5%, 3 of 5 neighbouring pairs stay apart; at 10%, 2 of 5 neighbouring pairs stay apart; at 20%, 0 of 5 neighbouring pairs stay apart.
Fig. 3 The six named proportions again, each bar now two criteria wide, for a listener timing each part to 5, 10 and 20 per cent. At the best timing anyone has, three of the five neighbouring pairs stay apart; at 10 per cent, two; at 20 per cent — the top of the range reported for an uncounted span of several seconds — none, and 1 : 1 overlaps 4 : 3.

The figure is worth reading for the pairs that go first rather than for the counts. At a criterion a listener could report with confidence, 4 : 3, 3 : 2 and the golden section have merged even at five per cent, and at twenty the only proportions still separate are ones a whole category apart — the equal division against 2 : 1 or 3 : 1. A proportion is a coarse attribute of the thing it describes, and on this criterion it is coarse even at its best.

Which named proportions survive

The counts say how many proportions a stretch of ratios holds. The more useful question for a reader of analyses is which particular proportions a listener could tell from their neighbours, and that has a closed-form answer: two ratios merge at the Weber fraction whose blur equals the distance between their logarithms.

Where each pair of neighbouring proportions merges. For each pair of neighbouring named proportions, the Weber fraction below which a listener can still tell them apart at d′ = 1 (the solid bar) and at d′ = 2 (the tick). 1 : 1 and 4 : 3: 21%; 4 : 3 and 3 : 2: 8.3%; 3 : 2 and the golden section: 5.4%; the golden section and 2 : 1: 15%; 2 : 1 and 3 : 1: 29%. The shaded ranges are the fractions reported for an interval near a second, several seconds, not counted, minutes, judged afterwards.
Fig. 4 For each pair of neighbouring named proportions, the Weber fraction below which a listener can still tell them apart at d=1d' = 1 (the bar) and at d=2d' = 2 (the tick). 1 : 1 and 4 : 3 merge at 21 per cent; 4 : 3 and 3 : 2 at 8.3; 3 : 2 and the golden section at 5.4; the golden section and 2 : 1 at 15; and 2 : 1 and 3 : 1 at 29.

Read against the three bands, the pairs sort into three groups.

3 : 2 and the golden section need a listener at the very bottom of the best band — timing each part to 5.4 per cent, which is how well a trained listener times an interval of about a second presented alone. At the fraction any section of music is judged with, the two proportions are one proportion. The golden section and 5 : 3 are closer still, and would need a Weber fraction of 2.1 per cent, which nobody has.

The golden section and 2 : 1 part company at fifteen per cent, in the middle of the band for an uncounted span of several seconds. A climax placed at 0.618 of a phrase and one placed at two thirds of it are distinguishable by a careful listener if the phrase lasts a few seconds. Across a movement lasting minutes, they are not.

1 : 1 against 4 : 3, and 2 : 1 against 3 : 1, survive into the minutes band, at 21 and 29 per cent. These are the distinctions a listener can make about the large sections of a piece: equal or a little unequal, unequal or very unequal. They are exactly the two distinctions the verdict above contained.

The two phrases this does separate

The arithmetic is not only a machine for refusing claims. Between 1 : 1 and 2 : 1 — where almost every phrase proportion in the classical repertoire lives — it says how many distinctions a phrase can carry, and it says it at a scale where timing is good.

How many proportions a listener can hold apart. The number of proportions that fit between 1 : 1 and 2 : 1 when each is one criterion's worth of ratio apart, for a listener who times each part with the Weber fraction on the horizontal axis. The curves are for d′ = 1 and d′ = 2. The shaded ranges are the fractions reported for an interval near a second (5% to 10%), several seconds, not counted (10% to 20%), minutes, judged afterwards (25% to 50%). At 5% there are 9.8; at 10% there are 4.9; at 25% there are 2.0 at d′ = 1.
Fig. 5 The number of proportions between 1 : 1 and 2 : 1 against the Weber fraction, with the same bands and the same two criteria. At 5 per cent there are 9.8, at 10 per cent 4.9, and at 25 per cent 2.0 — so across a span of minutes the stretch from an equal division to a two-to-one division holds one step.

One of these eight-bar phrases accelerates separated the sentence from the period by a single ratio of unit lengths, 0.67 against 1.00. Those two numbers are a factor of 1.5 apart, the logarithm of which is 0.405. At a Weber fraction of fifteen per cent — an eight-bar phrase lasting some seconds, not counted — that distance is 1.92 criteria: the sentence’s acceleration is audible as a change of proportion, and comfortably. Stretch the same two designs across sections lasting minutes, at thirty-five per cent, and the same distance is 0.84 criteria, below the line.

So the arithmetic keeps the distinction theory draws at the scale theory draws it. The sentence and the period are proportions a listener can hear because they are phrase-sized. A sonata movement whose exposition and development stood in the same two ratios would be making a distinction nobody in the hall could follow by timing alone.

A bar added is counted, not heard

Eighteenth-century theory had a vocabulary for phrases that are not four or eight bars long. Heinrich Christoph Koch described phrases lengthened by repeating a segment, inserting a parenthesis, or adding an appendix after the cadence, and the five-bar and seven-bar phrases of Haydn are the textbook examples of the result. The question those phrases put to the arithmetic is sharp: how long does an extension have to be before a listener hears the phrase as longer?

A phrase of 8 bars lengthened, and when the length is heard. How distinguishable a phrase of 8 bars becomes from the same phrase with whole bars added, measured as d′ for a listener who times both with a Weber fraction of 7%, 15%, 35%. The dashed line is the criterion d′ = 1. At 7% the extension has to reach 0.8 bars; at 15% the extension has to reach 1.9 bars; at 35% the extension has to reach 4.9 bars. Below that, a longer phrase can only be known by counting it.
Fig. 6 A phrase of eight bars lengthened by whole bars, and the dd' between it and the phrase it was, for a listener timing both to 7, 15 and 35 per cent. The dashed line is the criterion. At 7 per cent the extension is heard at 0.8 of a bar; at 15 per cent it has to reach 1.9 bars; at 35 per cent, 4.9.

At 108 beats a minute in four-four an eight-bar phrase lasts 17.8 seconds, which puts it in the band for an uncounted span of several seconds. At fifteen per cent a one-bar extension sits at d=0.56d' = 0.56, a little over half the criterion. A five-bar phrase where four were expected, or a nine-bar one where eight were, is not longer to a listener’s clock. It is longer to a listener’s count.

That is a claim with a mechanism already in place. The bar above the bar showed that the same computation that finds a beat in onsets finds a four-bar group when it is fed bars. A listener who has induced a four-bar hypermetre does not need to time a phrase to know that one bar too many arrived before the cadence: the cadence lands on the wrong beat of the hypermeasure. The extension is heard as a displacement in a metre, with the precision a metre has, rather than as a duration, with the precision a clock has.

And the requirement grows with the phrase, twice over.

A phrase of 16 bars lengthened, and when the length is heard. How distinguishable a phrase of 16 bars becomes from the same phrase with whole bars added, measured as d′ for a listener who times both with a Weber fraction of 15%, 25%, 35%. The dashed line is the criterion d′ = 1. At 15% the extension has to reach 3.8 bars; at 25% the extension has to reach 6.7 bars; at 35% the extension has to reach 9.9 bars. Below that, a longer phrase can only be known by counting it.
Fig. 7 The same question for a section of sixteen bars, which at 108 beats a minute lasts 35.6 seconds, timed to 15, 25 and 35 per cent. The added bars needed are proportional to the section — 3.8 at 15 per cent — and the section’s greater length also moves it toward the worse bands, where 6.7 bars are needed at 25 per cent and 9.9 at 35.

The first factor is the scalar property itself: a section twice as long needs an extension twice as long. The second is the band. A section lasting half a minute is no longer a span a listener times as well as a phrase, so the fraction it is judged with rises as well. An extension of a quarter of a section can pass as no extension at all, and the only listener who notices is the one counting.

A proportion in bars is not a proportion in seconds

Everything so far has taken the durations at face value. The analyses it is testing do not: they measure in bars and quote the result as though a bar were a unit of time, which at a fixed tempo it is and in performance it is not.

The size of the discrepancy is easy to state against the numbers above. An ending is a deceleration found performances slowing into a close along a curve, and a closing section played fifteen per cent slower than the section it balances has grown by fifteen per cent in seconds. The logarithm of 1.15 is 0.140. The whole distance between 3 : 2 and the golden section is 0.076. An ordinary ritardando moves a proportion further than the distance an analysis rests on, and it moves it in the direction that makes the later part longer.

So a golden section located by bar count is a golden section of the page. In performance, a closing section held back in the ordinary way puts the same bar somewhere else in time, and a listener hears the time.

What the model takes for granted

That the two durations are timed independently. The blur of a ratio is 2\sqrt{2} times the blur of a duration only if the two errors are unrelated. A listener’s clock can run fast for a whole piece, which would move both estimates together and leave the ratio better than drawn. It can also be distracted in one section and not the other, which does the reverse. The independent case is the neutral one and nothing here measures which way real listening leans.

That the comparison is part against part. A proportion like 0.618 of a movement compares a part with the whole, and the whole contains the part, so the two estimates share an error. How that changes the blur depends on how a listener estimates the whole — as one span, or as the sum of its parts — and nothing here decides which. It does not rescue the golden section, whose distance from 3 : 2 is what it is on either reading.

That the fractions apply to music. The bands are measured on intervals marked by tones and clicks, and a section of music is full of events. Filled durations are timed differently from empty ones, and a listener attending to melody and harmony has less attention for time. The first usually helps and the second usually hurts. The bands are drawn wide partly to hold both.

That nobody counts. Every figure above is a listener with a clock and no counter. Musicians count, and a hypermetre counts on a listener’s behalf. Where counting is available the resolution of a proportion is a different quantity altogether, discrete rather than blurred, and the extension figures are the first place that difference shows.

What a proportion on the page cannot establish

That a listener compares sections at all. The figures price the comparison a listener would have to make; they do not show that one is made. The evidence that large proportions are weakly attended is old and awkward for the analyses. Nicholas Cook asked listeners whether pieces ended in the key they began in, and found the judgement barely possible beyond about a minute; Karno and Konečni rearranged the sections of the first movement of Mozart’s G minor symphony and found listeners’ preferences almost untouched. Neither experiment is about duration, and both suggest a form’s large shape is not held in a form where proportion could be read off.

That a proportion does nothing if it is not heard. A composer can build a movement on a ratio as a way of deciding where things go, and the movement can be better for it without any listener perceiving the ratio. The arithmetic here is about perception and is silent about composition.

That listeners weigh sections by clock time. A section dense with events and one nearly empty can occupy the same number of seconds and be remembered as very different lengths. Every duration above is a clock reading, and memory is not a clock.

The golden section, and what the test is a test of

The proportional analysis this arithmetic bears on most directly is the one associated with Ernő Lendvai’s writings on Bartók, which located climaxes and structural divisions at golden sections of movements measured in bars, and with Roy Howat’s Debussy in Proportion, which argued for proportional systems of the same kind in Debussy. Howat was also among Lendvai’s sharpest critics, on the grounds that the bar counts had been chosen to fit.

The figures here take no side on whether the proportions are in the scores. They answer a narrower question that both sides’ arguments pass through without stopping: if the proportions are there, at what scale could a listener tell them from their neighbours? The answer is that a golden section is audible as distinct from 2 : 1 inside a phrase of a few seconds and nowhere larger, and audible as distinct from 3 : 2 nowhere at all in music. A proportion claimed at the scale of a movement, to three figures, is a claim about how the music was made rather than about how it is heard — which may be exactly what its advocates meant, and is not what it is usually taken to mean.

The same arithmetic is kinder to the classical theory of phrase. The distinctions Koch and his successors drew — balanced and unbalanced phrases, sentence and period, a phrase and its extension — are made at the scale where timing is good and counting is easy, and they are made in the coarse categories that scale supports.

Still open: what memory keeps of a section

Every duration here is the time a section takes, and a section is not remembered as a time. It is remembered as what happened in it, and a passage heard for the second time is remembered differently from the first, because a listener who already has the material does not need to store it again. A rondo’s refrain occupies three fifths of the clock and is new only once. Measured by what a listener has to keep rather than by the seconds that pass, the proportions of any form with returns in it are not the proportions the bar counts give — and the storage-size account of remembered duration says that is the measure retrospective judgements follow. How much of this is new already counts the new material bar by bar, and applying that count section by section is arithmetic.

Part 1 of 6

One essay in the series on proportion. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

DurationHypermetreMusical formNotationPerceptual presentPhraseProportionWeber fraction