Field

Form and structure

The shape a piece has in time, computed rather than labelled. A distance matrix finds the sections before anybody names them, a phrase turns out to be a number of seconds rather than of bars, and an ending is five separate signals that can each arrive without the others.
thirty-two-bar AABA, every bar against every other bar. A self-similarity matrix of 32 bars of thirty-two-bar AABA. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.5. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.

A piece is mostly itself again

Take a piece of music, encode each bar as the notes sounding in it, and compare every bar with every other bar. The picture that comes out has blocks and stripes in it, and those blocks and stripes are the form — arrived at by arithmetic that has never heard of an exposition, a chorus or a refrain.

Boundaries found by a local operator, at three kernel widths. Foote's checkerboard novelty computed on the self-similarity matrix of thirty-two-bar AABA, at kernel widths of 2, 4, 8 bars. The dashed verticals are where the encoding's sections actually change; nothing about them enters the computation. A peak is a place where the bars before resemble each other, the bars after resemble each other, and the two groups do not resemble each other.

The boundary is where the neighbourhood changes

A section boundary can be found by an operator that never sees a section. It walks the diagonal of a similarity matrix asking one local question — do the bars behind me resemble each other, do the bars ahead resemble each other, and do the two groups resemble each other — and where the answer is yes, yes, no, there is an edge. What it cannot find turns out to say more than what it can.

Redundancy in bits, and why the number needs a length beside it. Left: the LZ78 cost of each scheme divided by the cost of sending the same symbols flat, against how many bars are sent, with each bar coded as its chord and key. Every scheme is above 1 at a single chorus — the coder loses — and every one falls under it as the piece runs. Right: new dictionary phrases per bar at one chorus, which is the statistic that survives at short lengths.

How much of this is new

Repetition can be counted rather than looked at. Feed a piece's bars to a compressor and the bits it needs are a measure of how much of the piece is a repeat of an earlier part of itself. The measurement works, the number is real, and it turns out to be a statement about the description rather than about the music — which is the most useful thing it has to say.

A phrase is a number of seconds, and the bars follow the tempo. Phrase durations for 1, 2, 4, 8, 16-bar phrases at seven tempos, on a logarithmic seconds axis, with the 2 to 8 second window shaded. The window is a property of the listener and does not move; which bar count falls inside it is decided entirely by the tempo.

A phrase is a number of seconds

Musical phrases are described in bars, and four is the number everybody names. But the constraint that fixes a phrase is a property of the listener and is measured in seconds, so the bar count is whatever the tempo makes it. Across seven ordinary tempos the bar count that lands inside the window moves by a factor of eight, while the window itself does not move at all.

Two ways to fill eight bars, and only one of them accelerates. The period against the sentence, drawn as the lengths of their constituent units against position on a grid of eight bars. The ratio beside each row is the mean unit length in its second half divided by the mean in its first: the period at 1.00, the sentence at 0.67. A ratio below one is an acceleration — the unit shortening as the phrase approaches its arrival — and a ratio of one is a plan whose unit never changes length.

One of these eight-bar phrases accelerates

The sentence and the period both occupy eight bars, both end with a cadence, and both are recognised by ear rather than counted. What separates them is arithmetic. One halves its unit halfway through and the other does not, and the difference comes out as a single ratio — 0.67 against 1.00 — computed from nothing but the lengths of the parts.

Five signals, computed separately, and no total. The five components of closure for 6 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.

What makes an ending an ending

Cadences are ranked. The authentic one is strong, the plagal weaker, the deceptive weaker still — and the solver here measured the quantity that ranking is usually explained by and found it says something else entirely. What survives is not a weaker version of the ranking but a different kind of object, with five components and no total.

Five signals, computed separately, and no total. The five components of closure for 4 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.

An ending that exists so a bigger one can

Half of the cadences in tonal music are built to fail. A phrase that stopped convincingly at bar four would be a piece four bars long, so the ending at bar four is engineered to arrive and not to settle — and the components it withholds are exactly the ones its partner at bar eight supplies. Closure is nested, and the nesting is what turns two phrases into one thing.

The same induction, one level up. Bar-level onsets from thirty-two-bar AABA — a bar is marked where a section or a key begins — scored against hypermetres of 2, 3, 4, 6, 8 bars with the identical function the beat-level figures use. The best-fitting period is 8 bars, which at 108 beats a minute lasts 17.8 seconds.

The bar above the bar

A four-bar group is a bar whose beats are bars. That is not an analogy — it is the same computation, and the same metre-induction model produces one when it is handed bars instead of beats, unchanged. What decides where the hierarchy of levels stops is not in the arithmetic at all, and it is a number the phrase essay already measured.

The key plan is the shape. The key of a classical sonata-form movement against position in the movement, measured in steps along the chain of fifths from the home key. The positions are the proportions such a movement is described by rather than bar numbers from any one score. The furthest point is 4 steps out, sharing 3 of seven notes with home.

The key plan is the form

The large shape of a classical movement is not a shape at all, it is a journey — out to one key and back. Which key is not a matter of taste. Of the two keys that share six of their seven notes with home, only one introduces a note the home key does not use in any of its chords, and that note is the arriving key's own leading note. The departure is audible because of one accidental.

A cycle has no ending to compute, so it uses density instead. 5 layers over 32 cycles of a 12-step pattern, with each layer's entry and exit marked, and the onsets per step summed underneath. No chord changes and no cadence occurs; the closure vector is zero throughout, and every change a listener hears is a change in how many things are playing.

A cycle cannot cadence

Every component of closure is defined by a first time and a last time. Music built on a repeating cycle has neither, so the whole apparatus returns zero on it — not a small value, zero, at every setting. What such music uses instead is how many things are playing, and that is a curve which can be computed from the onsets and nothing else.

How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 1 case. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.427 and the smallest 0.427. 1 of 1 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.

How long until it comes back

A self-similarity matrix has a second reading that nobody looks for. Add up each diagonal instead of walking along one, and out falls repetition as a function of how long ago — a period, in bars, with no segmentation, no kernel width and no bar numbers anywhere in the answer. Five of the six schemes here report the length a listener would have named. The sixth reports something better.

The period, as the piece goes by. The strongest lag of thirty-two-bar AABA computed on only the bars heard so far, against how many bars that is. The final answer is 4 bars; it is revised 6 times on the way, and is not reached for the last time until bar 29 of 32, which is 91 per cent of the way through and 64 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have.

The form a first hearing cannot have

Every figure so far was computed with the whole piece in hand. Run the same methods over only the bars already heard and one of the two methods survives intact — the boundary operator turns out to be causal at a fixed delay of a few bars — while the other collapses. The period of a piece is not knowable until the piece is nearly over, and in two of the six schemes here not until its last bar.

Sensitivity and specificity on one dial. Aligned similarity — bar i against bar i+L, which is what a return is — for 4 eight-bar comparisons, as the key-invariance dial turns. One comparison is constructed: a literal repeat in the encoding, moved up a fifth, which is a stated manipulation because no scheme encoded here repeats a section in a new key. The shaded band is the margin between the two named comparisons, and it runs from 0.021 to 0.106.

The same thing somewhere else

A measure built on which notes are sounding calls a passage that comes back a fifth higher a stranger. There is a dial that fixes this, and turning it is supposed to be a trade — more sensitivity to a transposed return, less specificity against a coincidental one. It is not that trade. Two different statistics answer opposite ways, and the setting that would compromise between them is the worst one available.

thirty-two-bar AABA, as a strip of time. thirty-two-bar AABA laid out one cell per bar, coloured by section, with the roman numeral in each bar. the A section's turnaround is the ii-V every variant keeps; the bridge is a chain of applied dominants. At 108 beats a minute in 4/4 the whole of it lasts 71 seconds. Cut into 4 repeat units of 8 bars, 3 pairs of units agree on more than 50 per cent of their bars. 2 of them are not identical, and 2 of those 2 differ in a run of bars ending at the last bar of the unit; the changed bars are marked in orange.

Where a repeat is changed

Cut every scheme into its own repeat unit, compare each unit with every other, and ask where a repeat stops agreeing with what it repeats. The answer is that it stops at the end, in every case the corpus contains — and the number of cases the corpus contains depends entirely on where the threshold for "a repeat" is put. Moving it by nothing at all takes the count from two to twenty and the finding with it.

A tracker following a tempo that will not stay still. The tracker's error as a fraction of a beat, against beat number, for 3 rates of tempo change with a period-correction gain of 0.2. It settles at 0.025 of a beat behind at 0.5 per cent a beat, 0.092 of a beat behind at 2.0 per cent a beat, 0.206 of a beat behind at 5.0 per cent a beat. It does not lose the beat; it lags, by very nearly the rate divided by the period-correction gain, and the lag reaches a quarter of a beat at 6.3 per cent a beat — at which point the tracker is nearer the wrong onset than the right one.

A metre has to be able to change its mind

Replace the scoring function with a phase-corrected oscillator and the two failures reported earlier separate. The beat now survives two silent bars, drifting 23 milliseconds of a 560-millisecond beat. But it does not lose a moving tempo so much as lag behind it, by the rate over the correction gain — and a cadential ritardando that halves the tempo in eight beats is faster than the model can follow.

Three ways to arrive at the same final tempo. Tempo against position in the closing passage, ending at 35 per cent of the opening tempo, for curvature exponents 1, 2, 3. All three begin and end at the same tempo, so what separates them is the middle: at the halfway point they read 68 per cent for linear in score position, 75 per cent for constant deceleration, 80 per cent for q = 3. The straight line is the one nobody plays. Measured ritardandos fit the decelerating curves, which is the whole of Kronman and Sundberg's argument: a closing gesture has the shape of a body stopping rather than of a dial being turned, and the parameter that varies between performances is the final tempo rather than the shape.

An ending is a deceleration

Every performance slows down at the end and the slowing has a shape. Tempo read against score position is the velocity of a body stopping — a square root rather than a straight line — and the three candidate curves agree at both ends by construction, so the whole audible difference is in the middle, where they part by fifteen per cent of the passage's length.

The boundary operator run over only what has been heard. Foote's checkerboard novelty on thirty-two-bar AABA at a kernel width of 4 bars, computed twice: once with the whole piece available, and once using only the bars heard up to and including each bar. The kernel reaches 4 bars forward, so every cell it needs has been heard 3 bars after its centre — the retrospective curve replotted 3 bars to the right lands on the causal one, and the operator turns out to be causal at a fixed delay rather than blind. The dashed verticals are the encoding's real section boundaries and are not an input.

An ending that can be heard coming

Two measurements are both called hearing an ending coming and they point in opposite directions. By the halfway mark of an ordinary form almost nothing new arrives — and the cost of coding each bar has not fallen at all. Neither statistic says anything is about to stop, because no statistic over content can: predicting the next event well is not predicting that there will not be one.

A silence measured in two ways that were not chosen to agree. How far a carried beat drifts during a silence, in fractions of a beat, for beat periods of 200 ms, 550 ms, 2000 ms and a tempo estimate 5 per cent wrong — which is the published discrimination limen rather than a figure chosen here. Half a beat of drift is where the metre coming out of the silence is no longer the one that went in, and it is reached after 10 beats whatever the tempo, which is 2.0 seconds at 200 ms, 5.5 seconds at 550 ms, 20.0 seconds at 2000 ms. The shaded band is the psychological present, 2 to 8 seconds, measured by a completely different literature and used in this collection to bound a phrase. The two answers overlap: a silence under about 2 seconds is a rest inside something and one over about 5.5 is after it.

A silence long enough to be an ending

Four of the five closure components are present or absent. Silence is the one with a continuous scale, so it is the one that can be given a threshold — and the threshold arrives from two literatures that were not chosen to agree, landing between three and a half and five and a half seconds. In a large hall, the room's own decay uses up most of it.

How much of the rule a walk with no rule reproduces. Post-skip reversal in 20,000-note random walks with no melodic knowledge of any kind. An unbounded walk reverses after 50.0 per cent of leaps, which is the chance rate and is the check that the measurement is right. Confining it to 12 semitones raises that to 61.3 per cent. Reaching the 70 per cent that corpus studies report needs a central tendency of 0.95 — an almost deterministic pull back toward the middle at the edges of the range. A wall is not enough; there has to be a spring.

The leap that pays itself back

Every melody textbook teaches that a leap should be followed by a step in the opposite direction, and every corpus that has been counted agrees — around seven leaps in ten are answered that way. A random walk with two walls, no memory of the leap and no rule of any kind reverses after 61 per cent of them, and the residue is not a rule either. What is left when the walls are accounted for is a prediction the rule does not make, and it is the prediction that decides between them.

The arch is not a preference. Every sequence of 6 notes over 8 scale degrees — 262,144 of them, enumerated rather than sampled — classified by contour, under three constraints. With none, the nine classes are spread. Requiring the sequence to return to its starting degree leaves only the arch, the valley and the flat, at 42.2 per cent each for the first two. Requiring it to begin and end on the LOWEST degree leaves the arch alone, at 99.6 per cent. Nothing here prefers a rise followed by a fall; the constraint is that the melody comes home, and a melody that comes home from below has nowhere to go but up first.

The shape that survives everything else

Throw away a melody's key, its tuning, its instrument and the sizes of its intervals, and what is left is a string of pluses and minuses. That string is what a listener who cannot name a note still has, and it costs 37 per cent of the tune to keep. The arch that melodic shape is famous for is not in it as a preference: enumerate every six-note sequence that begins and ends on the lowest degree it uses and 99.6 per cent of them are arches, because a melody that comes home from below has nowhere to go first but up.

Where a twelfth comes from. The range of a walk with no walls, against how many notes it runs for, at three settings of the one parameter it has. The parameter is fitted to the post-skip reversal rate and to nothing else; the range is then read off. With no central tendency at all the walk passes two octaves by 60 notes and keeps going. At the setting that reproduces 70 per cent reversal — κ = 0.78 — the range is 11.9 semitones at thirty notes and 17.9 at a hundred and twenty. It grows logarithmically, so over the whole plausible length of a tune it sits between an octave and a fifteenth, and a twelfth is the middle of that. The three tunes carried here are marked and all three fall below the curve.

The twelfth, and where it comes from

Melodies occupy about an octave and a fifth, and an earlier essay set out to explain that by the singer's register break and found that it does not: the chest mechanism alone spans two octaves and a semitone. The answer is in a parameter the essay on leaps fitted and then put down. A walk with no walls whose central tendency reproduces the post-skip reversal rate has a range that grows logarithmically — six semitones at eight notes, twelve at thirty, eighteen at a hundred and twenty — so across every length a tune plausibly has, the span is between an octave and a fifteenth.

Where the page ends a phrase, and where the ear does. Twinkle, twinkle with two sets of phrase boundaries on it. The lower curve is a local boundary detector — a peak in how much the interval and the note length change from one to the next, with nothing in it about bar lines or harmony — and the marks above it are where the notation puts the phrase ends. It finds 100 per cent of them and 2 boundaries the page does not have. Where the two agree it is because a long note is sitting at the join; where they disagree the page is marking a grammatical unit and the detector is finding a perceptual one.

Where a phrase ends

Run a boundary detector over the three tunes used throughout and it agrees with the notated phrasing on one of them perfectly and on another almost not at all. The reason is which cue each tune uses: Twinkle's phrases all end on a long note, so a duration-weighted detector finds five of five with no false alarms; Ode to Joy's run on in crotchets and its phrasing is in the intervals, where a duration detector finds one of three and a pitch detector finds all three and eight others. No fixed weighting serves both, and the published one is worse on each tune than the single cue that tune uses.

What survives a change of encoding: verse and chorus. The same 32 bars of verse and chorus under five encodings, scored on the three things measured here measures. Mean off-diagonal similarity says how alike the piece looks to the arithmetic. Recall and precision are the novelty operator's boundaries against the 3 the section plan has, at a kernel of four bars. The period is the strongest peak of the lag profile, in bars. Under the bag of pitch classes every other figure uses, the piece is 90 per cent self-similar and the operator finds 0 per cent of the boundaries; under how far the root moved it finds 100 per cent. The period is the quantity that does not move.

The repeat that is not in the notes

Eight earlier essays compare bars by writing each one as a bag of pitch classes and taking a cosine. Nothing chose that encoding — the first used it and the other seven inherited it. Encode the same six schemes four other ways and one of the three findings survives untouched, one survives with different numbers, and one turns out to have been a statement about the encoding all along: the boundary operator finds none of the section edges in three schemes as a bag of pitch classes and every one of them as tonic, subdominant and dominant.

note length against the onsets. Every candidate metre's fit to a 16-step pattern with 4 onsets, plotted against how strongly note length is weighted. At a gain of zero the scoring is the onset-only one every earlier model used, and the winner is step 1 and step 4. At a gain of 0.05 the answer becomes step 4. 2 candidates are exactly flat — step 2 and step 3 have no onset on any strong position, so there is no credit for the cue to multiply and no weighting of it can move the line.

What the onsets left out

Eight essays induce a metre from a list of ones and zeros, and every failure they recorded was argued about as a failure of the rules. Two of the three are not. Note length is already in that list and the scoring throws it away: put it back and the son clave's two-way tie resolves to the notated downbeat. But the groove with its beat removed cannot be repaired by any cue at any strength, and the reason is arithmetic rather than empirical — the true phase has no onset on any of its strong positions, so there is no credit for a cue to multiply and its line is exactly flat.

The boundaries that survive each amount of smoothing. The local boundary strengths of Twinkle, twinkle read at every scale: the curve is smoothed with a Gaussian of the width on the horizontal axis and the peaks that survive are counted. Small scales give 11 boundaries and large ones give one, and the notation marks 5. The level with that many falls at a width of 2, where the model finds 100 per cent of the notated boundaries and 100 per cent of what it finds is notated — a comparison with no threshold in it, which is what the scale parameter buys.

A boundary at a stated level

A boundary detector run over three tunes agreed with the notation on one and barely at all on another, and left two things owing: a version with a scale parameter, and a version run on performance timings. Both are paid here, and they pay differently — the scale removes a free parameter from the comparison and does not rescue the hard case, while two per cent of rubato does.

Crescendo, and what the impression does. A crescendo of 20 dB over 8 seconds, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.02 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up.

Loud is relative, and it comes down slowly

The account of loudness had a model of a moment and the account of closure asked it for a model of a form. The published one exists and its content is a pair of numbers that are not the same: a listener's running impression of how loud the music is rises to meet a step in a fifth of a second and takes seven seconds to come back down. A twenty-decibel crescendo spread over eight seconds therefore buys almost no contrast at all, and the same twenty decibels taken as a step buys a factor of two.

Twinkle, twinkle, phrased at the level each tempo selects. The number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 5 boundaries, drawn as the flat line; the model matches it best at 160 beats per minute, where the present holds 8.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.

The level the tempo chooses

The boundary detector has a scale parameter and an earlier essay left it free, ending with the sentence that names this one: the scale is in notes and the psychological present is in seconds. The psychological present is two to eight seconds, a tempo converts one to the other, and the level a listener reads then stops being a parameter at all — which is a prediction with teeth, because the same tune at two tempos should be phrased differently at levels the arithmetic names in advance.

Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 3 parts to 8 adds 4.3 decibels of power and 0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 6.0 to 3.9 as the parts crowd into the same register.

The dynamics are in the score already

Count the parts in each bar, realise them in their ranges, put every partial in its critical band, sum the loudnesses and run the result through the two smoothers built earlier. What comes out is a dynamic curve for a piece with no performance in it anywhere — and it says that doubling the number of parts inside a fixed register adds three decibels of power and about one of loudness, because the extra parts land in bands that were already occupied. Let the register widen with the parts and the same arithmetic gives eight phon, which is what a tutti actually is.

Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness.

A final chord is not made loud by adding to it

An earlier essay on closure said the loudest cue an ending has needs a corpus rather than an arithmetic. The arithmetic was built one essay ago, so it does not. Run four ending textures through it and two things come out backwards: a final chord three parts thicker than the rest arrives *quieter* against the running impression than the passage it ends, and a texture that drops a part a bar does not get quieter at all until the bar where there is one part left.

Every assignment, at equal levels and at its own best balance. The 6 ways of putting 3 players on a chord, each drawn twice: hollow at equal levels, which is what an assignment ranking sees, and filled at the levels that balance the parts and then minimise roughness. Every scoring here is at the same total loudness, 23.3 sones, so two points are comparable. Solving the discrete problem first picks violin · clarinet · oboe; solving both at once picks violin · oboe · clarinet, and the two-stage answer costs 9.0 per cent more roughness. The orderings do not keep their places between the two columns, which is the whole of the argument: a ranking taken at equal levels is not a ranking.

Who plays what and how loud is one question

Two lines of argument, one about spectrum and one about loudness, each stopped at the same wall and each said so. One of them can choose who plays which note and has every player at the same level; the other can choose how loud each part is and has nobody assigned to anything. Put together they are a single problem with two kinds of variable, and solving it in stages picks a different answer from solving it at once — nine per cent rougher, at the same loudness, on an ordinary triad.

The tempo turns, and almost nothing moves. The earlier arrival reading — what is sounding at the final chord over what the listener has been hearing — swept over bar lengths from 0.5 to 5 seconds, which is 480 down to 48 beats a minute, at 3 closing lengths. Every curve is nearly flat. Across a tenfold change of tempo one gesture's reading moves by a factor of 1.201 and the other's by 1.098, while the gap between the two gestures — which is what that essay was measuring — is 1.228. The expectation was that the tempo would decide the answer, on the grounds that a two-second bar against a two-second release is a comparable pair. The premise is wrong in a way the sweep makes obvious: the thing being compared with the release is not a bar, it is the WHOLE ENDING, which is 2 to 8 bars long and is therefore far longer than the release at every tempo anybody plays. The running impression has caught up with the closing texture before the final chord arrives, at 0.5 seconds a bar and at 5, and what is left is the last bar's own jump.

The parameter that did not decide the answer

An earlier essay on closure ended by naming the tempo as the thing every number in it was resting on, and said it was the kind of parameter that had caused trouble before by turning out to decide the answer. Turned across a tenfold range at a closing gesture of fixed length it moves the reading by four per cent, against a twenty-three per cent gap between the gestures it is distinguishing. The parameter beside it in the same figure — how many bars the gesture occupies — moves it by twenty, and nobody had named that one at all.

How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.

A detector whose resolution the performance sets

The boundary detector lost its free parameter when the psychological present became a number of notes at a stated tempo, and what that held still was named at the time: a performance slows into a phrase end, so the number of notes inside the present is not the same everywhere in a tune — it falls exactly where a boundary is. Making the width follow the performance recovers half of what removing the parameter cost, and honestly leaves the other half.

A ritardando does not spend the diminuendo. The arrival reading under a deceleration into the ending, from no ritardando at all to a final tempo 30 per cent of the starting one — which stretches the closing bars from 12.0 seconds to 20.2. The expectation was that it would matter: a ritardando lengthens exactly the bars the gesture is happening in, so a diminuendo that would have been absorbed at a steady tempo gets more of the smoother's own time to be absorbed in. It moves the reading by 0.00 per cent. Every line here is flat to within the thickness of the line, which is the second time a tempo parameter has been swept here and found to do nothing.

The reading was a step response

Sweeping the tempo found it did not decide the answer. This one sweeps the deceleration across a factor of three and finds a null to five figures, and then sweeps the length of the closing gesture across a factor of forty-eight and finds it moves the reading by eight per cent — but not as a function of seconds. Sorted by seconds the twelve runs scatter; sorted by how many bars the instruction covers they fall into three tight groups. One sentence explains the null and the not-null together.

A written dynamic is an instruction to the listener's impression. Every earlier scoring holds one chord still. A passage is a succession, and the running impression of loudness carries a chord into the one after it, so what a marking asks for and what playing the marking produces are different things. Here is a five-chord passage with a written shape. Playing each chord at its own written loudness gives the running impression 2.4, 3.0, 4.2, 5.6, 4.0 sones against the 2.4, 3.0, 4.2, 5.6, 2.0 that were asked for — right until the last chord, where it misses by 2.0. Solving for levels that make the impression arrive at the marking does not fix it: the last chord's target is I, two parts, and it is unreachable — the correction runs to silence and the impression still sits 1.1 sones above. A subito piano after a full chord is not a level a player can produce. It is a rate of change, and the smoother's two-second release is what refuses it.

A subito piano is a rate, not a level

All three earlier essays score one chord held still. An orchestration is a succession, and the running impression carries a chord into the one after it — so a written dynamic is an instruction to the listener's impression rather than to the instantaneous sound, and there are markings that cannot be produced at all. The correction runs to silence and the impression still sits above the target.

One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics.

The dissonance arrives and the dynamic does not

A scoring decides two things at once and both of them have to be integrated by a listener before they exist. The loudness smoother's release is two seconds and the roughness window is thirty-seven milliseconds, and that ratio of fifty decides which of the two survives at the pace music is actually played. Nothing anybody performs is fast enough to blur a dissonance, and a great deal of it is fast enough to average a dynamic.

A rest is a diminuendo, and a long one. How far a listener's running impression of loudness falls during a silence, converted into the diminuendo that would have taken it the same distance. Half a second of nothing is worth 3.2 decibels, a second and a bit is worth 8.1, and two and a half seconds is worth 17. The marked line is three and a half seconds, which is where a gap starts to be heard as an ending rather than as a pause: at that length the reference has fallen by 24 decibels, which is more than a fortissimo to a pianissimo. A tempo swept over a factor of ten, a deceleration over a factor of three and a gesture length over a factor of forty-eight all returned the same reading to four significant figures. This one moves it by twenty-four decibels.

A rest is a diminuendo

Three parameters swept over factors of ten, three and forty-eight returned the same reading to four significant figures. The one manipulation left unswept moves it by twenty-four decibels: a silence. A listener's running impression decays at the loudness smoother's two-second release, so a general pause is a diminuendo nobody wrote, and at the length that makes a gap an ending it is worth more than any marking a composer has.

What the chord before takes out of the chord after. Five chords at 1.2 seconds each, with the roughness each one has on its own — its simultaneous masking and the threshold of hearing already applied — and the roughness it actually has once the chord in front of it has raised the threshold. Four of the five are untouched. The fifth, i, two parts, follows the only step in this passage that falls more than fifteen decibels, and it arrives into a hole: it is entirely below threshold for its first 13 milliseconds and takes 240 to get all of itself back. Masking can only remove partials, so it can only lower a roughness — and the chords it can reach are the ones a written dynamic has just made quiet, which are already the smooth ones. Across the passage the dissonance contrast goes from 3021 to 3113: the mask widens it by 3.0 per cent rather than eating it.

A soft chord has to fade in

Forward masking sits between the two integration times already in play — two hundred milliseconds against a thirty-seven millisecond roughness window and a two-second loudness release — and it was owed as the term that might eat the dissonance contrast. It does not. It widens it, by three per cent at a chorale's pace and fifty-nine at four chords a second, because it can only ever remove partials and it can only reach the chord a dynamic has already made quiet. What it does instead is stranger: one chord in the passage is entirely inaudible for its first twelve milliseconds and takes a quarter of a second to arrive whole.

A general pause of a bar is worth 4.9 decibels in a hall and 8.1 in silence. What a rest is worth as a diminuendo, in rooms with different reverberation times. The upper curve is the earlier figure, which assumed the sound stops when the players do; each curve below it lets the hall go on sounding, falling sixty decibels in its own reverberation time until it reaches the background. At the bar of silence a general pause usually is — about 1.2 seconds — the dry value is 8.1 decibels, shoebox concert hall keeps 60 per cent of it and gothic cathedral keeps 22. At 3.5 seconds, where a gap stops being a pause and becomes an ending, the same hall keeps 86 per cent. A long silence outlives any hall's tail and a short one does not, which is why the room costs the device most at exactly the length a composer writes it.

A rest needs a dry room

The twenty-four decibels a general pause is worth assume the sound stops when the players do. Put a hall under it and a bar of silence keeps 60 per cent of its value in a shoebox concert hall and 22 per cent in a cathedral, while a three-and-a-half-second one keeps 86 and 45 — because a hall's tail has a length and a rest either outlives it or does not. A written bar of silence is worth half its dry value at 2.76 seconds of reverberation, which falls between the concert hall and the stone church.

A bar of silence is worth 13.1 decibels written last and 6.0 written first. Four closing gestures, drawn against how many seconds of silence each contains. All four hold the same 12-part texture, write the same 6.0-decibel diminuendo over 4 seconds, and differ only in what order the diminuendo and the silence are written in. The quantity is how far the listener's running impression has fallen when the final chord arrives, in decibels of equivalent diminuendo. Written with the silence last, a general pause of a bar is worth 13.06 decibels and one of three and a half seconds is worth 28.3. Written with the silence first, both are worth 6.00 — exactly the diminuendo's own depth, because the music resuming after the silence puts the reference back at its own level. The silence written first is worth less than the silence written with no diminuendo at all, which reads 8.12 at a bar: a diminuendo placed after a general pause takes 2.12 decibels away and adds nothing.

A general pause is spent by the note after it

Whether a composer should write the pause before the diminuendo or after it looks like a question about how big the ensemble is. It is not. Forty decibels of ensemble are worth one decibel of silence, and the order is worth seven — because a running impression rises twenty times faster than it falls, so half a general pause is spent by ninety-seven milliseconds of sound.

An entrance stops being a loudness event and never stops being a colour one. The same oboe entering on the same note at the same level, against how many players were already sounding. Its contribution to the loudness falls from 15.5 phons to 0.32 — a factor of 48 — and crosses the one-phon difference limen at 5 players already playing. Its contribution to the roughness rises by a factor of 12.3 over the same range, because roughness is a sum over pairs and the entrant makes one new pair with everybody. Both curves are drawn as a share of their own largest value, since a phon and a squared pascal have no exchange rate. The claim is the two directions, not the crossing point of two units.

An entrance is a change of colour

Eight essays on orchestration move the assignment and hold the ensemble still, and a score does the opposite: it brings players in and takes them out. Loudness is a sum over parts and roughness is a sum over pairs, so the player who joins adds one term to the first and one to the second for everybody already there. What the entrance is worth in phons falls by a factor of forty-eight across the range an ensemble spans and crosses the difference limen at five players; what it is worth in roughness rises by twelve, and by a further factor of ten for every ten decibels the passage is played at.

Three clocks receive one entrance, and they do not agree about when. An oboe joining 5 players already sounding, at time zero, with each of the listener's three readings drawn as its own share of the change it eventually makes. The roughness window is 49 milliseconds wide and has half the change at 25; the short-term loudness smoother has half at 15; the long-term one, whose release is the two seconds an earlier essay is about, has half at 90. The two-second release is on the wrong side of the smoother to hide an entrance. Its attack is 99 milliseconds, so an entrance is received promptly and it is a departure that is not.

The release is on the wrong side

Whether the loudness model's two-second release makes an entrance inaudible has the answer no, for a reason the question did not anticipate. The smoother is asymmetric — ninety-nine milliseconds going up and two seconds coming down — so a rise is tracked twenty times faster than a fall, and an entrance is received promptly by every one of a listener's three readings. The colour of it arrives first, at twenty-five milliseconds against ninety, and the reading that moves with the ensemble is the one nobody would have picked.

The same player, arriving and leaving, read as a share of the change. One oboe joining 5 players and the same oboe leaving them again, with both loudness readings drawn as the share of their own change that has arrived. The entrance is half received in 90 milliseconds and the exit in 1.43 seconds, a factor of 15.9. The roughness readings, drawn faintly, are 25 and 25 milliseconds and lie on top of each other. A score that writes a diminuendo under a departing part is not softening the exit. It is doing the smoother's release for it, on a clock the smoother would otherwise take two seconds over.

A part that leaves is not a part that arrives

The same player, the same note, the same level, and the only difference is which way round it happens. A listener's loudness reading takes 1.43 seconds to receive half of a departure and 90 milliseconds to receive half of an arrival — a factor of sixteen with nothing asymmetric in the sound at all, since both readings integrate the same two states in the same order. The colour reading receives the two identically, because a window has no direction, so a departure is a change whose grain arrives at once and whose level takes most of two seconds.

Which chord of a passage has room for the part that is entering. An oboe entering on one note, tried at each chord of a five-chord passage, scored by how far its own partials sit above the threshold the ensemble already sounding puts over them. The best moment gives it 9.0 decibels of margin and the worst 0.3, a spread of 8.7 — and the best moment is not the quietest chord, which is vi, close below. Room for an entrance is spectral rather than dynamic. A chord with a hole in its written spacing need not have one in its spectrum, because the partials of its bass fill the middle whatever the notes above it do.

The chord that has room for an entrance

Three essays have made the ensemble something a score can change and none of them has asked when. The ensemble already sounding puts a masked threshold over whatever register an entering part takes, and that threshold is set by the voicing rather than by the dynamic — so the five chords of one passage differ by 8.7 decibels in how much of an entering oboe survives them, and the quietest chord of the five is the worst place in the passage to bring somebody in. Swept over the entrant's own pitch, the choice of moment is worth as much as the choice of register.

The schedule that hears every entrance best holds the high parts back. Six parts waiting to enter a five-chord passage over four sounding players, each entering once and staying: the schedule under which the least audible entrance is as audible as it can be made. brass on E3 enters at I, open with -1.4 decibels of mean margin over the mask; oboe on E4 enters at vi, close below with -2.6 decibels of mean margin over the mask; clarinet on G4 enters at I, open with -1.4 decibels of mean margin over the mask; voice on C5 enters at IV, close above with 2.1 decibels of mean margin over the mask; violin on G5 enters at V, bracketing with 6.9 decibels of mean margin over the mask; flue pipe on C6 enters at I, hollow with 4.0 decibels of mean margin over the mask. The least audible entrance is at -2.6 decibels and the margins sum to 7.6; of all 15625 schedules 0 have a better least audible entrance and 576 a larger sum.

Room is used up by whoever enters first

The chord with the most room for a part entering alone is a fact about that chord. It stops being a fact the moment two parts want it, because each part that comes in raises the mask over everybody after it. Given six parts waiting to enter a five-chord passage, choosing each part's moment the way one part's moment is chosen puts three of them into the same chord and lands in the bottom fifth of all 15,625 schedules. Placing them one at a time does no better. The schedule under which the least audible entrance is heard best is unique, and it brings the low and middle parts in while the texture is thin and holds the three highest back for the last three chords — because a high part keeps its room over a full texture and a middle part does not.

Six named proportions, as blurred as the durations that make them. Six proportions between two parts of a piece — 1 : 1, 4 : 3, 3 : 2, golden section, 2 : 1, 3 : 1 — placed on one axis by the logarithm of the ratio of the longer part to the shorter, and drawn as bars one criterion wide (d′ = 1) for a listener timing both parts with a Weber fraction of 7%, 15%, 35%. Bars that overlap are proportions that listener cannot tell apart. At 7%, 4 of 5 neighbouring pairs stay apart; at 15%, 3 of 5 neighbouring pairs stay apart; at 35%, 0 of 5 neighbouring pairs stay apart.

A proportion is only as fine as its two durations

Analyses of form measure proportions in bars and report them to three figures — a climax at 0.618, a section in the ratio 3 : 2. A listener has each part only as an estimate of how long it lasted, and a ratio of two estimates is blurred by both. Timed as well as anyone times a single second, eleven proportions fit between 1 : 1 and 3 : 1; timed from memory over minutes, two do. The golden section is told from 3 : 2 only below a Weber fraction of 5.4 per cent.

The same forms by the clock and by what is stored. Six forms, each drawn twice: its sections sized by their share of the bars, and sized by their share of what a listener has to store when a bar counts only if it is recognised from 1 bar of context. Returns are drawn pale with a dashed edge. twelve-bar blues: returns take 67 per cent of the clock and 13 per cent of the storage; thirty-two-bar AABA: returns take 25 per cent of the clock and 5 per cent of the storage; rondo, ABACA: returns take 40 per cent of the clock and 10 per cent of the storage; verse and chorus: returns take 50 per cent of the clock and 13 per cent of the storage; two eight-bar phrases: returns take 0 per cent of the clock and 0 per cent of the storage; a four-bar ostinato: returns take 88 per cent of the clock and 0 per cent of the storage.

A return is shorter than its first hearing

A rondo's refrain takes three fifths of the clock and a verse-and-chorus song is balanced to the bar. Count instead the bars a listener could not have predicted when they arrived, and the returns shrink to between a tenth and a quarter of what is kept — so a song equal by the clock is between three and seven times heavier in its first half. A coder that learns repeats one bar at a time says the halves are equal, and the two memories disagree by more than any proportion a listener could confuse.

A final chord stands above the impression for a fraction of a second. How far a final chord at the tutti's own level stands above the listener's running impression at the instant it is released, against how long it lasts, for four ways of arriving at it. Straight out of the tutti it stands above nothing at any length; after 1.2 s of silence the impression is 8.1 dB down, and the chord stands highest, 5.14 dB, when it lasts 54 ms; after 3.5 s of silence the impression is 23.7 dB down, and the chord stands highest, 14.42 dB, when it lasts 28 ms; after a 6 dB diminuendo the impression is 5.0 dB down, and the chord stands highest, 3.18 dB, when it lasts 42 ms. Every curve is level again by half a second, because the impression's attack of 99 ms catches the note's attack of 22 ms, so a held chord is released at the impression's level whatever preceded it.

A final chord stands out for a twentieth of a second

A general pause drives a listener's running impression down, and the final chord that follows is supposed to cash the fall in. It cashes in at most two thirds of it. The note's own loudness rises with a 22-millisecond constant and the impression with a 99-millisecond one, so after a bar of silence the chord stands furthest above the impression 54 milliseconds in, by 5.1 of the 8.1 decibels the silence bought, and after 206 milliseconds the two are within a phon of each other. A short stamp spends most of its life standing out; a chord held a second and a half spends a seventh of it.

A count is least reliable at both ends and best in the middle. How precisely a listener knows the length of a section they are counting, against how many units long it is, at three kinds of timing judgement. A slip — one unit miscounted, at 2% a unit — accumulates as a random walk, so its relative cost FALLS as the section lengthens. A lapse — the count lost altogether, at 1% a unit — compounds, so the chance of still having the count falls geometrically and a long enough section is certain to lose it. A listener who has lost the count is back to timing, so the two failures mix into a floor. Against a Weber fraction of 7.5% the count is worth most at 21 units, where it is 1.7 times finer than timing, and falls back under a quarter better by 98 units; Against a Weber fraction of 15% the count is worth most at 10 units, where it is 2.4 times finer than timing, and falls back under a quarter better by 101 units; Against a Weber fraction of 38% the count is worth most at 4 units, where it is 3.7 times finer than timing, and falls back under a quarter better by 102 units. The length at which it stops being worth much is nearly the same in all three, because it is set by the lapse rate alone.

A count is not an estimate

Both established routes to a proportion are estimates — a duration timed, blurred by a Weber fraction, and a duration stored, biased by what was new. A listener who has induced a hypermetre has a third, and it is exact until it fails. It fails two ways that pull opposite: a slip miscounts one unit and its relative cost falls as the section lengthens, while a lapse loses the count entirely and its chance compounds. The mixture has a floor at about four units, where counting is 3.7 times finer than timing, and it is worth almost nothing past a hundred.

Timing blurs a whole form evenly; counting sharpens it downward. A piece of 480 seconds divided 7 times, each level half the length of the one above, with how many proportions between 1 : 1 and 3 : 1 a listener can tell apart at each. Timed, the answer is 2.1 at the top and 5.2 at the bottom, a spread of 2.4 — because a timing judgement's Weber fraction is a step function of duration and almost every level of a piece falls in one step of it. Counted, in units of 2 seconds, the answer runs 2.2 to 10.3, a spread of 5.5. At no level does timing separate 3 : 2 from the golden section.

A form is sharp at the bottom and vague at the top

A movement is divided into sections, each into phrases, each into bars, and every level is a ratio of two estimates. Timed, the hierarchy is almost uniformly blunt — 2.1 distinguishable proportions at the top and 5.2 at the bottom, because a Weber fraction is a step function of duration and six of a piece's seven levels fall in one step of it. Counted, the same hierarchy runs from 2.2 to 12.2 and sharpens monotonically downward. At no level of either does timing separate 3 : 2 from the golden section.

The golden section and an equal division are one judgement. Where a boundary falls in a piece, as a share of its length, with the band a listener cannot tell from the golden section shaded. A stretch of minutes is judged with a Weber fraction of about 38%, so one criterion's worth of ratio spread around 0.618 covers everything between 0.492 and 0.730 — a quarter of the piece wide, and containing the halfway point. 1 : 1 and 4 : 3 and 3 : 2 and golden section and 2 : 1 are inside it. A claim that a climax falls at the golden section rather than at the middle is, at this resolution, not a claim about anything a listener could hear.

A golden section is a coin toss with six coins

An analysis that reports a climax at 0.618 of a piece has not tested one prediction; it has looked at a piece with several defensible boundaries and reported whichever landed nearest. The rate at which that happens under no hypothesis is one line of arithmetic, and the tolerance it needs is not a number chosen on the page — it is the blur a listener's own timing puts on the judgement. Over a stretch of minutes that blur covers everything from 0.492 to 0.730 of the piece, which contains the halfway point, and six candidate boundaries produce a hit eighty per cent of the time.

The breath is the looser ceiling nearly everywhere. How long a trained singer can hold a phrase on one breath, across a compass and at four dynamics, against the 8-second ceiling the psychological present puts on the same phrase. The flow through the folds rises with pitch and with loudness, so the breath ceiling falls both ways: at 60 decibels it runs 32.6 seconds at the bottom of the compass to 21.2 at the top; at 70 decibels it runs 23.1 seconds at the bottom of the compass to 15.0 at the top; at 80 decibels it runs 16.4 seconds at the bottom of the compass to 10.6 at the top; at 90 decibels it runs 11.6 seconds at the bottom of the compass to 7.5 at the top. The shaded line is the listener's ceiling and it does not move. The breath binds only where the two lines cross — 1 of the 40 cells drawn, all of them loud and high. So the constraint everybody names when asked why a phrase is the length it is, is almost never the constraint that decides it.

The ceiling everybody names is the loose one

Ask why phrases are the length they are and the answer given is the breath. It is arithmetic — usable lung volume over the air a note costs per second — and it comes out between fifteen and twenty-three seconds at a comfortable dynamic and between seven and twelve at a loud one. The ceiling the present moment imposes, the two-to-eight seconds inside which a stretch is heard as one thing rather than as a series, is two to three times tighter at almost every note and dynamic. A singer in an adagio is not running out of breath at the phrase end. They are running out of present.

A louder final chord stands higher and still stands for a fraction of a second. How far a final chord stands above the listener's running impression at the instant it is released, against how long it lasts. The chord is a struck six-note tonic; "a step louder" is the hammer velocity doubled, which raises its loudness 7.12 phons above the tutti's. At the tutti's level, after 1.2 s of silence, the impression is 8.09 dB below the chord and the chord stands highest, 5.14 dB, at 54 ms; a step louder, straight out of the tutti, the impression is 7.12 dB below the chord and the chord stands highest, 4.51 dB, at 36 ms; a step louder, after 1.2 s of silence, the impression is 15.21 dB below the chord and the chord stands highest, 9.46 dB, at 38 ms. Every curve returns to zero by half a second: the impression climbs to whatever level the chord is played at, so a louder mark is not a stand that lasts but a deeper fall to climb out of, and a silence and a mark add as depths.

A louder final chord is a deeper silence and a brighter sound

A final chord marked a step louder than the passage was supposed to stand above a listener's running impression for as long as it sounded, since the impression can climb no higher than the chord. It climbs exactly that high, and the stand closes in half a second as it always did. What a louder mark actually buys is depth — about seven phons, the same depth a second of silence buys — and a spectrum whose balance point sits most of a whole tone higher, which, unlike the stand, lasts for the whole chord.

With 4 of six required at the end, the best schedule hands parts over. Six parts entering, leaving and re-entering a five-chord passage over four sounding players, in the walk through all 64 sets of sounding parts that makes the least audible entrance as audible as possible, with no memory of the chord before and at least 4 of the six sounding at the last chord. I, open: clarinet on G4 enters at 0.6 dB; vi, close below: oboe on E4 enters at 0.3 dB, and clarinet leaves; IV, close above: brass on E3 enters at -0.7 dB, voice on C5 enters at 2.2 dB, and oboe leaves; V, bracketing: violin on G5 enters at 7.2 dB; I, hollow: flue pipe on C6 enters at 4.0 dB. The least audible entrance is -0.72 dB against -2.65 for the best schedule in which nobody leaves; the walk has 2 exits and 6 entrances.

An exit is worth nothing until the tutti is given up

Six parts entering a five-chord passage have a best schedule when each enters once and stays, and letting parts leave and come back was supposed to improve it. Searched over every set of sounding parts at every chord, it improves it by exactly nothing, with or without the chord before still masking — as long as all six must be playing at the end. Let one part be missing from the final chord and the weakest entrance gains 1.4 decibels; let two be missing and it gains 1.9, by a relay in which the parts with least room come in, are heard for one chord, and give way.

Checkpoints sharpen the middle of a form and leave its top vague. A piece of 480 seconds divided 7 times, with how many proportions between 1 : 1 and 3 : 1 a listener tells apart at each level: timed, counted in 2-second units, and counted with a second count of 16-second phrases that can mend a lapse in the first. 480 s: 2.1 timed, 2.2 counted, 3.9 with 0 per cent of lapses shared and 2.5 with 50 per cent of lapses shared; 240 s: 2.1 timed, 2.5 counted, 7.1 with 0 per cent of lapses shared and 3.1 with 50 per cent of lapses shared; 120 s: 2.1 timed, 3.1 counted, 13.2 with 0 per cent of lapses shared and 4.0 with 50 per cent of lapses shared; 60 s: 2.1 timed, 4.1 counted, 19.4 with 0 per cent of lapses shared and 5.4 with 50 per cent of lapses shared; 30 s: 2.1 timed, 5.4 counted, 18.1 with 0 per cent of lapses shared and 7.2 with 50 per cent of lapses shared; 15 s: 5.2 timed, 12.2 counted, 12.2 with 0 per cent of lapses shared and 12.2 with 50 per cent of lapses shared; 7.5 s: 5.2 timed, 10.3 counted, 10.3 with 0 per cent of lapses shared and 10.3 with 50 per cent of lapses shared. With the two counts failing independently, the level of 60 seconds goes from 4.1 to 19.4, and the whole piece only from 2.2 to 3.9.

Checkpoints sharpen the middle of a form, not its top

A listener who counts bars loses the count somewhere in a long section and is thrown back on timing the whole of it. A listener who also counts phrases can mend the lapse at the last phrase. If the two counts fail independently, the level a minute long goes from four distinguishable proportions to nineteen; the whole eight-minute piece goes only from two to four, because thirty phrases are long enough to lose a count as well. And if a fifth of lapses take both counts at once, three quarters of the gain is gone.

All fields · All essays