Concept

Perceptual present — where it appears

The span, of a few seconds, inside which a series of events is heard as one thing rather than as a series. Phrase lengths, the longest usable silence and the drift of a carried beat all land inside it, from unrelated measurements.

Named by 19 essays across 3 fields — each of them below, with the objects they name alongside it.

A phrase is a number of seconds, and the bars follow the tempo. Phrase durations for 1, 2, 4, 8, 16-bar phrases at seven tempos, on a logarithmic seconds axis, with the 2 to 8 second window shaded. The window is a property of the listener and does not move; which bar count falls inside it is decided entirely by the tempo.

A phrase is a number of seconds

Musical phrases are described in bars, and four is the number everybody names. But the constraint that fixes a phrase is a property of the listener and is measured in seconds, so the bar count is whatever the tempo makes it. Across seven ordinary tempos the bar count that lands inside the window moves by a factor of eight, while the window itself does not move at all.

form · Phrase
The same induction, one level up. Bar-level onsets from thirty-two-bar AABA — a bar is marked where a section or a key begins — scored against hypermetres of 2, 3, 4, 6, 8 bars with the identical function the beat-level figures use. The best-fitting period is 8 bars, which at 108 beats a minute lasts 17.8 seconds.

The bar above the bar

A four-bar group is a bar whose beats are bars. That is not an analogy — it is the same computation, and the same metre-induction model produces one when it is handed bars instead of beats, unchanged. What decides where the hierarchy of levels stops is not in the arithmetic at all, and it is a number the phrase essay already measured.

form · Metre induction
The period, as the piece goes by. The strongest lag of thirty-two-bar AABA computed on only the bars heard so far, against how many bars that is. The final answer is 4 bars; it is revised 6 times on the way, and is not reached for the last time until bar 29 of 32, which is 91 per cent of the way through and 64 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have.

The form a first hearing cannot have

Every figure so far was computed with the whole piece in hand. Run the same methods over only the bars already heard and one of the two methods survives intact — the boundary operator turns out to be causal at a fixed delay of a few bars — while the other collapses. The period of a piece is not knowable until the piece is nearly over, and in two of the six schemes here not until its last bar.

form · Repetition
The ranking is settled either side of one narrow band. Remembered repetition — each bar's best match to an earlier bar, discounted by exp(−Δt/τ) with Δt in seconds — for 6 schemes at 108 beats a minute, against the decay constant τ on a logarithmic axis. The order of the schemes changes only between 8 and 13 seconds; outside that band it is fixed, so an estimate of τ wrong by any amount that stays outside it leaves the ranking alone.

A return has to be remembered

A stripe four bars off the diagonal and a stripe twenty-four bars off it are the same ink and are not the same experience. Convert the lag axis to seconds, discount every comparison by how long ago it was, and the ranking of these six schemes by how repetitive they are changes — and the decay constant and the tempo turn out to enter the arithmetic as one number rather than two.

perception · Repetition
Three ways to arrive at the same final tempo. Tempo against position in the closing passage, ending at 35 per cent of the opening tempo, for curvature exponents 1, 2, 3. All three begin and end at the same tempo, so what separates them is the middle: at the halfway point they read 68 per cent for linear in score position, 75 per cent for constant deceleration, 80 per cent for q = 3. The straight line is the one nobody plays. Measured ritardandos fit the decelerating curves, which is the whole of Kronman and Sundberg's argument: a closing gesture has the shape of a body stopping rather than of a dial being turned, and the parameter that varies between performances is the final tempo rather than the shape.

An ending is a deceleration

Every performance slows down at the end and the slowing has a shape. Tempo read against score position is the velocity of a body stopping — a square root rather than a straight line — and the three candidate curves agree at both ends by construction, so the whole audible difference is in the middle, where they part by fifteen per cent of the passage's length.

form · Closure
A silence measured in two ways that were not chosen to agree. How far a carried beat drifts during a silence, in fractions of a beat, for beat periods of 200 ms, 550 ms, 2000 ms and a tempo estimate 5 per cent wrong — which is the published discrimination limen rather than a figure chosen here. Half a beat of drift is where the metre coming out of the silence is no longer the one that went in, and it is reached after 10 beats whatever the tempo, which is 2.0 seconds at 200 ms, 5.5 seconds at 550 ms, 20.0 seconds at 2000 ms. The shaded band is the psychological present, 2 to 8 seconds, measured by a completely different literature and used in this collection to bound a phrase. The two answers overlap: a silence under about 2 seconds is a rest inside something and one over about 5.5 is after it.

A silence long enough to be an ending

Four of the five closure components are present or absent. Silence is the one with a continuous scale, so it is the one that can be given a threshold — and the threshold arrives from two literatures that were not chosen to agree, landing between three and a half and five and a half seconds. In a large hall, the room's own decay uses up most of it.

form · Closure
What a contour costs to remember. A melody of n notes over 8 degrees carries 3 bits a note. Its contour carries fewer, and fewer than the number of distinct contours suggests, because the contours are not equally likely: at 6 notes there are 243 of them but the entropy is 6.59 bits, an effective alphabet of 96. Each further note adds 1.28 bits of contour against three of melody, so the shape keeps a stable 37 per cent of what is there however long the tune.

The part of the tune that is kept

Contour survives transposition, retuning, a change of instrument and a doubling of every interval, and the usual explanation is that it is what a listener retains. That can be counted rather than assumed. A six-note melody over eight degrees carries eighteen bits; its contour carries 6.59 — not the 7.92 the number of distinct shapes suggests, because the shapes are wildly unequal — and the effective alphabet is ninety-six out of two hundred and forty-three. Each further note adds 1.28 bits of shape against three of melody, and at about nine notes a contour is specific enough to pick one tune out of a thousand.

perception · Melody
Twinkle, twinkle, phrased at the level each tempo selects. The number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 5 boundaries, drawn as the flat line; the model matches it best at 160 beats per minute, where the present holds 8.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.

The level the tempo chooses

The boundary detector has a scale parameter and an earlier essay left it free, ending with the sentence that names this one: the scale is in notes and the psychological present is in seconds. The psychological present is two to eight seconds, a tempo converts one to the other, and the level a listener reads then stops being a parameter at all — which is a prediction with teeth, because the same tune at two tempos should be phrased differently at levels the arithmetic names in advance.

form · Phrase
The tempo turns, and almost nothing moves. The earlier arrival reading — what is sounding at the final chord over what the listener has been hearing — swept over bar lengths from 0.5 to 5 seconds, which is 480 down to 48 beats a minute, at 3 closing lengths. Every curve is nearly flat. Across a tenfold change of tempo one gesture's reading moves by a factor of 1.201 and the other's by 1.098, while the gap between the two gestures — which is what that essay was measuring — is 1.228. The expectation was that the tempo would decide the answer, on the grounds that a two-second bar against a two-second release is a comparable pair. The premise is wrong in a way the sweep makes obvious: the thing being compared with the release is not a bar, it is the WHOLE ENDING, which is 2 to 8 bars long and is therefore far longer than the release at every tempo anybody plays. The running impression has caught up with the closing texture before the final chord arrives, at 0.5 seconds a bar and at 5, and what is left is the last bar's own jump.

The parameter that did not decide the answer

An earlier essay on closure ended by naming the tempo as the thing every number in it was resting on, and said it was the kind of parameter that had caused trouble before by turning out to decide the answer. Turned across a tenfold range at a closing gesture of fixed length it moves the reading by four per cent, against a twenty-three per cent gap between the gestures it is distinguishing. The parameter beside it in the same figure — how many bars the gesture occupies — moves it by twenty, and nobody had named that one at all.

form · Closure
How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.

A detector whose resolution the performance sets

The boundary detector lost its free parameter when the psychological present became a number of notes at a stated tempo, and what that held still was named at the time: a performance slows into a phrase end, so the number of notes inside the present is not the same everywhere in a tune — it falls exactly where a boundary is. Making the width follow the performance recovers half of what removing the parameter cost, and honestly leaves the other half.

form · Phrase
How long a bar of unequal beats can be, at a present of 3.5 seconds. Two quantities against the number of subdivisions in the bar. The bars are how many genuinely distinct unequal metres that length admits — groupings of twos and threes, up to rotation, discarding any that repeats a shorter grouping — and they run from one at 5 to 28 at 23. The line is how good a beat the best of those metres can manage once the whole bar is required to fit inside a psychological present of 3.5 seconds. It is flat at 0.865 up to a bar of 16 units, which is where the bar at the best subdivision first overruns the present, and falls after it: 17 at 0.830, 18 at 0.797, 19 at 0.765, 20 at 0.735. The supply of metres is still growing where the quality has begun to fall, so the lengths a tradition can use are a bounded prefix of an unbounded list.

How long a limping bar can be

The bound on an unequal beat turned out to be arithmetic, and the tempo window was left with only the tempo to decide. It decides nothing: every metre built from twos and threes gets the same answer, because the window is asked a yes-or-no question. Graded instead, an unequal beat costs 0.135 of the window's own preference at every bar length — and the constraint that does depend on length is the one nobody applied, that the whole bar has to fit inside the psychological present. At the subdivision that suits both beats best, a bar of sixteen units just fits and a bar of seventeen does not, which is where the supply of distinct metres has only started to grow.

rhythm · Additive metre
Where a bar of 25 and a bar of 9 first disagree. The additive metre 2+2+2+3+2+2+2+3+2+2+3 — 25 units, an onset at the head of every group — with the accents it predicts drawn above the accents predicted by reading it as a repeating bar of 9, which is the cut of it that agrees longest. The two rows are identical for 24 consecutive steps and differ for the first time at step 25, where the shorter reading expects an accent and the metre does not supply one. Nothing before that step distinguishes the two hypotheses, so a listener who has not heard 25 consecutive steps has no evidence either way — whatever they are disposed to hear.

A twenty-five is a nine until its last unit

Every account of long additive metres says they are heard as groups of shorter ones, and the metre-induction model had never been pointed at the claim. Pointed at it, the model does not prefer the group — it prefers the long bar outright, and would go on preferring it more the longer anybody listened. What it cannot do is start: the evidence that separates a bar of twenty-five from a bar of nine does not exist until the whole bar has been heard, and at the tempo an unequal metre is best played at the psychological present holds sixteen units.

rhythm · Additive metre
An accent moves the longest separable bar from 16 units to 18, and no further. For every bar length from 9 to 25 units, over all 1820 arrangements of twos and threes that are not a repeat of a shorter bar, the fewest and the most consecutive steps before the whole bar beats every shorter cut of it. 9: onsets alone 9 to 14, long beat predicted 7 to 12, downbeat predicted 7 to 8; 10: onsets alone 10 to 13, long beat predicted 8 to 11, downbeat predicted 8 to 9; 11: onsets alone 11 to 18, long beat predicted 9 to 16, downbeat predicted 9 to 10; 12: onsets alone 12 to 17, long beat predicted 10 to 15, downbeat predicted 10 to 11; 13: onsets alone 13 to 22, long beat predicted 11 to 20, downbeat predicted 11 to 12; 14: onsets alone 14 to 23, long beat predicted 12 to 21, downbeat predicted 12 to 13; 15: onsets alone 15 to 26, long beat predicted 13 to 24, downbeat predicted 13 to 14; 16: onsets alone 16 to 25, long beat predicted 14 to 23, downbeat predicted 14 to 15; 17: onsets alone 17 to 30, long beat predicted 15 to 28, downbeat predicted 15 to 16; 18: onsets alone 18 to 29, long beat predicted 16 to 27, downbeat predicted 16 to 17; 19: onsets alone 19 to 34, long beat predicted 17 to 32, downbeat predicted 17 to 18; 20: onsets alone 20 to 35, long beat predicted 18 to 33, downbeat predicted 18 to 19; 21: onsets alone 21 to 38, long beat predicted 19 to 36, downbeat predicted 19 to 20; 22: onsets alone 22 to 37, long beat predicted 20 to 35, downbeat predicted 20 to 21; 23: onsets alone 23 to 42, long beat predicted 21 to 40, downbeat predicted 21 to 22; 24: onsets alone 24 to 41, long beat predicted 22 to 39, downbeat predicted 22 to 23; 25: onsets alone 25 to 46, long beat predicted 23 to 44, downbeat predicted 23 to 24. A present of 3.5 seconds holds 16.0 steps at 218 milliseconds a step, so the longest bar some arrangement of which separates inside it is 16 units on onsets alone, 18 with the long beat predicted and 18 with the downbeat predicted.

The accent buys two units, however loud it is

A twenty-five cannot be told from a group of shorter bars on its onsets until more steps have gone by than a listener's present holds, and the obvious objection is that nobody plays an aksak bar as bare onsets: the long beat is louder, and the bar's first beat is marked. So how loud does an accent have to be? The question has a surprising answer. Loudness is not the variable. The existing accent cue changes nothing, and delays the answer where it changes anything. An accent that a reading has to predict works at any strength at all, and at no strength does more than a fixed amount: on the long beat it buys the two steps of a short beat, and on the downbeat it takes every arrangement to one floor — the bar less its last beat — which no cue carried by the notes can break. The longest bar that can be heard as one moves from sixteen units to eighteen.

rhythm · Additive metre
Six named proportions, as blurred as the durations that make them. Six proportions between two parts of a piece — 1 : 1, 4 : 3, 3 : 2, golden section, 2 : 1, 3 : 1 — placed on one axis by the logarithm of the ratio of the longer part to the shorter, and drawn as bars one criterion wide (d′ = 1) for a listener timing both parts with a Weber fraction of 7%, 15%, 35%. Bars that overlap are proportions that listener cannot tell apart. At 7%, 4 of 5 neighbouring pairs stay apart; at 15%, 3 of 5 neighbouring pairs stay apart; at 35%, 0 of 5 neighbouring pairs stay apart.

A proportion is only as fine as its two durations

Analyses of form measure proportions in bars and report them to three figures — a climax at 0.618, a section in the ratio 3 : 2. A listener has each part only as an estimate of how long it lasted, and a ratio of two estimates is blurred by both. Timed as well as anyone times a single second, eleven proportions fit between 1 : 1 and 3 : 1; timed from memory over minutes, two do. The golden section is told from 3 : 2 only below a Weber fraction of 5.4 per cent.

form · Proportion
A listener who knows every metre recognises none of them inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps from the downbeat before every other one of the 1820 arrangements of twos and threes has been contradicted by the stream, on onsets alone, with the long beats accented and with the downbeat accented, against the 16 steps a present of 3.5 seconds holds. onsets alone: recognised within the present for 0 of 1820; long-beat accent: recognised within the present for 0 of 1820; downbeat accent: recognised within the present for 85 of 1820. The dashed line is the present.

Knowing every metre is slower than knowing none

A long aksak bar cannot be told from its shorter cuts by induction before one step into its last beat, and no accent carried by the notes moves that floor. The obvious escape is a listener who knows the repertoire and recognises the metre instead. Recognition among all 1,820 arrangements of twos and threes never beats the floor, is never quicker than induction, and is slower for half the metres: a nine induced in 9 steps is recognised in 27. What breaks the floor is a small repertoire that leaves out the metre's own longest cut — with the cut known, no repertoire of any size does.

rhythm · Additive metre
Come in part-way with the downbeat accented, and no bar of sixteen units or more is recognised inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps a listener who knows every arrangement of twos and threes needs to recognise the metre and where its bar begins, with the downbeat accented: coming in at a sample of steps inside the bar, against hearing it from its written downbeat. On onsets alone, or with the long beats accented, a metre entered part-way is never told from its rotations. 9: from inside the bar 10 to 17, 16 of 20 inside the present; from the downbeat 10 to 10; 10: from inside the bar 11 to 19, 12 of 20 inside the present; from the downbeat 11 to 11; 11: from inside the bar 12 to 21, 9 of 18 inside the present; from the downbeat 12 to 12; 12: from inside the bar 13 to 23, 8 of 24 inside the present; from the downbeat 13 to 13; 13: from inside the bar 14 to 25, 8 of 28 inside the present; from the downbeat 14 to 14; 14: from inside the bar 15 to 27, 4 of 28 inside the present; from the downbeat 15 to 15; 15: from inside the bar 16 to 29, 4 of 32 inside the present; from the downbeat 16 to 16; 16: from inside the bar 17 to 31, 0 of 32 inside the present; from the downbeat 17 to 17; 17: from inside the bar 18 to 32, 0 of 36 inside the present; from the downbeat 18 to 18; 18: from inside the bar 19 to 32, 0 of 36 inside the present; from the downbeat 19 to 19; 19: from inside the bar 20 to 37, 0 of 40 inside the present; from the downbeat 20 to 20; 20: from inside the bar 21 to 35, 0 of 40 inside the present; from the downbeat 21 to 21; 21: from inside the bar 22 to 40, 0 of 44 inside the present; from the downbeat 22 to 22; 22: from inside the bar 23 to 39, 0 of 44 inside the present; from the downbeat 23 to 23; 23: from inside the bar 24 to 44, 0 of 48 inside the present; from the downbeat 23 to 24; 24: from inside the bar 23 to 39, 0 of 48 inside the present; from the downbeat 23 to 25; 25: from inside the bar 22 to 44, 0 of 52 inside the present; from the downbeat 23 to 25. In all, 61 of 590 entries are recognised within the 16 steps of a 3.5-second present.

A dancer who comes in late needs the downbeat marked

Every window for recognising an aksak metre so far started at its written downbeat. A dancer joining a dance already going has not heard the downbeat, and the arithmetic of that is blunt: a metre entered part-way is, onset for onset, each of its own rotations heard from their downbeats, and the rotations are metres too — 2+2+3 and 3+2+2 are counted differently. So on onsets, and with the long beats accented, no metre is ever told from its rotations. Only an accented downbeat tells them apart, and with it a listener who knows thirty metres recognises 54 per cent of them inside the present from a random entry, against 1 per cent without.

rhythm · Additive metre
Timing blurs a whole form evenly; counting sharpens it downward. A piece of 480 seconds divided 7 times, each level half the length of the one above, with how many proportions between 1 : 1 and 3 : 1 a listener can tell apart at each. Timed, the answer is 2.1 at the top and 5.2 at the bottom, a spread of 2.4 — because a timing judgement's Weber fraction is a step function of duration and almost every level of a piece falls in one step of it. Counted, in units of 2 seconds, the answer runs 2.2 to 10.3, a spread of 5.5. At no level does timing separate 3 : 2 from the golden section.

A form is sharp at the bottom and vague at the top

A movement is divided into sections, each into phrases, each into bars, and every level is a ratio of two estimates. Timed, the hierarchy is almost uniformly blunt — 2.1 distinguishable proportions at the top and 5.2 at the bottom, because a Weber fraction is a step function of duration and six of a piece's seven levels fall in one step of it. Counted, the same hierarchy runs from 2.2 to 12.2 and sharpens monotonically downward. At no level of either does timing separate 3 : 2 from the golden section.

form · Proportion
The breath is the looser ceiling nearly everywhere. How long a trained singer can hold a phrase on one breath, across a compass and at four dynamics, against the 8-second ceiling the psychological present puts on the same phrase. The flow through the folds rises with pitch and with loudness, so the breath ceiling falls both ways: at 60 decibels it runs 32.6 seconds at the bottom of the compass to 21.2 at the top; at 70 decibels it runs 23.1 seconds at the bottom of the compass to 15.0 at the top; at 80 decibels it runs 16.4 seconds at the bottom of the compass to 10.6 at the top; at 90 decibels it runs 11.6 seconds at the bottom of the compass to 7.5 at the top. The shaded line is the listener's ceiling and it does not move. The breath binds only where the two lines cross — 1 of the 40 cells drawn, all of them loud and high. So the constraint everybody names when asked why a phrase is the length it is, is almost never the constraint that decides it.

The ceiling everybody names is the loose one

Ask why phrases are the length they are and the answer given is the breath. It is arithmetic — usable lung volume over the air a note costs per second — and it comes out between fifteen and twenty-three seconds at a comfortable dynamic and between seven and twelve at a loud one. The ceiling the present moment imposes, the two-to-eight seconds inside which a stretch is heard as one thing rather than as a series, is two to three times tighter at almost every note and dynamic. A singer in an adagio is not running out of breath at the phrase end. They are running out of present.

form · Phrase
A listener who hears only the landmarks recognises the metre sooner. The share of metres recognised within the 16 steps of a 3.5-second present after coming in at a random step, against how many metres the listener knows, for five listeners: every onset with nothing marked, every onset with the long beats accented, only the downbeats and long beats with nothing marked, every onset with the downbeat accented, and only the downbeats and long beats with the downbeat marked. Every onset, nothing marked: 5 known, 49% within the present, 3% never; 10 known, 19% within the present, 8% never; 30 known, 1% within the present, 17% never; 100 known, 0% within the present, 34% never. Every onset, long beats accented: 5 known, 68% within the present, 3% never; 10 known, 39% within the present, 8% never; 30 known, 6% within the present, 17% never; 100 known, 0% within the present, 34% never. Landmarks only, nothing marked: 5 known, 69% within the present, 0% never; 10 known, 61% within the present, 3% never; 30 known, 26% within the present, 3% never; 100 known, 9% within the present, 9% never. Every onset, downbeat accented: 5 known, 88% within the present, 0% never; 10 known, 74% within the present, 0% never; 30 known, 54% within the present, 0% never; 100 known, 14% within the present, 0% never. Landmarks only, downbeat marked: 5 known, 91% within the present, 0% never; 10 known, 82% within the present, 0% never; 30 known, 56% within the present, 0% never; 100 known, 23% within the present, 0% never.

A late dancer needs the landmarks, not the rhythm

A listener who joins an additive-metre dance part-way recognises it far more reliably with the downbeat accented. Strip the stream down to its landmarks — the onsets that begin a bar or a long beat, with every other onset removed — and the listener does as well or better: knowing a hundred metres, 23 per cent are recognised within the present against 14 with every onset. Unmarked, the landmarks still beat every onset with the long beats accented. Neither half does it alone; what identifies a metre from a late entry is where its long beats sit relative to its bar.

rhythm · Additive metre

Named alongside it

The objects these essays reach for when they reach for this one.

PhraseTempoAdditive metreAksakInferenceMetreMetrical levelTactusDownbeatDurationGroupingClosure

All concepts