The collection

Every essay — page 19

Page 19 of 21, continuing through the fields in the same order.

Pitch and tuning Intervals and chords Scales and modes Harmony and voice leading Rhythm and metre Timbre and acoustics Perception and the listener Instruments and their design Form and structure Series Objects Sounds Search

Instruments and their design

Where the sound came from before any of the above. A stopped tube has only odd partials, a hammer at one seventh silences the seventh, and a bass string is wound because a plain one would be longer than the room.

Two registers 50 microseconds apart on one string. The partials of a 355-millimetre string plucked at 32 and 50 millimetres from the nut, drawn twice: once with the two quills releasing together, once with the far one releasing 0.05 milliseconds later, which is 0.026 of this string's period. A delayed release turns partial n through 2·pi·f_n·dt, so the rotation is proportional to the partial number: the fundamental is turned 9 degrees and partial 24 is turned 226. Simultaneous, the pair is missing partials 17 and 20 — holes it digs for itself where the two combs are equal and opposite. Staggered, it is missing none of them: a rotation of anything at all takes two amplitudes out of opposition. The partials neither comb can excite at all — 7, 11, 22 — are filled either way, because where one comb is zero the sum is the other one whatever its phase. The fundamental's gain over the far register alone falls from 4.36 decibels to 4.34.

The interval between two quills

Two jacks on one key are voiced separately and do not let go at the same instant. That interval turns each partial of the later pluck through an angle proportional to its number — so it leaves the fundamental alone and inverts the twentieth partial, and the holes the pair digs for itself vanish at a hundredth of a period. What a regulator can tolerate turns out to be one fixed fraction of a period at every pitch, which on a five-octave instrument is a factor of sixteen in milliseconds.

7 figures
Where a bar and its pipe stop being two things. The two normal modes of a bar and a resonator tuned to it, against how strongly they are coupled, at 262 hertz. Below a threshold the pair has one frequency and two different decay rates — the pale curves, which are the damping splitting rather than the pitch — and above it the frequencies separate. The threshold is exact and it is not a matter of degree: it is where the coupling rate equals half the difference between the two damping rates, which for a bar of Q 197 against a tube of Q 80 is a coupling of 0.37 per cent. A marimba's own coupling is 0.62 per cent — 1.67 times the threshold, and not free: it is fixed by how much louder the tube makes the note, since the coupling that splits the pair is the coupling that carries the energy out. So the resonator model's assumption that the tube is a filter downstream of the bar is wrong at middle C, and it is wrong by less than a factor of two.

A bar and its pipe are one object

Three earlier essays treat a marimba's resonator as a filter the bar's output passes through, and both of them said in their own caveats that the coupling was not modelled. It is here, and the debt was right: the coupling is 1.67 times the threshold at which the pair acquires two frequencies instead of two decay rates, so the tube is not downstream of anything. Every consequence of that is smaller than the peaks it would have to be seen between.

7 figures
Where a woodwind's A♭3 leaves it, below its corner and above it. The same fingering — 6 holes open on a 15-millimetre bore 567 millimetres long, sounding A♭3 at 207 hertz — drawn twice, with each opening's circle scaled by the share of the radiated power that leaves through it. At 400 hertz 78 per cent of it leaves through the first open hole, the bell takes 0 per cent, and the number of apertures really doing the radiating is 1.6; At 2600 hertz 6 per cent of it leaves through the first open hole, the bell takes 49 per cent, and the number of apertures really doing the radiating is 3.3. The lower frequency is below this fingering's corner and the higher one above it: below the corner the instrument is a short tube with one opening at the end of it, and above the corner it is the whole lattice at once. The power-weighted station — where a listener would say the sound is coming from — moves from 392 millimetres to 515.

Where a woodwind actually sounds from

Every number so far is read at the mouthpiece, and the corner's whole musical meaning is at the other end. Run the same solver forwards and it gives the flow leaving every hole — from which a clarinet's radiating aperture turns out to be a function of fingering and of frequency, but not the way it was predicted to: the fingering sets how far the aperture opens, almost exactly to the number of open holes, and barely moves the frequency at which it does.

7 figures
One doubling, held down a phrase. Where a single held arrangement of 4 players on 3 notes stands among the 36 at each chord of a 5-chord passage, best at the top, with what each chord would rather have named along the bottom. The held answer is flue pipe · trumpet · clarinet+violin, and it is the chord's own first choice at 4 of 5 of them. Holding it costs 16.2 per cent of the passage's roughness against re-scoring every chord — which is 2.4 per cent of the range the choice actually spans, since the arrangements at one chord differ by a factor of 7.8 on average. The cost is not spread over the passage: 1 chord carries nearly all of it.

An orchestrator doubles a line, not a chord

Three earlier essays made the objective a functional over a passage and a later one went back to holding one chord still. Put the doubling back into time and the retreat turns out to have been cheap: one arrangement held down a five-chord phrase is that phrase's own best answer at four of its five chords and costs 2.4 per cent of the range the choice spans — while the forward mask named earlier as the third temporal constant reaches for twenty milliseconds rather than two hundred, and cannot change the answer at any pace at all.

7 figures

Form and structure

The shape a piece has in time, computed rather than labelled. A distance matrix finds the sections before anybody names them, a phrase turns out to be a number of seconds rather than of bars, and an ending is five separate signals that can each arrive without the others.

thirty-two-bar AABA, every bar against every other bar. A self-similarity matrix of 32 bars of thirty-two-bar AABA. Each cell is the cosine similarity of two bars' pitch-class vectors, with the chord's own notes weighted 1 and the rest of the key 0.5. Similarity is quantised into four bands for drawing and anything under 0.35 is left as paper. Nothing in the computation knows what a section is; the blocks and stripes are what the arithmetic returns.

A piece is mostly itself again

Take a piece of music, encode each bar as the notes sounding in it, and compare every bar with every other bar. The picture that comes out has blocks and stripes in it, and those blocks and stripes are the form — arrived at by arithmetic that has never heard of an exposition, a chorus or a refrain.

8 figures
Boundaries found by a local operator, at three kernel widths. Foote's checkerboard novelty computed on the self-similarity matrix of thirty-two-bar AABA, at kernel widths of 2, 4, 8 bars. The dashed verticals are where the encoding's sections actually change; nothing about them enters the computation. A peak is a place where the bars before resemble each other, the bars after resemble each other, and the two groups do not resemble each other.

The boundary is where the neighbourhood changes

A section boundary can be found by an operator that never sees a section. It walks the diagonal of a similarity matrix asking one local question — do the bars behind me resemble each other, do the bars ahead resemble each other, and do the two groups resemble each other — and where the answer is yes, yes, no, there is an edge. What it cannot find turns out to say more than what it can.

8 figures
Redundancy in bits, and why the number needs a length beside it. Left: the LZ78 cost of each scheme divided by the cost of sending the same symbols flat, against how many bars are sent, with each bar coded as its chord and key. Every scheme is above 1 at a single chorus — the coder loses — and every one falls under it as the piece runs. Right: new dictionary phrases per bar at one chorus, which is the statistic that survives at short lengths.

How much of this is new

Repetition can be counted rather than looked at. Feed a piece's bars to a compressor and the bits it needs are a measure of how much of the piece is a repeat of an earlier part of itself. The measurement works, the number is real, and it turns out to be a statement about the description rather than about the music — which is the most useful thing it has to say.

8 figures
A phrase is a number of seconds, and the bars follow the tempo. Phrase durations for 1, 2, 4, 8, 16-bar phrases at seven tempos, on a logarithmic seconds axis, with the 2 to 8 second window shaded. The window is a property of the listener and does not move; which bar count falls inside it is decided entirely by the tempo.

A phrase is a number of seconds

Musical phrases are described in bars, and four is the number everybody names. But the constraint that fixes a phrase is a property of the listener and is measured in seconds, so the bar count is whatever the tempo makes it. Across seven ordinary tempos the bar count that lands inside the window moves by a factor of eight, while the window itself does not move at all.

6 figures
Two ways to fill eight bars, and only one of them accelerates. The period against the sentence, drawn as the lengths of their constituent units against position on a grid of eight bars. The ratio beside each row is the mean unit length in its second half divided by the mean in its first: the period at 1.00, the sentence at 0.67. A ratio below one is an acceleration — the unit shortening as the phrase approaches its arrival — and a ratio of one is a plan whose unit never changes length.

One of these eight-bar phrases accelerates

The sentence and the period both occupy eight bars, both end with a cadence, and both are recognised by ear rather than counted. What separates them is arithmetic. One halves its unit halfway through and the other does not, and the difference comes out as a single ratio — 0.67 against 1.00 — computed from nothing but the lengths of the parts.

5 figures
Five signals, computed separately, and no total. The five components of closure for 6 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.

What makes an ending an ending

Cadences are ranked. The authentic one is strong, the plagal weaker, the deceptive weaker still — and the solver here measured the quantity that ranking is usually explained by and found it says something else entirely. What survives is not a weaker version of the ranking but a different kind of object, with five components and no total.

5 figures
Five signals, computed separately, and no total. The five components of closure for 4 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.

An ending that exists so a bigger one can

Half of the cadences in tonal music are built to fail. A phrase that stopped convincingly at bar four would be a piece four bars long, so the ending at bar four is engineered to arrive and not to settle — and the components it withholds are exactly the ones its partner at bar eight supplies. Closure is nested, and the nesting is what turns two phrases into one thing.

6 figures
The same induction, one level up. Bar-level onsets from thirty-two-bar AABA — a bar is marked where a section or a key begins — scored against hypermetres of 2, 3, 4, 6, 8 bars with the identical function the beat-level figures use. The best-fitting period is 8 bars, which at 108 beats a minute lasts 17.8 seconds.

The bar above the bar

A four-bar group is a bar whose beats are bars. That is not an analogy — it is the same computation, and the same metre-induction model produces one when it is handed bars instead of beats, unchanged. What decides where the hierarchy of levels stops is not in the arithmetic at all, and it is a number the phrase essay already measured.

7 figures
The key plan is the shape. The key of a classical sonata-form movement against position in the movement, measured in steps along the chain of fifths from the home key. The positions are the proportions such a movement is described by rather than bar numbers from any one score. The furthest point is 4 steps out, sharing 3 of seven notes with home.

The key plan is the form

The large shape of a classical movement is not a shape at all, it is a journey — out to one key and back. Which key is not a matter of taste. Of the two keys that share six of their seven notes with home, only one introduces a note the home key does not use in any of its chords, and that note is the arriving key's own leading note. The departure is audible because of one accidental.

5 figures
A cycle has no ending to compute, so it uses density instead. 5 layers over 32 cycles of a 12-step pattern, with each layer's entry and exit marked, and the onsets per step summed underneath. No chord changes and no cadence occurs; the closure vector is zero throughout, and every change a listener hears is a change in how many things are playing.

A cycle cannot cadence

Every component of closure is defined by a first time and a last time. Music built on a repeating cycle has neither, so the whole apparatus returns zero on it — not a small value, zero, at every setting. What such music uses instead is how many things are playing, and that is a curve which can be computed from the onsets and nothing else.

7 figures
How long until it comes back. Mean similarity along each diagonal of the self-similarity matrix, minus the matrix's own mean off-diagonal similarity, against lag in bars, for 1 case. Lags run to half the length of each scheme, because a longer diagonal holds too few pairs to average. All rows share one vertical scale and the spread of each is printed beside it; the largest is 0.427 and the smallest 0.427. 1 of 1 cases with any spread at all put their strongest lag at the scheme's own repeat unit or a multiple of it.

How long until it comes back

A self-similarity matrix has a second reading that nobody looks for. Add up each diagonal instead of walking along one, and out falls repetition as a function of how long ago — a period, in bars, with no segmentation, no kernel width and no bar numbers anywhere in the answer. Five of the six schemes here report the length a listener would have named. The sixth reports something better.

8 figures
The period, as the piece goes by. The strongest lag of thirty-two-bar AABA computed on only the bars heard so far, against how many bars that is. The final answer is 4 bars; it is revised 6 times on the way, and is not reached for the last time until bar 29 of 32, which is 91 per cent of the way through and 64 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have.

The form a first hearing cannot have

Every figure so far was computed with the whole piece in hand. Run the same methods over only the bars already heard and one of the two methods survives intact — the boundary operator turns out to be causal at a fixed delay of a few bars — while the other collapses. The period of a piece is not knowable until the piece is nearly over, and in two of the six schemes here not until its last bar.

9 figures
Sensitivity and specificity on one dial. Aligned similarity — bar i against bar i+L, which is what a return is — for 4 eight-bar comparisons, as the key-invariance dial turns. One comparison is constructed: a literal repeat in the encoding, moved up a fifth, which is a stated manipulation because no scheme encoded here repeats a section in a new key. The shaded band is the margin between the two named comparisons, and it runs from 0.021 to 0.106.

The same thing somewhere else

A measure built on which notes are sounding calls a passage that comes back a fifth higher a stranger. There is a dial that fixes this, and turning it is supposed to be a trade — more sensitivity to a transposed return, less specificity against a coincidental one. It is not that trade. Two different statistics answer opposite ways, and the setting that would compromise between them is the worst one available.

6 figures
thirty-two-bar AABA, as a strip of time. thirty-two-bar AABA laid out one cell per bar, coloured by section, with the roman numeral in each bar. the A section's turnaround is the ii-V every variant keeps; the bridge is a chain of applied dominants. At 108 beats a minute in 4/4 the whole of it lasts 71 seconds. Cut into 4 repeat units of 8 bars, 3 pairs of units agree on more than 50 per cent of their bars. 2 of them are not identical, and 2 of those 2 differ in a run of bars ending at the last bar of the unit; the changed bars are marked in orange.

Where a repeat is changed

Cut every scheme into its own repeat unit, compare each unit with every other, and ask where a repeat stops agreeing with what it repeats. The answer is that it stops at the end, in every case the corpus contains — and the number of cases the corpus contains depends entirely on where the threshold for "a repeat" is put. Moving it by nothing at all takes the count from two to twenty and the finding with it.

6 figures
A tracker following a tempo that will not stay still. The tracker's error as a fraction of a beat, against beat number, for 3 rates of tempo change with a period-correction gain of 0.2. It settles at 0.025 of a beat behind at 0.5 per cent a beat, 0.092 of a beat behind at 2.0 per cent a beat, 0.206 of a beat behind at 5.0 per cent a beat. It does not lose the beat; it lags, by very nearly the rate divided by the period-correction gain, and the lag reaches a quarter of a beat at 6.3 per cent a beat — at which point the tracker is nearer the wrong onset than the right one.

A metre has to be able to change its mind

Replace the scoring function with a phase-corrected oscillator and the two failures reported earlier separate. The beat now survives two silent bars, drifting 23 milliseconds of a 560-millisecond beat. But it does not lose a moving tempo so much as lag behind it, by the rate over the correction gain — and a cadential ritardando that halves the tempo in eight beats is faster than the model can follow.

8 figures
Three ways to arrive at the same final tempo. Tempo against position in the closing passage, ending at 35 per cent of the opening tempo, for curvature exponents 1, 2, 3. All three begin and end at the same tempo, so what separates them is the middle: at the halfway point they read 68 per cent for linear in score position, 75 per cent for constant deceleration, 80 per cent for q = 3. The straight line is the one nobody plays. Measured ritardandos fit the decelerating curves, which is the whole of Kronman and Sundberg's argument: a closing gesture has the shape of a body stopping rather than of a dial being turned, and the parameter that varies between performances is the final tempo rather than the shape.

An ending is a deceleration

Every performance slows down at the end and the slowing has a shape. Tempo read against score position is the velocity of a body stopping — a square root rather than a straight line — and the three candidate curves agree at both ends by construction, so the whole audible difference is in the middle, where they part by fifteen per cent of the passage's length.

6 figures
The boundary operator run over only what has been heard. Foote's checkerboard novelty on thirty-two-bar AABA at a kernel width of 4 bars, computed twice: once with the whole piece available, and once using only the bars heard up to and including each bar. The kernel reaches 4 bars forward, so every cell it needs has been heard 3 bars after its centre — the retrospective curve replotted 3 bars to the right lands on the causal one, and the operator turns out to be causal at a fixed delay rather than blind. The dashed verticals are the encoding's real section boundaries and are not an input.

An ending that can be heard coming

Two measurements are both called hearing an ending coming and they point in opposite directions. By the halfway mark of an ordinary form almost nothing new arrives — and the cost of coding each bar has not fallen at all. Neither statistic says anything is about to stop, because no statistic over content can: predicting the next event well is not predicting that there will not be one.

7 figures
A silence measured in two ways that were not chosen to agree. How far a carried beat drifts during a silence, in fractions of a beat, for beat periods of 200 ms, 550 ms, 2000 ms and a tempo estimate 5 per cent wrong — which is the published discrimination limen rather than a figure chosen here. Half a beat of drift is where the metre coming out of the silence is no longer the one that went in, and it is reached after 10 beats whatever the tempo, which is 2.0 seconds at 200 ms, 5.5 seconds at 550 ms, 20.0 seconds at 2000 ms. The shaded band is the psychological present, 2 to 8 seconds, measured by a completely different literature and used in this collection to bound a phrase. The two answers overlap: a silence under about 2 seconds is a rest inside something and one over about 5.5 is after it.

A silence long enough to be an ending

Four of the five closure components are present or absent. Silence is the one with a continuous scale, so it is the one that can be given a threshold — and the threshold arrives from two literatures that were not chosen to agree, landing between three and a half and five and a half seconds. In a large hall, the room's own decay uses up most of it.

5 figures
How much of the rule a walk with no rule reproduces. Post-skip reversal in 20,000-note random walks with no melodic knowledge of any kind. An unbounded walk reverses after 50.0 per cent of leaps, which is the chance rate and is the check that the measurement is right. Confining it to 12 semitones raises that to 61.3 per cent. Reaching the 70 per cent that corpus studies report needs a central tendency of 0.95 — an almost deterministic pull back toward the middle at the edges of the range. A wall is not enough; there has to be a spring.

The leap that pays itself back

Every melody textbook teaches that a leap should be followed by a step in the opposite direction, and every corpus that has been counted agrees — around seven leaps in ten are answered that way. A random walk with two walls, no memory of the leap and no rule of any kind reverses after 61 per cent of them, and the residue is not a rule either. What is left when the walls are accounted for is a prediction the rule does not make, and it is the prediction that decides between them.

9 figures
The arch is not a preference. Every sequence of 6 notes over 8 scale degrees — 262,144 of them, enumerated rather than sampled — classified by contour, under three constraints. With none, the nine classes are spread. Requiring the sequence to return to its starting degree leaves only the arch, the valley and the flat, at 42.2 per cent each for the first two. Requiring it to begin and end on the LOWEST degree leaves the arch alone, at 99.6 per cent. Nothing here prefers a rise followed by a fall; the constraint is that the melody comes home, and a melody that comes home from below has nowhere to go but up first.

The shape that survives everything else

Throw away a melody's key, its tuning, its instrument and the sizes of its intervals, and what is left is a string of pluses and minuses. That string is what a listener who cannot name a note still has, and it costs 37 per cent of the tune to keep. The arch that melodic shape is famous for is not in it as a preference: enumerate every six-note sequence that begins and ends on the lowest degree it uses and 99.6 per cent of them are arches, because a melody that comes home from below has nowhere to go first but up.

9 figures