Theme

What the listener supplies

A downbeat, a key, a category, a preference — none of them is in the sound, and every one of them can be measured. These essays are about the part of music that is contributed by the person listening, which is also the part that varies between listeners and between traditions.
A 7-semitone sequence at 120 ms a tone. Tones drawn as pitch against time, one bar per tone. The events are the same in both readings of this pattern; what changes is whether a listener assigns them to one line that leaps back and forth or to two lines that each stay put. Nothing in the drawing decides which, and nothing in the sound does either. Perception and the listener

The ear builds objects, and sometimes offers a choice

What arrives at an ear is one pressure signal. What a listener gets is a set of separate things — a violin, a voice, a car outside. The assignment is a construction, and the clearest evidence is that it can be flipped by changing nothing but the speed: one sequence of tones is a single line when slow and two lines when fast, with a wide region in between where the listener may choose.

Where A has been. Documented pitch standards and surviving instruments, plotted as cents from A440. The extremes are 392 Hz and 465 Hz, which is 296 cents apart — 3.0 semitones, close enough to a minor third that a piece written at one and played at the other is in a different key. Nothing here is a preference; each is a decision somebody recorded. Pitch and tuning

A memory for the note itself, and it is dated

Absolute pitch is usually described as a rare perceptual gift. It is better described as a memory for a convention — and conventions have dates. Possessors trained on A=440 mis-name Baroque pitch by a semitone, their own labels drift sharp with age, and meanwhile most listeners without it start familiar songs within a semitone of the record.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval. Perception and the listener

The chord is still major, and that is why temperament works

A major third can be seventeen cents wrong and still be a major third. That tolerance is not a failure of hearing — it is the reason the whole subject of tuning is a discussion rather than a catastrophe. Every temperament ever proposed moves intervals around inside their categories, and the one thing none of them may do is push one across a boundary.

The probe-tone profile, major key. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap. Perception and the listener

Counting produced the hierarchy

Ask listeners how well each of the twelve notes fits after a passage in C major and the answers are not a smooth gradient. They fall into four groups with no overlap at all: the tonic, then the rest of the tonic triad, then the rest of the scale, then everything else — categories the subject had names for centuries before anybody ran the experiment.

Two models of consonance, and where they disagree. Every chord scored twice: horizontally by summed Plomp–Levelt roughness, vertically by the largest integer needed to write it as members of one harmonic series. Both are supposed to be measuring consonance and both are computed here from the chord itself. They correlate, but not tightly enough to be the same claim, and the chords furthest from the diagonal are the ones any experiment has to be run on. Perception and the listener

Consonance is half learned, and this is the half

This site's founding claim is that consonance is small whole numbers. Two computable models say so and they disagree about which chords — which is already awkward. The cross-cultural evidence is worse: listeners with little exposure to Western music discriminate roughness exactly as anyone does, match octaves exactly as anyone does, and rate consonant and dissonant chords as equally pleasant.

The metrical levels of 120 bpm, against the window. The range of inter-onset intervals that can be heard as a beat at all, from about 100 to 2000 milliseconds, with the preferred rate near 550. Each mark is one metrical level of a piece at 120 beats a minute. Which of them a listener taps is decided by which falls nearest the preferred rate, not by which one the notation calls the beat. Rhythm and metre

The beat has a preferred rate, and it is not the notation's

Metre is a hierarchy of levels and only one of them is tapped. Which one is decided by the listener's own clock rather than by the time signature: the window in which a series of events can be a beat at all runs from about a tenth of a second to two seconds, with a marked preference near half a second — so at the extremes of tempo the notated beat and the felt one part company, predictably.

A bell tuned to 294 Hz, and the note it is heard at. The partials of a well-tuned church bell, as ratios to the prime, with the three that imply the strike note marked. The nominal, twelfth and double octave sit at 2, 3 and 4, which is a harmonic series on 1 — so they imply a fundamental at 294 Hz, an octave below the loudest partial the bell has. The tierce at 1.2 is a MINOR third above the prime, which is why a bell has a minor quality by construction rather than by choice. Instruments and their design

A bell has no fundamental

The note a listener names when a church bell is struck is not any partial the bell has. Its nominal, twelfth and double octave sit at 2, 3 and 4 times the prime, which is a harmonic series on a pitch an octave below the loudest thing in the sound — and that pitch is supplied by the listener. Founders have been tuning it by ear since the fifteenth century.

What is struck, and where its partials land. Partial ratios for an ideal string, an ideal membrane, a kettledrum, drawn on a logarithmic axis so that a whole-number ratio is a fixed distance from the last. An ideal string: 1, 2, 3, 4, 5, 6. An ideal membrane: 1, 1.59, 2.14, 2.29, 2.65, 2.92. A kettledrum: 1, 1.50, 1.99, 2.44, 2.89. The membrane's are Bessel zeros — a derivation, and they land nowhere near the whole numbers. The kettledrum's have been pulled toward 2 : 3 : 4 : 5 by the enclosed air, on a fundamental an octave below the lowest partial present. Instruments and their design

What a drum is doing instead

An ideal membrane's modes are zeros of Bessel functions — 1, 1.59, 2.14, 2.30 — which support no common fundamental, so a drum rings without a note. A timpani is that problem solved — the kettle's air and the radiation load drag four modes onto 1, 1.5, 2, 2.5, and the pitch a timpanist tunes is the fundamental those four imply and none of them is.

Boundaries found by a local operator, at three kernel widths. Foote's checkerboard novelty computed on the self-similarity matrix of thirty-two-bar AABA, at kernel widths of 2, 4, 8 bars. The dashed verticals are where the encoding's sections actually change; nothing about them enters the computation. A peak is a place where the bars before resemble each other, the bars after resemble each other, and the two groups do not resemble each other. Form and structure

The boundary is where the neighbourhood changes

A section boundary can be found by an operator that never sees a section. It walks the diagonal of a similarity matrix asking one local question — do the bars behind me resemble each other, do the bars ahead resemble each other, and do the two groups resemble each other — and where the answer is yes, yes, no, there is an edge. What it cannot find turns out to say more than what it can.

A phrase is a number of seconds, and the bars follow the tempo. Phrase durations for 1, 2, 4, 8, 16-bar phrases at seven tempos, on a logarithmic seconds axis, with the 2 to 8 second window shaded. The window is a property of the listener and does not move; which bar count falls inside it is decided entirely by the tempo. Form and structure

A phrase is a number of seconds

Musical phrases are described in bars, and four is the number everybody names. But the constraint that fixes a phrase is a property of the listener and is measured in seconds, so the bar count is whatever the tempo makes it. Across seven ordinary tempos the bar count that lands inside the window moves by a factor of eight, while the window itself does not move at all.

Two ways to fill eight bars, and only one of them accelerates. The period against the sentence, drawn as the lengths of their constituent units against position on a grid of eight bars. The ratio beside each row is the mean unit length in its second half divided by the mean in its first: the period at 1.00, the sentence at 0.67. A ratio below one is an acceleration — the unit shortening as the phrase approaches its arrival — and a ratio of one is a plan whose unit never changes length. Form and structure

One of these eight-bar phrases accelerates

The sentence and the period both occupy eight bars, both end with a cadence, and both are recognised by ear rather than counted. What separates them is arithmetic. One halves its unit halfway through and the other does not, and the difference comes out as a single ratio — 0.67 against 1.00 — computed from nothing but the lengths of the parts.

Five signals, computed separately, and no total. The five components of closure for 6 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted. Form and structure

What makes an ending an ending

Cadences are ranked. The authentic one is strong, the plagal weaker, the deceptive weaker still — and the solver here measured the quantity that ranking is usually explained by and found it says something else entirely. What survives is not a weaker version of the ranking but a different kind of object, with five components and no total.

Five signals, computed separately, and no total. The five components of closure for 4 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted. Form and structure

An ending that exists so a bigger one can

Half of the cadences in tonal music are built to fail. A phrase that stopped convincingly at bar four would be a piece four bars long, so the ending at bar four is engineered to arrive and not to settle — and the components it withholds are exactly the ones its partner at bar eight supplies. Closure is nested, and the nesting is what turns two phrases into one thing.

The same induction, one level up. Bar-level onsets from thirty-two-bar AABA — a bar is marked where a section or a key begins — scored against hypermetres of 2, 3, 4, 6, 8 bars with the identical function the beat-level figures use. The best-fitting period is 8 bars, which at 108 beats a minute lasts 17.8 seconds. Form and structure

The bar above the bar

A four-bar group is a bar whose beats are bars. That is not an analogy — it is the same computation, and the same metre-induction model produces one when it is handed bars instead of beats, unchanged. What decides where the hierarchy of levels stops is not in the arithmetic at all, and it is a number the phrase essay already measured.

How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; I – IV – V – vi totals 7.3 bits over 3 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire. Harmony and voice leading

The chord that did not come

A deceptive cadence is described as a surprise, and the explanation offered is that the wrong chord arrived. Measured against the tonal hierarchy already in use, the wrong chord is the second best-fitting triad in the key — and two of its three voices do exactly what they would have done in the right one. The surprise is not statistical. It is one voice, and it is the bass.

How often the chord changes, and what a room allows. Chord changes a second implied by each style's stated rate and tempo, on a logarithmic axis, with the rate above which a room leaves more than one earlier chord above 20 dB marked for six rooms. The style rates are conventions rather than corpus measurements and the figure says so; the room rates are arithmetic from the reverberation time. Harmony and voice leading

How often the chord changes

Two pieces can use the same chords in the same order and be nothing alike, because a progression says which chords and not how fast. Harmonic rhythm is the second variable, it runs over a factor of thirty between the styles that use it, and both of its limits are set by things that are not harmony — a listener's memory at the slow end and a building at the fast one.

A cycle has no ending to compute, so it uses density instead. 5 layers over 32 cycles of a 12-step pattern, with each layer's entry and exit marked, and the onsets per step summed underneath. No chord changes and no cadence occurs; the closure vector is zero throughout, and every change a listener hears is a change in how many things are playing. Form and structure

A cycle cannot cadence

Every component of closure is defined by a first time and a last time. Music built on a repeating cycle has neither, so the whole apparatus returns zero on it — not a small value, zero, at every setting. What such music uses instead is how many things are playing, and that is a curve which can be computed from the onsets and nothing else.

The 7-note sets in which every interval occurs a different number of times. Each set of 7 notes containing C whose six interval counts are all different, with the counts printed. Every one of them is a rotation of one of two shapes, and only one of the two has steps a scale could use — the other is a run of semitones with the gap at the end. Scales and modes

Every interval a different number of times

Count the intervals inside a major scale and the six answers are 2, 5, 4, 3, 6 and 1 — six different numbers, no two alike. That is not decoration. It means the number of notes a key shares with a transposition of itself identifies the distance uniquely, so a listener who can only count common tones can still tell exactly how far a modulation went.

The period, as the piece goes by. The strongest lag of thirty-two-bar AABA computed on only the bars heard so far, against how many bars that is. The final answer is 4 bars; it is revised 6 times on the way, and is not reached for the last time until bar 29 of 32, which is 91 per cent of the way through and 64 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have. Form and structure

The form a first hearing cannot have

Every figure so far was computed with the whole piece in hand. Run the same methods over only the bars already heard and one of the two methods survives intact — the boundary operator turns out to be causal at a fixed delay of a few bars — while the other collapses. The period of a piece is not knowable until the piece is nearly over, and in two of the six schemes here not until its last bar.

The ranking is settled either side of one narrow band. Remembered repetition — each bar's best match to an earlier bar, discounted by exp(−Δt/τ) with Δt in seconds — for 6 schemes at 108 beats a minute, against the decay constant τ on a logarithmic axis. The order of the schemes changes only between 8 and 13 seconds; outside that band it is fixed, so an estimate of τ wrong by any amount that stays outside it leaves the ranking alone. Perception and the listener

A return has to be remembered

A stripe four bars off the diagonal and a stripe twenty-four bars off it are the same ink and are not the same experience. Convert the lag axis to seconds, discount every comparison by how long ago it was, and the ranking of these six schemes by how repetitive they are changes — and the decay constant and the tempo turn out to enter the arithmetic as one number rather than two.

What 2 partials imply, and how many answers there are. The partials are at 32.7, 49.0 Hz. Each row is a harmonic series they are consistent with to within 30 cents: the filled dots are the partials in their assigned slots and the open dots are slots the series predicts that nothing occupies. There are 5 such series with harmonic numbers up to 16, the best fitting them to 1.4 cents with a fundamental of 16.34 Hz and no unoccupied slot. The rest sit at one half, one third, one quarter, one fifth of it and their arithmetic is exactly as exact; what separates them is the count of empty slots, which is why a residue pitch built from few partials is reported an octave out by a minority of listeners and not by the rest. Instruments and their design

Two pipes for a note neither makes

The resultant stop sounds a sixteen-foot pipe and a ten-and-two-thirds-foot pipe and asks the listener for a thirty-two-foot note. It is the missing fundamental built on purpose from the fewest partials that can imply anything — and counting how many fundamentals two partials actually imply is the arithmetic behind three centuries of builders disagreeing about whether it works.

The most a 6 cm cone can make of a low note. The maximum sound pressure level at one metre from a circular radiator of effective radius 3.2 cm, moving 1.5 mm at its limit, in a system resonating at 250 Hz. Below resonance the cone is already at that limit and the pressure a piston makes goes as the square of frequency, so the curve falls at twelve decibels an octave: 66 dB at 40 Hz, 78 dB at 80 Hz, 94 dB at 200 Hz. The 40 Hz figure is 28 decibels below the 200 Hz one, and that gap is arithmetic about a radius and a displacement rather than a property of any particular loudspeaker. The dots are the harmonics of a 41.2 Hz note with a one-over-n source spectrum; the loudest of them is the 6th. Instruments and their design

The bass a small loudspeaker does not make

A three-inch cone at its excursion limit produces sixty-six decibels at forty hertz, which the ear converts to seventeen phons — barely above nothing. The note is heard anyway, because its harmonics are radiated and its fundamental is supplied by the listener. Computing what arrives turns the residue from a curiosity into a design decision, and finds that the fundamental of a low note is not the loudest part of it on any system a listener is likely to own.

The thirteen intervals at C4, ordered by roughness, against Fux's species, 1725. Every interval within an octave above 262 Hz, ordered by the Plomp–Levelt roughness of a string spectrum — from 0.0029 for the unison to 0.2635 for the minor second — beside whether Fux's species, 1725 counts it a consonance and how many of the thirteen may follow it under that rule set. The best single threshold on roughness misclassifies 2 of the thirteen, among them the major third and the minor third, and the gap it falls in is 4.2 per cent of the roughness either side of it. Harmony and voice leading

A dissonance is what has to be resolved

The perfect fourth is a consonance between two upper voices and a dissonance against the bass, and the same three notes are involved either way. Score both arrangements with the roughness model and the one the rules call a dissonance comes out thirty per cent smoother. Whatever the rule is tracking, it is not the sound.

One pattern, four metres. The same 6-step onset pattern read under 3 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are the 3 is the beat 8, the 2 is the beat 4, every step is a beat 8, so the 3 is the beat and every step is a beat tie and the model does not choose. Nothing about the sound differs between these readings; the bar line is supplied by the listener. Rhythm and metre

Which of the two is the beat

A polyrhythm is notated as a bar of two with three laid across it. Run the metre rules used here over the composite and they choose the three — at every ratio tried, without exception. Add the one other rule the rules have, and the answer changes hands at a bar of 1,013 milliseconds, which is 118 to the minute.

Syncopation against a bar of 8. An 8-step pattern with 3 onsets, against the metrical weights of its bar. A position's weight is zero on the downbeat and one lower at each level down the subdivision tree, drawn here as the depth of the bar hanging beneath it. A note on a weak position followed by a rest on a stronger one costs the difference. the tresillo, in a bar of eight scores 2 — the note at step 4 against the rest at step 5, costing 2. Rhythm and metre

Syncopation is a number about the metre

Longuet-Higgins and Lee price a syncopation at the metrical weight a note skips over. That makes it computable, and it makes it a property of a pair rather than of a rhythm — the son clave scores 4 read from step one, 2 from step three and 8 from step four, and not one onset has moved. Worse, the induction rules pick very nearly the reading that scores lowest.

One pattern, four metres. The same 16-step onset pattern read under 4 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are downbeat on step 1 0, downbeat on step 2 -12, downbeat on step 3 -12, downbeat on step 4 0, so downbeat on step 1 and downbeat on step 4 tie and the model does not choose. Nothing about the sound differs between these readings; the bar line is supplied by the listener. Perception and the listener

The beat that is never sounded

A listener who has heard four bars of a groove and then hears two bars with the downbeats taken out does not move the downbeat. This site's rule set does, every time, on every pattern tried — and the direction it moves in says exactly what kind of model would be needed instead.

The seven modes, brightest first. The same seven pitch classes started on each of its degrees in turn, ordered by how many of their notes are raised. Each row differs from the one below it by exactly one note, and that note moves down one semitone each time. Scales and modes

Nothing in the census knows which note is home

All four properties six earlier essays are about are invariant under rotation and transposition — one property tuple over all eighty-four rotations and transpositions of the diatonic set. So the census that separates 349 shapes cannot separate a major scale from its own Aeolian mode, and everything that makes one note a tonic is outside it.

The price of a tonic. Every note of Dorian is given the same duration except its tonic, which is lengthened; the horizontal axis is the share of the total that goes to it. The key-finder answers with the parent key until 25.0 per cent of the time is spent on the modal tonic, and with D minor above it. At the left-hand edge every note has equal weight, which is the pitch-class set itself — and with every weight identical the correlation is not merely low but undefined, because a flat histogram has no variance to correlate with anything. Perception and the listener

What a tonic costs in seconds

The standard key-finding algorithm cannot be run on a pitch-class set at all — a flat histogram has no variance and the correlation is undefined. Give it durations and it answers with the parent key for all seven modes identically, and it takes between 15.8 and 30.0 per cent of the total time spent on one note before it names that note instead.

How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is. Perception and the listener

How much evidence a modulation needs

Run a key-finder bar by bar over a progression that moves to the dominant at bar six. With four bars of history the answer becomes the new key at bar seven and holds. With three bars it never gets there at all, and reports E minor and B minor on the way. The window decides the lag as much as the music does.

Long-term average spectra: an orchestra, playing forte against a trained operatic soloist. Each source's mean spectrum over a long passage, in decibels below its own strongest region, on a logarithmic frequency axis. An orchestra, playing forte peaks at 250 Hz and is 30 dB down by 3,150 Hz; a trained operatic soloist peaks at 250 Hz and is 11 dB down by 3,150 Hz. The shapes are the same until about 1 kHz and separate above it: at 3153 Hz the difference is 19.0 decibels, which is the largest anywhere in the range. Nothing here is about level. Both curves are drawn against their own peaks, so what is being compared is shape. Timbre and acoustics

One voice over ninety players

A soloist heard over a full orchestra is not louder than it and could not be. What the trained voice does instead is put a peak of energy at three kilohertz, which is where the orchestra's spectrum has already fallen away and where the ear's own threshold happens to be lowest. Nineteen decibels of advantage, in a place nobody is competing for.

a major triad, C–E–G. The partials are at 261.6, 329.6, 392.0 Hz. Each row is a harmonic series they are consistent with to within 30 cents: the filled dots are the partials in their assigned slots and the open dots are slots the series predicts that nothing occupies. There are 2 such series with harmonic numbers up to 16, the best fitting them to 10.1 cents with a fundamental of 65.54 Hz and no unoccupied slot. The rest sit at one half of it and their arithmetic is exactly as exact; what separates them is the count of empty slots, which is why a residue pitch built from few partials is reported an octave out by a minority of listeners and not by the rest. Intervals and chords

The root an ear supplies

A major triad's notes fit 4:5:6 with nothing missing and a fundamental two octaves below the bass. A minor triad's fit two different series with two different answers a major sixth apart, and the model cannot choose between them. The ambiguity theorists argued about for two centuries is a computable quantity, and the spectrum invented to remove it is one no object produces.

How much of the rule a walk with no rule reproduces. Post-skip reversal in 20,000-note random walks with no melodic knowledge of any kind. An unbounded walk reverses after 50.0 per cent of leaps, which is the chance rate and is the check that the measurement is right. Confining it to 12 semitones raises that to 61.3 per cent. Reaching the 70 per cent that corpus studies report needs a central tendency of 0.95 — an almost deterministic pull back toward the middle at the edges of the range. A wall is not enough; there has to be a spring. Form and structure

The leap that pays itself back

Every melody textbook teaches that a leap should be followed by a step in the opposite direction, and every corpus that has been counted agrees — around seven leaps in ten are answered that way. A random walk with two walls, no memory of the leap and no rule of any kind reverses after 61 per cent of them, and the residue is not a rule either. What is left when the walls are accounted for is a prediction the rule does not make, and it is the prediction that decides between them.

The arch is not a preference. Every sequence of 6 notes over 8 scale degrees — 262,144 of them, enumerated rather than sampled — classified by contour, under three constraints. With none, the nine classes are spread. Requiring the sequence to return to its starting degree leaves only the arch, the valley and the flat, at 42.2 per cent each for the first two. Requiring it to begin and end on the LOWEST degree leaves the arch alone, at 99.6 per cent. Nothing here prefers a rise followed by a fall; the constraint is that the melody comes home, and a melody that comes home from below has nowhere to go but up first. Form and structure

The shape that survives everything else

Throw away a melody's key, its tuning, its instrument and the sizes of its intervals, and what is left is a string of pluses and minuses. That string is what a listener who cannot name a note still has, and it costs 37 per cent of the tune to keep. The arch that melodic shape is famous for is not in it as a preference: enumerate every six-note sequence that begins and ends on the lowest degree it uses and 99.6 per cent of them are arches, because a melody that comes home from below has nowhere to go first but up.

The vowel in "hod", sung at 110 Hz. The partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung. Timbre and acoustics

The sound a listener knows best

A voice is recognisable across every vowel it says, across two octaves of pitch, down a bad telephone line and in a whisper where there is no pitch at all. Nothing that survives all of that can be a frequency. What survives is a ratio: the resonances of a vocal tract are set by its length, so a shorter tract multiplies every formant by the same factor, and identity is a scale on the spectral envelope rather than a position within it. Between an adult man and a child the whole pattern moves by a fifth, and the vowel does not change at all.

The same eight notes, read three ways. The scale, eight quavers, scored against every triad and seventh at every root. Barred as written the best reading is C major7 at 0.850; with the barline one quaver later it is D minor7 at 0.850. With no metre — every note weighted the same — 4 readings tie at 0.625 and the passage has no best analysis at all. The notes are identical in all three. What changed is where the bar starts, which is not a fact about harmony. Harmony and voice leading

Which notes are the chord

A progression is a list of chords, and before there is a list something has to decide which of the notes sounding are chord tones and which are passing. Take the eight notes of a scale as eight quavers and score every triad and seventh at every root: barred as written the best reading is C major seventh, with the barline moved by one quaver it is D minor seventh, and with no metre at all three readings tie exactly and the passage has no best analysis. Same eight notes in all three. Harmonic analysis is a function of a variable that is not harmony.

One pattern, four metres. The same 12-step onset pattern read under 3 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are 3/4 -4, 6/8 12, 2/4 -4, so 6/8 wins. The page prints 3/4, which the model does not prefer — it ranks 6/8 above it. Nothing about the sound differs between these readings; the bar line is supplied by the listener. Rhythm and metre

The time signature is a claim

A bar line is not a measurement. It is a claim about where the accents are, made before the sound exists, and there is a metre-induction model here that can be handed the same onsets and asked whether it agrees. On a hemiola it does not: the page says three and the model says six, by a margin of twelve against minus four. And on a bar of seven the signature is not even a candidate — what decides the reading is the beaming, which the signature does not contain.

What a contour costs to remember. A melody of n notes over 8 degrees carries 3 bits a note. Its contour carries fewer, and fewer than the number of distinct contours suggests, because the contours are not equally likely: at 6 notes there are 243 of them but the entropy is 6.59 bits, an effective alphabet of 96. Each further note adds 1.28 bits of contour against three of melody, so the shape keeps a stable 37 per cent of what is there however long the tune. Perception and the listener

The part of the tune that is kept

Contour survives transposition, retuning, a change of instrument and a doubling of every interval, and the usual explanation is that it is what a listener retains. That can be counted rather than assumed. A six-note melody over eight degrees carries eighteen bits; its contour carries 6.59 — not the 7.92 the number of distinct shapes suggests, because the shapes are wildly unequal — and the effective alphabet is ninety-six out of two hundred and forty-three. Each further note adds 1.28 bits of shape against three of melody, and at about nine notes a contour is specific enough to pick one tune out of a thousand.

The pitch moves and the repetition rate does not. Three partials around harmonic 10 of 200 hertz, shifted together by up to 200 hertz, with three curves. The flat line is the envelope repetition rate, which the shift cannot move at all. The rising line is the shift divided by the harmonic number — 200 hertz becoming 220.0 — which is what a harmonic template predicts and is what listeners report. The third curve is the next-best template, which overtakes the first partway along: the pitch is ambiguous, and it drops back rather than rising indefinitely. Perception and the listener

The pitch that moves the wrong distance

Take three partials two hundred hertz apart and move every one of them up by forty. The spacing has not changed, so anything reading the pitch off how often the waveform repeats must give the same answer as before. The pitch moves to 204 — the shift divided by the harmonic number — which is what a harmonic template predicts and what listeners report. Push the shift to a hundred and a second reading overtakes the first, so there are two pitches and neither is the spacing. This is the measurement that closes the question, and it closes it by ruling out one mechanism rather than by choosing between the two that are left.

The mode is its tritone. Every set of degrees that lies inside at least one of the seven modes — 510 of them — scored by how many modes it leaves standing, with the 15 minimal sets that decide printed in full. All 15 have two members and every one of them either is a tritone or contains the tritone above the tonic. The only degree every mode has is the tonic itself; there is no degree belonging to just one mode. Averaged over every order the degrees could arrive in, 5.67 of the seven have to have been heard before the mode is settled — most of the scale, because the deciding pair arrives when it arrives. Scales and modes

The one note that decides the mode

A key signature names the seven and not the rotation, so the mode has to be heard. Enumerate every set of degrees that lies inside at least one of the seven modes — 510 of them — and count how many modes each leaves standing. Fifteen minimal sets decide, every one has two members, and every one is the mode's own tritone. The only degree all seven share is the tonic; no degree belongs to a single mode; and averaged over the orders the degrees could arrive in, 5.67 of the seven have to have been heard first.

Where the page ends a phrase, and where the ear does. Twinkle, twinkle with two sets of phrase boundaries on it. The lower curve is a local boundary detector — a peak in how much the interval and the note length change from one to the next, with nothing in it about bar lines or harmony — and the marks above it are where the notation puts the phrase ends. It finds 100 per cent of them and 2 boundaries the page does not have. Where the two agree it is because a long note is sitting at the join; where they disagree the page is marking a grammatical unit and the detector is finding a perceptual one. Form and structure

Where a phrase ends

Run a boundary detector over the three tunes used throughout and it agrees with the notated phrasing on one of them perfectly and on another almost not at all. The reason is which cue each tune uses: Twinkle's phrases all end on a long note, so a duration-weighted detector finds five of five with no false alarms; Ode to Joy's run on in crotchets and its phrasing is in the intervals, where a duration detector finds one of three and a pitch detector finds all three and eight others. No fixed weighting serves both, and the published one is worse on each tune than the single cue that tune uses.

note length against the onsets. Every candidate metre's fit to a 16-step pattern with 4 onsets, plotted against how strongly note length is weighted. At a gain of zero the scoring is the onset-only one every earlier model used, and the winner is step 1 and step 4. At a gain of 0.05 the answer becomes step 4. 2 candidates are exactly flat — step 2 and step 3 have no onset on any strong position, so there is no credit for the cue to multiply and no weighting of it can move the line. Form and structure

What the onsets left out

Eight essays induce a metre from a list of ones and zeros, and every failure they recorded was argued about as a failure of the rules. Two of the three are not. Note length is already in that list and the scoring throws it away: put it back and the son clave's two-way tie resolves to the notated downbeat. But the groove with its beat removed cannot be repaired by any cue at any strength, and the reason is arithmetic rather than empirical — the true phase has no onset on any of its strong positions, so there is no credit for a cue to multiply and its line is exactly flat.

A model that cannot say two keys does not say it is unsure. I – IV – V – I played in C major and in a second major key at the same time, with the second key moved round the circle of fifths. For each separation: the correlation the standard key-finder gives its single best answer, and the correlation reached by the best PAIR of key profiles — a hypothesis the finder does not have. The pair recovers both keys that are sounding at every separation, 7 of 7. The single answer names neither of them at 4 of the 7, and its confidence does not fall when it is wrong: at four steps apart it reports E minor at r = 0.886, against 0.959 for the same progression in one key. Harmony and voice leading

Two keys at once

Every key figure so far assumes one key is sounding, and the standard key-finder has no value it can return that means two. Play one progression in C and the same progression a major third away at the same time, and the model does not report uncertainty: it reports E minor, at a correlation of 0.886, against 0.959 for the same progression in one key. It is as confident as it ever is, and neither key it names is being played. Give it the missing hypothesis — pairs of key profiles rather than single ones — and it recovers both keys at every separation, all seven of seven.

Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 2.8 cents at ten to one and 3.6 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it. Perception and the listener

The boundary that barely moves

Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.

How fast two keys can alternate before the finder stops following. The share of bars a moving key-finder names correctly, once its reading is shifted back by its own lag, against how many bars each key holds for. One line per window. Below a block of three bars the second key is never named at all — 2 of the sweep's readings report a single key for the whole passage — and above about twice the window the tracking is over ninety per cent. The lag itself is about half the window: 0 bars at a window of 3, 0 bars at a window of 4, 3 bars at a window of 8. Harmony and voice leading

The alternation a key-finder cannot follow

Two keys sounding together are not in the key-finder's vocabulary, and an earlier essay ended by pointing at the other case and saying what was missing: a passage whose alternation rate can be varied while everything else is held still. Built, it gives a rule with three numbers in it — the second key is never named below a block of three bars, the tracking clears ninety per cent above twice the window, and the reading is late by half the window throughout.

What the joint search changes, and what it never changes. Over 552 constructed passages of eight slots with rests, how often the joint reading differs from the pipeline's. The chord differs in 29 per cent and the barline in 31, with both differing in 20. The key differs in 0 per cent — never — because the key is read from a pitch-class histogram, which does not know where the bar starts or which notes are chord tones. Two of the three decisions are entangled and the third is not. Harmony and voice leading

Three decisions that constrain each other

Every model here decides one thing at a time — the key from the pitch classes, the metre from the onsets, the chords from the metre — and an earlier essay ended by saying a listener does all three at once. Resolving them jointly costs a hundred and fifty-seven times the search and changes the reading of two passages in five. It never once changes the key, and the reason it cannot is the reason the whole account is built the way it is.

Which note a chord would rather have twice. Every complete four-part voicing of each chord inside the SATB ranges, grouped by which member sounds twice and scored for roughness — 480 voicings for a triad. The order for a major triad is root < fifth < third, which is the rule every part-writing treatise states. For a minor triad it is fifth < root < third, which is not. The numbers printed under each bar are the mean error, in cents, with which the four sounding notes fit a single harmonic series, and that measure separates the three far more sharply than roughness does. Intervals and chords

The note that sounds twice

A triad has three notes and a four-part texture has four voices, so one note is doubled — and the voicing model used here leaves the choice free because the rules have an opinion about it. Asked properly, the arithmetic agrees with the treatises for the first time in nine essays: root, then fifth, then third. For a minor triad it does not agree, and for a symmetric chord it correctly has nothing to say.

One envelope, and the three places a listener might be said to hear it. The amplitude envelope of a note with a 90 millisecond exponential attack, with the three criteria the literature offers drawn across it. The heard moment is 8.0 ms at the detection criterion, 28 ms at the perceptual-onset criterion and 94 ms at the perceptual-attack criterion. The physical onset is at zero on this axis and no criterion puts the heard moment there. The buttons play this attack against a two-millisecond one, started at the same instant. Rhythm and metre

A note is heard after it starts

Every rhythm essay until now has treated a note's onset as the moment it happens. It is not: the instant a listener aligns a note with a beat is later than its physical start by an amount the note's own attack decides, and for a sung or bowed note that amount is about thirty milliseconds — the size of the whole quantity six essays on microtiming set out to measure.

How much correlation it would take to matter. The limen of a 7-semitone interval at a note length of 0.25 seconds, against the correlation between the two notes' errors. The independent model at the left gives 9.44 cents. Halving that needs a correlation of 0.75; a fifth off it needs 0.31. The curve is √(1 − ρ) and nothing else, so the correlation required for a stated improvement is arithmetic — which turns the question from “does a key help?” into “by how much, and here is the number it must reach”. Intervals and chords

How much an anchor would have to be worth

Two pitch errors added in quadrature assume an independence nobody measured — a listener inside a key hears a note as a scale degree, and a shared reference is exactly a correlated error. Turning the dial is not evidence. What is evidence is that the dial is not free: a shared error cancels out of a difference completely, so a listener's single-note limen and their interval limen give the two components with nothing left over, and halving the interval limen needs a correlation of exactly 0.75.

Every result so far, against the listener's own noise. Three findings drawn against the one parameter all of them assume: how finely the listener resolves a pitch. At 11 cents — a trained listener, and the value every earlier essay used — the octave holds 6 nameable categories, twelve equal ones are named right 91 per cent of the time, and a 20-to-one expectation moves a boundary by 3.6 cents. At 35 cents it is 2 categories, 72 per cent, and 37 cents. The capacity falls roughly as one over sigma and the shift rises as its square, so the three curves separate rather than moving together. Scales and modes

The listener the model was never run for

Five earlier essays rest on one number — how finely a listener resolves a pitch — and every one of them used a trained listener's eleven cents. The model's dependence on it is not gentle: the capacity goes as its reciprocal and the expectation shift as its square, so an untrained listener at thirty-five cents has two nameable categories per octave rather than six, and a foreign tuning system is not mis-transcribed by them but absorbed.

How much of each spectrum a listener can assemble into one note. Each partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there. Perception and the listener

The spectrum that will not fuse

A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.

Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 3 parts to 8 adds 4.3 decibels of power and 0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 6.0 to 3.9 as the parts crowd into the same register. Form and structure

The dynamics are in the score already

Count the parts in each bar, realise them in their ranges, put every partial in its critical band, sum the loudnesses and run the result through the two smoothers built earlier. What comes out is a dynamic curve for a piece with no performance in it anywhere — and it says that doubling the number of parts inside a fixed register adds three decibels of power and about one of loudness, because the extra parts land in bands that were already occupied. Let the register widen with the parts and the same arithmetic gives eight phon, which is what a tutti actually is.

Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness. Form and structure

A final chord is not made loud by adding to it

An earlier essay on closure said the loudest cue an ending has needs a corpus rather than an arithmetic. The arithmetic was built one essay ago, so it does not. Run four ending textures through it and two things come out backwards: a final chord three parts thicker than the rest arrives *quieter* against the running impression than the passage it ends, and a texture that drops a part a bar does not get quieter at all until the bar where there is one part left.

What the notes in between do to the anchor the interval is measured against. How finely a 7-semitone interval can be judged when its two notes are separated by other notes rather than by silence, under the two published accounts. Confirming material restates the key and refreshes the shared reference, so the correlation climbs from 0.5 toward a ceiling and the limen falls to 4.81 cents. Overwriting material competes for the same memory, so the correlation decays to 0.04 and the limen rises to 9.28. By 8 notes the two accounts differ by 4.5 cents, which is 47 per cent of the limen with no anchor at all — and no experiment here distinguishes them. Intervals and chords

The notes in between

Every figure until now is about two notes with nothing between them, and a melody is notes with other notes between them. Two published accounts of what the intervening material does predict opposite signs — one says the key is restated and the shared reference is refreshed, the other says each note competes for the same memory and it decays. By eight notes they differ by four and a half cents, which is nearly half the limen the interval would have with no anchor at all.

The one number the ordered key-finder was tuned on. For each rate of alternation between two keys, the cost of changing key at which the model stops hearing two keys and starts hearing borrowed chords in one. The threshold rises with the period — 0.95 at 1 bar, 0.95 at 2 bars, 0.95 at 4 bars, 2.00 at 8 bars, 3.50 at 16 bars — so the parameter and the rate trade off against each other exactly. The value tuned earlier, 2.2, sits above every threshold on this axis, which means its verdict about fast alternation was a consequence of the tuning rather than a finding about the music. Filled means the model names two keys; hollow means it names one and calls the rest borrowings. Scales and modes

A modulation and a borrowing are one number apart

The key-finder that keeps the order has one tuned parameter, and said so. Sweep it and the parameter turns out to be the whole verdict: below a threshold the model hears two keys alternating, above it one key with borrowed chords. The threshold rises with how slowly the keys alternate — and the value it chose sits above every threshold in range, so its finding about fast alternation was a consequence of the tuning.

Two kinds of evidence, each measured against its own chance. Every key as a point: across, how many standard deviations its cadence count is above what the same bars resampled would give; up, the same for its pitch-class profile correlation. The winner on cadences is key C at 6.2 standard deviations and on profile C at 1.1, and they agree. Standard deviations above chance are the same unit whatever produced them, which is the commensuration this figure is for — and it costs something: it assumes the two chances are equally interesting, which is a weighting in disguise. The obvious null does not work at all for one of the two: shuffling the bars leaves a pitch-class histogram exactly as it was, so its spread is zero and every key scores nothing against it. Harmony and voice leading

A count and a correlation

Cadence evidence is a count of ordered pairs and profile evidence is a correlation with a template, and the cadence essay refused to total them because they are not in the same units. Score each against its own chance and they are — standard deviations above chance are the same unit whatever produced them. Then the trouble moves: the obvious null does nothing at all to one of the two, because shuffling the bars leaves a pitch-class histogram exactly as it was.

How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; I – IV – V – vi totals 7.3 bits over 3 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire. Harmony and voice leading

Surprise is a number

The chord that did not come was described rather than measured. Its measure is the information content of what did arrive, and a model of the probability has been to hand since the key-finding essays — eight root-motion weights, ordinal and stipulated. Reading them as a distribution prices a deceptive cadence at 2.71 bits against a perfect one's 1.85, and turns up the fact that the largest of the eight had never been read by anything.

The fingerboard, with the hardest place on it. Every written pitch from G3 to G6 on every string that can reach it, shaded by the width of Schelleng's bow-force window there — dark is narrow, which is a note that is hard to start. Two effects are multiplied and neither earlier figure could show the other: the body's admittance, which depends on the frequency, and the bowing fraction and string impedance, which depend on where the hand is. The worst place is C♯4 on the G3 string, 6 semitones up it, at a window of 12.2 against 471 at the easiest — a factor of 39 across the instrument. It sits where the body's A0 resonance crosses the heaviest string played high, which is the compounding this figure was drawn to find: neither variable alone puts a minimum there. Instruments and their design

The hardest place on the fingerboard

Two things narrow a bow's window and each has been drawn alone. The bridge's admittance is a function of frequency; the bowing fraction is a function of where the left hand is. They meet on a real fingerboard, and multiplying the two curves gives a map with a worst place on it — C♯ on the G string, sixth position, where the body's air resonance crosses the heaviest string played short. The window there is twelve, against three hundred and eighty at the easiest.

Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together. Perception and the listener

The partials that do not die together

The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.

The same interval, started on each of the twelve. An interval of 7 semitones started on each pitch class of a major key, against how strongly the key specifies its two notes — the mean of the probe-tone profile at each. The interval account says the listener encodes a distance, so the key cannot enter and the prediction is a horizontal line at 5.4 cents. The degree account says the listener refers each note to the key, so its precision on a note falls as the key's specification of that note weakens; scaled to agree at the most stable start, it rises from 5.4 cents on C to 8.5 on E♭. Every earlier figure measures a quantity the second account says is not being formed at all. Intervals and chords

The quantity a rival account says is not there

Two earlier essays measure how much two notes' errors are correlated through a shared anchor, and price what that correlation would be worth. There is a rival account in which a listener refers each note to a key and never forms the distance at all — under which the correlation is not small, it is a description of something that is not happening. The two accounts agree on almost everything and disagree on one manipulation, and the manipulation costs an afternoon.

How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation. Form and structure

A detector whose resolution the performance sets

The boundary detector lost its free parameter when the psychological present became a number of notes at a stated tempo, and what that held still was named at the time: a performance slows into a phrase end, so the number of notes inside the present is not the same everywhere in a tune — it falls exactly where a boundary is. Making the width follow the performance recovers half of what removing the parameter cost, and honestly leaves the other half.

The surprise of each chord, against the uncertainty it arrived into. The information content of each step — minus the log of its probability under a distribution that multiplies the root-motion weight by how well the destination triad's notes fit the key — with the entropy of the moment before it drawn behind. I – IV – V – I: I→IV 1.71 bits, IV→V 2.80 bits, V→I 1.60 bits, against a mean uncertainty of 2.63; I – IV – V – vi: I→IV 1.71 bits, IV→V 2.80 bits, V→vi 2.61 bits, against a mean uncertainty of 2.63. A surprise larger than the entropy it arrived into is an outcome the model was not expecting even given how uncertain it was; one below it is an outcome the model had already mostly bet on. An earlier essay produced the first of those numbers and had no way to produce the second, because a set of preferences is not a distribution and only a distribution has an entropy. Harmony and voice leading

A chord, given a key and a predecessor

A chord's improbability has been priced from its root motion alone, which left one multiplication unmade: a chord is also improbable because its notes do not fit the key, and that number has been available since the probe-tone profile. Multiplied and renormalised, the two give a conditional distribution — and a distribution has an entropy, which is the quantity a surprise has to be read against and which a list of preferences cannot supply.

A cycle whose position is in the instrumentation. 3 isochronous layers over a cycle of 16 steps, at periods 16, 8, 4. Every layer on its own is perfectly symmetric and tells a listener nothing about where they are; the combination gives 4 distinct signatures over 16 steps, and hearing one of them leaves 2.81 bits unknown. The information is in which instruments sound rather than in where the onsets fall, which is a different answer from the one a single timeline gives — and it needs no asymmetry anywhere. The cost is 7 strokes a cycle, 0.44 to the step, spread over 3 players. Rhythm and metre

A cycle that says where it is

Euclidean timelines were asked how quickly they tell a listener where in the cycle they are, and answered it with rotational asymmetry: a symmetric pattern never locates at all. A colotomic cycle answers the same question with nothing asymmetric in it. Several isochronous layers at nested periods — a gong every sixteen, a kempul every eight, a kenong every four — put the position in which instruments sound, and the position is legible from a single stroke.

What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged. Perception and the listener

The exchange rate nobody has

There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.

Expectation as a curve, and what a change costs where it lands. Every quantity so far is attached to a chord change: a list of surprises, one per event. A listener's expectation is continuous — it sharpens through a bar and collapses when the change arrives — and the two ingredients for it are already here, the harmonic rhythm and the metrical beat weights. The curve is the hazard: given that the chord has not changed yet, the chance that it changes on this beat. It runs from 0.043 on the weakest beat to 0.290 on the downbeat, a ratio of 6.7, against 0.125 if every beat were alike. The marked beats are where the changes actually arrive, and their timing bill is 3.6 bits against 6.0 for a listener with no metre — so these changes are 1.7 times cheaper to expect than a metreless listener would find them. That term is new: an earlier essay prices which chord arrived and this prices when, and a listener meets the sum. Harmony and voice leading

Expectation is a curve, not a list

Every quantity so far is attached to a chord change: a list of surprises, one per event. A listener's expectation is continuous, sharpening through a bar and collapsing when the change arrives — and the two ingredients for it were already here, in two other accounts. What comes out is a second surprise, for when a chord arrives rather than for which one it is.

The same tune read at six widths of the psychological present. A later essay made the detector's smoothing width a function of position, which removed its last free parameter but one — and the one it cannot remove is the width of the psychological present, because that is a fact about listeners rather than a choice. So the honest object is not a reading but a family of them, one per width. A short present finds 6 boundaries and a long one finds 2, and the family agrees on 0 of them. The fixed-width control, at its own best width, scores 0.67 against the adaptive readings' 0.67, 0.75, 0.33, 0.33, 0.40, 0.40 — so the adaptation does not win, which is what that essay reported too. What the family adds is the ordering: a boundary in every row is a different claim from one in a single row, and a single reading has no way to say so. Rhythm and metre

A family of readings

Removing the detector's free parameter, and then its constant tempo, cost persistence both times — the property that made its boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than one, indexed by the width of the psychological present, which is the one parameter that cannot be removed, because it is a fact about listeners.

An ensemble finding an asynchrony nobody told it about. An earlier essay produced a map of required leads — which notes of a scoring have to be played early, and by how much — and nothing tells the players those numbers, because they are a property of the instruments' attacks rather than of the music. So an ensemble has to find them, and the mechanism is already here: each player hears sounds rather than onsets and moves their next onset toward the mean of the others'. The spread of arrival times starts at 20 milliseconds and settles at 4, crossing 5 milliseconds after 5 beats — about 1.3 bars of four. The leads it converges on match that map to within 0.2 milliseconds, which is what makes this a convergence rather than a coincidence: the fixed point of players listening to each other is every player leading by their own attack. Rhythm and metre

How many bars an ensemble needs

The map of required leads is something nobody tells the players, because the leads are a property of the instruments' attacks. So an ensemble has to find them, and the mechanism is the one the microtiming essays describe: each player hears sounds rather than onsets and moves toward the others. It converges on the map to within a fifth of a millisecond, in five beats, and there is a best correction gain.

How much of the reading comes from what has not happened yet. Every margin reported earlier is two-sided: the best path through a key at a bar is the best score into it plus the best score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line, because the dynamic program already had both halves separately. The mean margin falls from 3.90 bits with hindsight to 1.79 without it, so 54 per cent of this passage's certainty is retrospective. The two passes never disagree about which key is best here, so the hindsight buys confidence rather than a different answer. This is the quantity every earlier essay has assumed and none has measured. Scales and modes

How much of the reading arrives late

Every margin reported earlier is two-sided: the best path through a key at a bar is the score into it plus the score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line. On a thirty-two-bar song it removes more than half the certainty, and on a passage built to be ambiguous it changes the key named at nine bars out of eleven.

The period is still there, and it is wider. The autocorrelation of a 12-partial complex on 220 hertz, drawn twice: steady, and averaged over one cycle of a 71-cent vibrato. A vibrato moves every partial by the same number of cents, so the complex is exactly harmonic at every instant and nothing is mistuned — what moves is the period the extractor is looking for. The peak survives. It loses 6 per cent of its height above the surrounding lags and gains 11 per cent in width, because the vibrato swings the period by 0.37 milliseconds against a peak 0.90 wide. Its maximum also moves, to 2.4 cents sharp of the still tone's, which is a prediction with a sign in it. Instruments and their design

The pitch that does not wobble

Three earlier essays have treated a vibrato as a modulation of roughness. The reason singers use one is what it does to the note, and there is an extractor here that turns a set of partials into a pitch and has never been asked what it does with partials that will not hold still. The period survives, at a cost that rises with the extent — and the practice stops within a hair of where the cost becomes total.

One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics. Form and structure

The dissonance arrives and the dynamic does not

A scoring decides two things at once and both of them have to be integrated by a listener before they exist. The loudness smoother's release is two seconds and the roughness window is thirty-seven milliseconds, and that ratio of fifty decides which of the two survives at the pace music is actually played. Nothing anybody performs is fast enough to blur a dissonance, and a great deal of it is fast enough to average a dynamic.

The cue that settles it. Every spectrum to hand, arbitrated by the earlier competition and then again with the onset cue added at equal weight. 3 of the 5 change their verdict, and all 3 change the same way — from splitting into two streams to staying as one: a piano string, a bell, a bar. Nothing changes the other way, because the onset cue on a struck source votes for fusion on every partial and can only ever push toward one stream. The bell is the case worth naming: its partials are wildly inharmonic and it is heard as one sound, which is a fact the harmonicity cue alone cannot produce. Perception and the listener

The cue that settles it

Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.

A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance. Perception and the listener

A page has two decibels

The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.

A listener has a quarter of an analyst's confidence and the same answer. The mean margin between a passage's best two key readings, against how many bars a listener's memory of the evidence takes to halve. The two-sided reading — the one that uses bars that have not happened yet — sits at 3.90; the forward pass with perfect recall at 1.79; a forward pass whose evidence halves every 6.6 bars at 1.00. The dots' size is how often that reading names the same key as the two-sided one: 100 per cent at perfect recall and 81 at a one-bar half-life. So forgetting costs a great deal of confidence and very little accuracy — the key is robust and the certainty is not. Harmony and voice leading

The listener who forgets

Setting an analyst's reading of a key against a listener's measures what arrives late. Both passes assume perfect recall of their own half — which is as wrong going forward as knowing the future is going back. Put a decay on the forward pass and a listener with a memory of a few bars keeps a quarter of the confidence and nine tenths of the answers.

The number nobody has moves the size and not the order. The mean total surprise per chord change, against how much the two surprises share. At zero they are independent and the total is their sum; at one they are the same event and the total is the larger of the two. The mean falls by a factor of 1.53 across that whole range, which is the size of the thing a corpus would settle. The ordering of the events by total surprise does not move at all until the very end: 5 of the 6 correlations swept give exactly the ordering independence gives, and only perfect dependence changes it, by 3 places out of 7. The most surprising event in the passage is the same one at every correlation. So the corpus three separate accounts have recorded wanting would change what this figure reports and not what it concludes. Harmony and voice leading

Two surprises and one event

A chord change is surprising twice over — in which chord it is, and in when it comes — and a listener meets one event. Adding two surprises needs to know how much they share, which is a fact about a repertoire nobody has. Sweeping it instead: the total moves by half, the ordering does not move at all, and the most surprising moment in a passage is the same one whatever the answer turns out to be.

The same eight notes are four times as much to read. How many bits each note of a line carries, taken as minus the log of the probability of the interval that reached it, under the distribution of melodic steps measured over the tunes used throughout. A scale costs 1.76 bits a note and a wide leaps costs 7.02 — a factor of 4.0 at the same number of notes on the page. Every quantity computed until now counts notes, and the page cannot tell these apart: eight quavers are eight quavers of horizontal space whichever line they spell. Scales and modes

A reader does not read notes

Eleven earlier essays count notes, and the page cannot tell one line of eight quavers from another. A reader can: a scale of eight is one object where eight leaps are eight. Measured against the melodic interval distribution, the same eight notes are four times as much to read — and the eye–hand span, the best-measured quantity in the reading literature, is four notes of a tune and one of a leaping line.

A rest is a diminuendo, and a long one. How far a listener's running impression of loudness falls during a silence, converted into the diminuendo that would have taken it the same distance. Half a second of nothing is worth 3.2 decibels, a second and a bit is worth 8.1, and two and a half seconds is worth 17. The marked line is three and a half seconds, which is where a gap starts to be heard as an ending rather than as a pause: at that length the reference has fallen by 24 decibels, which is more than a fortissimo to a pianissimo. A tempo swept over a factor of ten, a deceleration over a factor of three and a gesture length over a factor of forty-eight all returned the same reading to four significant figures. This one moves it by twenty-four decibels. Form and structure

A rest is a diminuendo

Three parameters swept over factors of ten, three and forty-eight returned the same reading to four significant figures. The one manipulation left unswept moves it by twenty-four decibels: a silence. A listener's running impression decays at the loudness smoother's two-second release, so a general pause is a diminuendo nobody wrote, and at the length that makes a gap an ending it is worth more than any marking a composer has.

A clarinet's partials, each on its own resonance. The 8 partials of the clarinet's chalumeau D that ride an impedance peak, each building toward its steady amplitude as 1 − exp(−t/τ) with τ = Q/πf from that peak's own Q. The time constants run from 11.9 milliseconds to 64.5, so the partials do not arrive at different times — they all begin the instant the reed does — and what differs is how fast each approaches its final level. The horizontal bars are how far apart the first and last are at three criteria: 5.5 ms at 10 per cent, 36.4 ms at 50 per cent, 121.0 ms at 90 per cent. A twenty-millisecond asynchrony is the threshold for hearing a partial out of a note, and this note crosses it at 32 per cent of steady amplitude — so whether a blown note's onset cue is unanimous or divided is decided entirely by how far along a partial has to be before it counts as having started. Perception and the listener

A blown note does not start late, it starts slowly

Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.

Eleven partials is one partial too many. What fraction of a spectrum the harmonicity census finds fused, against how many partials it is asked to census. At ten a perfect harmonic series fuses 10 of 10 and the fundamental it finds is the right one. At eleven it fuses 5 of 11 and the fundamental jumps to exactly 2.00 — the octave above. The cause is the cap the census carries for a reason established earlier: without it a bell fuses perfectly at a fundamental nobody could hear, so the search refuses any fundamental more than about ten harmonics below the top partial. At eleven partials the first thing that cap excludes is the series' own fundamental, and the census then takes the octave and calls every odd partial inharmonic. So the number of partials and the cap are the same number, and nothing had ever said so, because every earlier figure censuses ten. Perception and the listener

Eleven partials is one too many

Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.

Where a cycle of 16 at 16, 8, 4 outruns the listener's memory. The residual uncertainty a listener is left with once the evidence has stopped accumulating, against how long one turn of a 16-step cycle takes. The listener's memory of a step halves after 3.5 seconds throughout; what changes is how many steps that is. At a cycle of 1.6 seconds it is 35 steps and every design reaches certainty, which is the regime a clave is played in. At 60 seconds it is 0.93 steps and none of them does: layers at 16, 8, 4 settles at 1.89 bits, son clave settles at 2.27 bits, the bossa-nova pattern settles at 2.38 bits, the best single line of 7 settles at 1.85 bits. That is the range a gong cycle occupies, and it is the design that wins there. Rhythm and metre

The cycle that outruns the memory

A timeline and a colotomy were compared at equal strokes and the comparison had no clock in it. A memory span is a number of seconds and a cycle is a number of steps, so the two only meet through a tempo — and at a clave's two seconds a listener's memory covers twenty-eight steps and forgets nothing, while at a gong cycle's forty it covers 1.4 and forgets almost everything. The single line is the better locator up to twenty-three seconds a cycle and the layered code is better after it, which is very close to where each is actually used.

An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none. Perception and the listener

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

Which bars the key is decided by, and which bars it is believed on. Every bar of a 32-bar scheme removed in turn, with what its absence costs. The column is how far the passage's mean margin falls without that bar — how much of the model's certainty it supplies. The dot is how many bars are then read as a different key — how much of the answer it supplies. The 18 bars an analysis would point at — a section opening or closing, a dominant, the chord a dominant resolves to — average 0.079 bits of certainty and 0.50 bars moved; the 14 ordinary bars average 0.047 and 0.07. So the structural bars carry 1.7 times as much certainty as the ordinary ones, and 7.0 times of the answer. Those are different quantities, and the second is the one an analysis is about: an ordinary bar can carry a great deal of a passage's certainty and none of its reading. Harmony and voice leading

The bars a key is made of

A discount that treats every bar alike is the wrong shape for a memory, so take each bar away in turn and see what it was worth. The bars an analysis points at carry seven times as much of the answer as the ordinary ones and slightly less of the certainty — and a memory built only of them reads the passage worse than a memory with no structure in it at all.

A third of the timing surprise is paid for by nothing happening. Every beat of a bar of 4/4 at 1 chord change a bar, with what it costs a listener in expectation. The lower block is the arrival — the chance the chord changes here times what that change costs to be surprised by. The upper block is the hold, which is paid when the chord does not change and is charged at every beat rather than at every event. Over the bar the arrivals come to 2.61 bits and the non-arrivals to 1.29, so the term never spent before is 33 per cent of the total rather than most of it. The two together are exactly the binary entropy of each beat's own hazard, which is why the bar's whole timing bill is 3.90 bits and cannot be raised by rearranging where the changes fall. Harmony and voice leading

The surprise of nothing happening

A hazard charges a listener twice — once when the chord changes and once, quietly, at every beat it does not. The second term was expected to dominate, because there are more beats than changes. It is a third of the bill at one chord a bar and never reaches a half at any rate a metre survives, because the cost of a beat where nothing happens is second-order small.

What a chord change costs where it actually lands. Every beat of a bar of 4/4 at 1 change a bar, priced as a listener meets it: the pale block is what has already been paid waiting through the beats the chord did not come on, and the dark block is the arrival itself. Their sum is what it costs to be surprised by a change here. The last column is the remaining case, never priced before: no change in the bar at all, at 1.61 bits and a probability of 0.33. The 9 costs are a proper distribution — 1.000 — which is the check that this is one model rather than two. And the spread is the finding: the arrival term alone puts a factor of 6.7 between the best and worst beat, and counting the waiting makes it 19. Harmony and voice leading

Where the chord actually lands

Every timing surprise so far is evaluated on a downbeat: the curve spans a factor of 6.7 and every number is read at its peak. Charge the waiting as well as the arrival and the costs over a bar become a proper distribution, the spread between the best and worst beat rises to a factor of 19.5, and a third of the probability sits on a bar in which nothing changes at all.

How much of A4's pitch error a key could possibly remove. The largest correlation two notes' pitch errors can have at A4, against how long each note lasts. A note's error has two parts and only one of them is the listener's: the steady-tone limen of 4.04 cents, which context might reduce, and the bound a note of length T puts on its own frequency, which context cannot touch. Taking them in quadrature, the shareable fraction is the curve. At a quarter-second note it is 0.21, so the one-half priced earlier is not available at all until each note lasts 486 milliseconds — which is exactly the crossover found by a different route, because a correlation of a half is the two parts being equal. The step is the convention used here, which takes the larger of the two rather than their sum and therefore says the shareable fraction below the crossover is zero. Intervals and chords

The part of the error a key cannot touch

Four earlier essays turn one dial — the correlation between two notes' pitch errors — and apply it to the whole of a note's limen. Half of that limen is not the listener's: a note of finite length does not carry its frequency more finely than 1/2T, and no context can put information into a signal that is not there. So the correlation has a ceiling, it is 0.21 at a quarter-second note at A4 and 0.07 at A2, and the figure that prices a correlation of one half is drawn where one half is unavailable.

A struck note's two ends are the same for every loss law. The partial levels of a string spectrum struck at 80 decibels on 130.8 hertz, and what is left of it when the fundamental itself falls under the threshold of hearing, for three laws relating a partial's decay rate to its number. The left panel is every one of them: a loss law cannot change the spectrum at the instant of the strike, because no time has passed. The other three are every one of them too: whatever the law, the note ends with nothing above the threshold. So both ends of the slide are shared, and everything that distinguishes an exponent of 0.5 from an exponent of 1 from an exponent of 2 is in the middle. Timbre and acoustics

The middle nobody could have guessed

A struck note has no steady state, only a slide from one spectrum to another — so the question is what the middle carries that the ends do not. The answer is exact rather than statistical: every loss law in the family leaves the strike with the same spectrum and ends in the same silence, so both endpoints carry precisely nothing about which of them it is. The whole difference is 41.3 decibels, and it peaks 0.38 seconds in, seven per cent of the way through the note.

Where a section's fluctuation stops being a beat, on 220 hertz. Two rates up the spectrum of a 220 hertz note sung by a section whose voices are spread by 15 cents. The rising line is the beat rate between a typical pair of them, which grows with the partial because a mistuning in cents is a difference in hertz that scales with frequency; it reaches the 15 hertz at which a beat stops being a beat by partial 5.5, at 1217 hertz. The flat line is the amplitude modulation the vibrato imposes through the formants, which is 6.0 hertz at every partial because the vibrato modulates every partial by the same number of cents at the same rate. The two are equal at 487 hertz. Above 1217 hertz the beating has become roughness and the only fluctuation left is the vibrato's — and that frequency is the same one an octave up, where it is partial 2.8 instead. Instruments and their design

The rate that does not rise with the partial

Twelve earlier essays give every vibrato the same six hertz, and the measured spread is 5.5 to 7.5. Putting the two fluctuations a choir contains on one axis shows why the rate matters: the beating between mistuned voices rises with the partial and leaves the range a listener follows as fluctuation at 1,217 hertz, while the vibrato's own modulation is six hertz at every partial. Above that frequency a section fluctuates by vibrato alone — and if every singer had the same rate, it would barely fluctuate at all.

A notehead in four parts costs 1.30 bits and one in two parts costs 1.89. What one notehead asks of a reader, against how many parts are on the page, for a progression realised by the voice-leading solver used here at 2 semitones of motion a voice a chord. The horizontal rule is a note of a single melody under the measure established earlier, 1.89 bits, which is what a texture costs when its parts have to be read one at a time. The bars are the harmonic reading: the chord, charged at the worst case of 2.81 bits for one of seven diatonic degrees, plus the logarithm of how many voicings of it the previous chord could legally have moved to. At two parts there is no bar, because a duet has no complete voicing of any triad — it cannot state the harmony and has to be read as 3.79 bits of two independent lines. Every thicker texture is cheaper a notehead than the thin one, and the four-part figure is an upper bound. Harmony and voice leading

Four parts are easier to read than two

Twelve earlier essays read one line, and a score is several at once. Measured through the voice-leading model, a notehead of a four-part chorale asks a reader for 1.30 bits and a note of an independent line asks 1.89 — so twice the ink is less than three quarters of the load. The reason is a boundary those essays already established: a duet has no complete voicing of any triad at all, so two parts cannot be read from their harmony and have to be read as two melodies.

A head model a few millimetres out reads every azimuth but the front. The azimuth a listener reports against the azimuth a source is at, for internal head radii from 8.22 to 9.28 centimetres against a true radius of 8.75. The delay a source produces is (r/c)(θ + sin θ) and the listener inverts it with the radius they believe they have, so their answer solves θ̂ + sin θ̂ = (r/r̂)(θ + sin θ). Every curve passes exactly through the origin: on the median plane there is no delay and therefore no error, whatever the head model is. The error grows with azimuth and is largest at the side. An internal head 5.3 millimetres too small runs out of azimuth at 82 degrees: beyond that the world is delivering a delay larger than any its owner's model can produce, and every source out there collapses onto the side. Perception and the listener

Where a wrong head gives itself away

Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.

Thirty cents out of tune is heard as 8 on a 125 ms note and 25 on a 1 s one. How far out of tune a note sounds against how far out of tune it is, at 4 note lengths, at 440 hertz. The key is treated as a prior over pitch: a mixture of Gaussians on the twelve scale degrees, weighted by Krumhansl and Kessler's probe-tone profile and given the width the degree account already uses. The likelihood is the note's own effective limen, which for a short note is the Fourier bound 1/2T. The estimate is the posterior mean, and the shrinkage toward a prior is one line of arithmetic. At 1 s a thirty-cent mistuning is heard as 25.0 cents and at 125 ms as 7.8. Every curve turns back up near the middle of the semitone, because past there the nearest degree is the other one and the pull reverses. The buttons sound at A4, which is the pitch the figure is computed at, at the shortest note length it draws. Intervals and chords

A short note is heard more in tune than it is

Four earlier essays treat a key as something that reduces the noise in a pitch judgement. Treat it instead as a prior and the prediction changes kind: not a smaller error but a systematic bias, pulling a short note toward the nearest scale degree by an amount the Fourier bound sets. Thirty cents out of tune on an eighth-of-a-second note is heard as eight. And the part the debt got wrong is the part that matters — the bias does not vanish on a long note. It stops at 17 per cent at A4 and at 48 per cent at A2, because the likelihood's width has a floor that no duration removes.

The mixture at which the chain acquires a tonic. Eight turns of i–iv–v–i in A minor, natural throughout, read at every mixture of the two emissions: nought is the original set overlap against the triad on each degree, one is the measured probe-tone profile rotated to each candidate key. The blocks along the top are the name the model gives, the line below is how far ahead of its best rival that name is. Below a mixture of 0.55 every bar is called C major, which is the collection and not the key; at and above it every bar is called A minor. The margin collapses to 0.33 bits at the crossing and recovers to 4.56 — higher than the 3.38 it started at, because a profile has an opinion about this passage and an overlap does not. Harmony and voice leading

A tonic bought with the function

Putting the measured probe-tone profile inside the ordered key-finder is one term, and it does what was predicted: the natural-minor passage is named A minor at every bar instead of C major. It also does two things nobody predicted. It renames a scheme that has been read in the wrong key at every bar since the day it was written, and it destroys the chord's function while it is buying the key.

The reading the joint search was never offered. The best chord at each mixture of the two segmentation cues, and what the same weighting gives the same notes shuffled into a different order. Both fall along the axis, and most of the fall is the ruler rather than the music: a metrical weighting over a bar of eight spans a factor of eight and a three-to-one duration spans three, so the weighted note mass is 2.1 times more concentrated at the left of the figure than at the right, and a concentrated mass is easier for four notes to cover. What is not the ruler is the gap. It is widest at a mixture of 0.75, where the reading is D minor7 at 2.15 standard deviations above its own null, against 1.12 for C major7 at a mixture of nought. The joint search holds this axis at nought, so D minor7 is not among the hypotheses it considers. Harmony and voice leading

A fourth decision, and two that were never made

The joint search resolves key, metre and segmentation together and holds the segmentation's cue mixture at zero. Adding the mixture is one loop, and reading the search in order to add it turns up something worse than a missing axis: on the passages it is drawn on, the key it reads is the same key at all forty-eight of its hypotheses and the metre scores every barline identically. The fourth axis then cannot be ranked at all until each reading is measured against its own null, because a mixture changes the ruler and not only the answer.

Eighty-one chords the expectation model cannot tell apart. Every voicing of a dominant seventh on G inside the three octaves above its own root, placed by how rough it is and how far its outer voices are apart, and coloured by which member of the chord is at the bottom. The roughness runs from 0.269 to 1.449, a factor of 5.4, computed from each voicing's own spectrum under Plomp and Levelt's roughness model. The identity surprise the expectation model assigns is 3.51 bits for every one of the 81, because it is a function of a scale degree and its predecessor and there is no register anywhere in it. What separates them is spacing rather than inversion: roughness falls as the outer voices spread apart, correlating -0.42 with the span, and is indifferent to which member of the chord is at the bottom at 0.02. The seventh in the bass is not what makes a chord rough; a fourth and a third packed together at the bottom of the range is. Harmony and voice leading

Eighty-one chords, one number

A dominant seventh has eighty-one arrangements inside three octaves and their roughness spans a factor of five and a half. The tonal-expectation model gives every one of them the same 3.51 bits, because its states are scale degrees and there is no register anywhere in them. Conditioning the surprise on the voicing costs no corpus — and the arithmetic says the conditioning belongs beside the probability rather than inside it, for three reasons that can each be computed.

Where a bar of 25 and a bar of 9 first disagree. The additive metre 2+2+2+3+2+2+2+3+2+2+3 — 25 units, an onset at the head of every group — with the accents it predicts drawn above the accents predicted by reading it as a repeating bar of 9, which is the cut of it that agrees longest. The two rows are identical for 24 consecutive steps and differ for the first time at step 25, where the shorter reading expects an accent and the metre does not supply one. Nothing before that step distinguishes the two hypotheses, so a listener who has not heard 25 consecutive steps has no evidence either way — whatever they are disposed to hear. Rhythm and metre

A twenty-five is a nine until its last unit

Every account of long additive metres says they are heard as groups of shorter ones, and the metre-induction model had never been pointed at the claim. Pointed at it, the model does not prefer the group — it prefers the long bar outright, and would go on preferring it more the longer anybody listened. What it cannot do is start: the evidence that separates a bar of twenty-five from a bar of nine does not exist until the whole bar has been heard, and at the tempo an unequal metre is best played at the psychological present holds sixteen units.

A bar of silence is worth 13.1 decibels written last and 6.0 written first. Four closing gestures, drawn against how many seconds of silence each contains. All four hold the same 12-part texture, write the same 6.0-decibel diminuendo over 4 seconds, and differ only in what order the diminuendo and the silence are written in. The quantity is how far the listener's running impression has fallen when the final chord arrives, in decibels of equivalent diminuendo. Written with the silence last, a general pause of a bar is worth 13.06 decibels and one of three and a half seconds is worth 28.3. Written with the silence first, both are worth 6.00 — exactly the diminuendo's own depth, because the music resuming after the silence puts the reference back at its own level. The silence written first is worth less than the silence written with no diminuendo at all, which reads 8.12 at a bar: a diminuendo placed after a general pause takes 2.12 decibels away and adds nothing. Form and structure

A general pause is spent by the note after it

Whether a composer should write the pause before the diminuendo or after it looks like a question about how big the ensemble is. It is not. Forty decibels of ensemble are worth one decibel of silence, and the order is worth seven — because a running impression rises twenty times faster than it falls, so half a general pause is spent by ninety-seven milliseconds of sound.

A note on the downbeat costs 1.89 bits and one on the offbeat 5.70. What each position in a bar of 4/4 asks of a reader, by two routes. The solid bar counts where the notes of this collection's own three tunes actually fall — 28, 2, 26, 4, 28, 0, 16, 0 notes at the 8 positions — and takes minus the log of the frequency. The rule across each bar is the same quantity from the stated metrical weights, 1, 0.15, 0.5, 0.15, 0.85, 0.15, 0.5, 0.15, normalised and logged the same way. Nothing makes the two agree. They put the eight positions in the same order, and they price the tunes' own rhythm a fifth of a bit apart — while differing by more than a whole bit about the quaver after the downbeat, which two notes in a hundred and four ever use. The two positions these tunes never touch at all are drawn at the floor, which is the same floor the melodic measure gives an interval nobody plays. A weight was always a probability waiting to be read as one. Rhythm and metre

Where the note is costs more than which note it is

Thirteen earlier essays measure a page, and the three that price a reader price only its pitches — every line they measure is a run of equal notes. A metrical weight normalised by its own sum is a probability, and minus its logarithm is bits — the same substitution made earlier for intervals. Measured over the tunes used throughout it comes out at 2.23 bits a note against the pitches' 1.89, so the larger half of a reader's load is where the note is.

The same interval, mistuned by the same amount, at each of its two ends. A C to G in the major key, played 25 cents wrong, with the departure carried by the lower note, split between the two, and carried by the upper note. All three are the same interval size; what differs is which note is off the scale. The share of the departure that survives into what a listener hears is 32 per cent when the lower note carries it and 56 when the upper does. The middle bar is the mean of the other two to within a hundredth, so the averaging is linear and the asymmetry is the whole of the effect. Two things produce it: the prior is 9.0 cents wide at the C and 10.0 at the G, and the likelihood is 13.2 cents wide at the lower pitch and 8.8 at the higher. Intervals and chords

An interval is two posteriors subtracted

Treating a key as a prior over one note predicts that an interval's pull is not the single-note pull doubled, because the two degrees are not equally weighted. Half of that is wrong: splitting a mistuning between the two notes gives exactly the mean of what each end gives alone, to a thousandth, at every one of the twenty-one intervals in the scale. What is not the mean is which end carries it — and the pull turns out to be largest not on the shortest notes but on notes of about an eighth of a second, where the likelihood is a quarter of a semitone wide.

Which intervals in a key can be mistuned invisibly, and from which end. Every interval between two degrees of the major scale, ranked by how differently its two ends treat a 25-cent departure on notes of 0.25 seconds. The C to B is the most lopsided, at 47 points: a mistuning on its upper note reaches the listener nearly 2.5 times as strongly as the same mistuning on its lower one. The D to E is the most even, at -0. A negative bar is an interval whose LOWER note is the one that carries a mistuning into the listener, which happens whenever the lower degree is the less specified of the two. No account of interval perception predicts a table like this, because an interval is usually treated as one quantity rather than as a difference of two estimates. Intervals and chords

Which end the mistuning is on

Twenty-one intervals in the major scale, each with two ends, and the same twenty-five cents reaches a listener at anywhere between 32 and 79 per cent of its size depending on which of the two notes carries it. The most lopsided is the tonic to the leading note, where a departure on the upper note arrives two and a half times as strongly as the same departure on the lower. Two mechanisms produce it and they can be separated by one flag: two thirds of the asymmetry is register and one third is the key.

Every product of a just interval is a harmonic of the note it implies. Each interval drawn as two harmonics of a fundamental it does not contain — the lower note is harmonic q and the upper harmonic p — with its three combination tones placed on the same numbering: the difference tone at p − q, the cubic product below the pair at 2q − p, and the one above at 2p − q. minor second 16:15: 1, 14, 17; major second 9:8: 1, 7, 10; minor third 6:5: 1, 4, 7; major third 5:4: 1, 3, 6; fourth 4:3: 1, 2, 5; fifth 3:2: 1, 1, 4; minor sixth 8:5: 3, 2, 11; major sixth 5:3: 2, 1, 7; minor seventh 9:5: 4, 1, 13; major seventh 15:8: 7, 1, 22. The shaded column is the fundamental itself. The difference tone sits on it for every interval up to the fifth, the cubic product for the fifth and every interval above except the minor sixth, whose products are its fundamental's octave and twelfth. Intervals and chords

The tone on the root changes hands at the fifth

Every combination tone of a just interval is a harmonic of a fundamental neither note contains, and which harmonic is fixed by the ratio. The difference tone lands on that fundamental for every interval up to the fifth; the cubic product lands on it for the fifth and every interval above except the minor sixth. So the loud product names the root of a narrow interval and the quiet one names the root of a wide one — and a just major seventh's difference tone is a note seven harmonics up that no keyboard has.

A major triad's combination tones, against its own notes. The three notes of a major triad on C4 in root position, voiced C4–E4–G4, as tall lines, and every combination tone its pairs make, as short ones: difference tones lowest, cubic products taller. In just intonation 2 cubic products land exactly on a note of the chord, and none comes within forty hertz of one. In equal temperament no cubic product lands on a note of the chord, and the nearest miss is 5.63 hertz. Intervals and chords

A major triad's combination tones are its own notes

Play a just major triad of pure tones and two of the ear's cubic products land exactly on its root and its fifth. The reason is a condition rather than a coincidence — a chord's cubic products fall on its own notes when its middle note is the mean of the outer two in hertz — and it holds for the major triad in root position and in the six-four, and for no minor triad in any position or tuning. Equal temperament misses the landing by one number, 5.6 hertz on middle C, which is a beat that belongs to no pair of notes in the chord.

A scale in parallel thirds has a line underneath it that nobody plays. A major scale on C4 harmonised in parallel diatonic thirds, with the difference tone f₂ − f₁ of each pair drawn as a third line. In five-limit just intonation that line is C2, A1, C2, F2, G2, F2, G2, C3. In equal temperament it moves to C♯2, A1, B1, F♯2, A♭2, E2, F♯2, C♯3, departing from the just line by +67, +33, −82, +69, +65, −80, −84, +67 cents. Intervals and chords

The bass line under a passage in thirds

A major scale harmonised in parallel thirds gives the ear a difference tone under every pair, and in five-limit just intonation those tones are a diatonic bass line — C, A, C, F, G, F, G, C — made of the scale's own notes. Tempered, the same line moves only by whole tones, a neutral third and a fourth stretched to 650 cents, and wobbles by up to 84 cents from note to note. In sixths the bass is drawn by the other product, because the cubic product of a pair is the difference tone of the same pair inverted.

Every question is answered in a corner, not on a ridge. The degree share times the minor share, over the plane of two cues: how much of the emission is the probe-tone profile, across, and how much a bass note is worth, down, with every chord's root in the bass and a bass rule that rewards the triad rooted on the bass. It runs from 0% at a profile share of 0 and a bass weight of 0 to 87% at 0.6 and 1.5. Bass 0: 0%, 0%, 0%, 0%, 41%, 32%, 46%. Bass 0.25: 0%, 0%, 0%, 0%, 52%, 55%, 44%. Bass 0.5: 0%, 0%, 0%, 0%, 68%, 62%, 51%. Bass 1: 0%, 0%, 0%, 0%, 78%, 77%, 77%. Bass 1.5: 0%, 0%, 0%, 0%, 87%, 80%, 86%. Bass 2: 0%, 0%, 0%, 0%, 87%, 87%, 87%. Scales and modes

Two cues meet in a corner

The profile finds a key's tonic and the bass finds a chord's degree, and until now each was swept with the other held at nothing. Swept together across 42 settings, the plane they make is not the ridge that was predicted. The tonic is a step in one direction, at a profile share of 0.55, and the bass cannot move it; the degree is a slope in the other, rising to 89 per cent as the bass is weighted, and the profile barely touches it. Every question is answered only in a corner of the plane — and the one place the two cues overlap is the one piece of music both can rescue.

The ninety per cent was a ceiling. The share of 188 scheme bars read right on both key and degree with a bass note worth 1.5, for three bass lines — every chord's root, a line moving to the nearest chord tone, every chord's fifth — and two ways of using the bass: rewarding the triad rooted on it, or any triad containing it. Wide bars are with no profile in the emission, narrow bars with the profile alone. root line, root rule: 89% and 86%; root line, member rule: 60% and 29%; smooth line, root rule: 68% and 46%; smooth line, member rule: 61% and 47%; sixfour line, root rule: 3% and 17%; sixfour line, member rule: 52% and 30%. The reading with no bass at all is 47%. Scales and modes

A bass line is not a list of roots

Every bass note the key-finder has been given was its chord's root, and under that line a bass cue reads 89 per cent of scheme bars on the right degree. Give the same chords an economical bass that moves to the nearest chord tone, as a keyboard reduction would, and nearly half of them are inverted. The cue that rewards the triad rooted on the bass then reads 68 per cent at best and worse as it is trusted more; the cue that rewards any triad containing the bass cannot be fooled and stops at 61. The same inverted line does one thing the roots never did: it puts the leading note of each new key at the bottom, and finds the rondo's modulations.

How often the metre, the chords and their product find the barline, chords at 1. Constructed passages of four bars of eight quavers, 100 at each setting, with the barline at the first slot. Rhythm regularity is how much likelier a note is on a strong slot than a weak one; chord regularity is how much likelier a note is to be a tone of its bar's chord than a random scale tone. rhythm 0: metre finds it 10%, chords find it 41%, product finds it 16%; rhythm 0.25: metre finds it 34%, chords find it 35%, product finds it 56%; rhythm 0.5: metre finds it 49%, chords find it 21%, product finds it 66%; rhythm 0.75: metre finds it 50%, chords find it 17%, product finds it 56%; rhythm 1: metre finds it 50%, chords find it 11%, product finds it 45%. Harmony and voice leading

The chords never move the barline

Every hypothesis the joint search had drawn was one bar long, and on one bar with a note in every slot the metre cannot choose a barline at all. Four bars with rests in them make the barline a decision the metre and the chords both have an opinion about, and the prediction was that the chords would move the barline more often than the barline moves the chords. It is the other way round, completely: whenever the two prefer different barlines the search takes the metre's, on up to 72 per cent of passages, and in fifteen hundred passages the chords never once move it. What the chords decide is the one thing the metre cannot see — whether the bar starts on the downbeat or half a bar later — and they decide it right a little over two times in three at best.

The product and the sum of standard scores, finding the barline, chords at 1. Constructed passages of four bars of eight quavers, 100 at each setting, with the barline at the first slot. Rhythm regularity is how much likelier a note is on a strong slot than a weak one; chord regularity is how much likelier a note is to be a tone of its bar's chord than a random scale tone. rhythm 0: product finds it 16%, sum of z finds it 22%, metre finds it 10%, chords, z 35%; rhythm 0.25: product finds it 56%, sum of z finds it 58%, metre finds it 34%, chords, z 38%; rhythm 0.5: product finds it 66%, sum of z finds it 61%, metre finds it 49%, chords, z 19%; rhythm 0.75: product finds it 56%, sum of z finds it 58%, metre finds it 50%, chords, z 14%; rhythm 1: product finds it 45%, sum of z finds it 48%, metre finds it 50%, chords, z 12%. Harmony and voice leading

The chords are a weak witness to the barline

Scaled by its own range, the metre overrules the chords every time the two disagree about where a bar begins. The obvious repair is to score each reading against its own chance — the metre against the same number of notes placed at random, the chords against the passage's notes shuffled across its bars — and add the standard scores. It changes very little: the search finds the barline within six points of where the product found it, and the chords gain the power to move the barline only on passages whose rhythm says nothing, where random notes move it nearly as often. The null's real result is the size of the two witnesses. At the written barline the metre stands up to 5.9 standard deviations above chance, and the chords, with every note a tone of its bar's chord, stand 1.55 above it at best.

An accent moves the longest separable bar from 16 units to 18, and no further. For every bar length from 9 to 25 units, over all 1820 arrangements of twos and threes that are not a repeat of a shorter bar, the fewest and the most consecutive steps before the whole bar beats every shorter cut of it. 9: onsets alone 9 to 14, long beat predicted 7 to 12, downbeat predicted 7 to 8; 10: onsets alone 10 to 13, long beat predicted 8 to 11, downbeat predicted 8 to 9; 11: onsets alone 11 to 18, long beat predicted 9 to 16, downbeat predicted 9 to 10; 12: onsets alone 12 to 17, long beat predicted 10 to 15, downbeat predicted 10 to 11; 13: onsets alone 13 to 22, long beat predicted 11 to 20, downbeat predicted 11 to 12; 14: onsets alone 14 to 23, long beat predicted 12 to 21, downbeat predicted 12 to 13; 15: onsets alone 15 to 26, long beat predicted 13 to 24, downbeat predicted 13 to 14; 16: onsets alone 16 to 25, long beat predicted 14 to 23, downbeat predicted 14 to 15; 17: onsets alone 17 to 30, long beat predicted 15 to 28, downbeat predicted 15 to 16; 18: onsets alone 18 to 29, long beat predicted 16 to 27, downbeat predicted 16 to 17; 19: onsets alone 19 to 34, long beat predicted 17 to 32, downbeat predicted 17 to 18; 20: onsets alone 20 to 35, long beat predicted 18 to 33, downbeat predicted 18 to 19; 21: onsets alone 21 to 38, long beat predicted 19 to 36, downbeat predicted 19 to 20; 22: onsets alone 22 to 37, long beat predicted 20 to 35, downbeat predicted 20 to 21; 23: onsets alone 23 to 42, long beat predicted 21 to 40, downbeat predicted 21 to 22; 24: onsets alone 24 to 41, long beat predicted 22 to 39, downbeat predicted 22 to 23; 25: onsets alone 25 to 46, long beat predicted 23 to 44, downbeat predicted 23 to 24. A present of 3.5 seconds holds 16.0 steps at 218 milliseconds a step, so the longest bar some arrangement of which separates inside it is 16 units on onsets alone, 18 with the long beat predicted and 18 with the downbeat predicted. Rhythm and metre

The accent buys two units, however loud it is

A twenty-five cannot be told from a group of shorter bars on its onsets until more steps have gone by than a listener's present holds, and the obvious objection is that nobody plays an aksak bar as bare onsets: the long beat is louder, and the bar's first beat is marked. So how loud does an accent have to be? The question has a surprising answer. Loudness is not the variable. The existing accent cue changes nothing, and delays the answer where it changes anything. An accent that a reading has to predict works at any strength at all, and at no strength does more than a fixed amount: on the long beat it buys the two steps of a short beat, and on the downbeat it takes every arrangement to one floor — the bar less its last beat — which no cue carried by the notes can break. The longest bar that can be heard as one moves from sixteen units to eighteen.

Six named proportions, as blurred as the durations that make them. Six proportions between two parts of a piece — 1 : 1, 4 : 3, 3 : 2, golden section, 2 : 1, 3 : 1 — placed on one axis by the logarithm of the ratio of the longer part to the shorter, and drawn as bars one criterion wide (d′ = 1) for a listener timing both parts with a Weber fraction of 7%, 15%, 35%. Bars that overlap are proportions that listener cannot tell apart. At 7%, 4 of 5 neighbouring pairs stay apart; at 15%, 3 of 5 neighbouring pairs stay apart; at 35%, 0 of 5 neighbouring pairs stay apart. Form and structure

A proportion is only as fine as its two durations

Analyses of form measure proportions in bars and report them to three figures — a climax at 0.618, a section in the ratio 3 : 2. A listener has each part only as an estimate of how long it lasted, and a ratio of two estimates is blurred by both. Timed as well as anyone times a single second, eleven proportions fit between 1 : 1 and 3 : 1; timed from memory over minutes, two do. The golden section is told from 3 : 2 only below a Weber fraction of 5.4 per cent.

The same forms by the clock and by what is stored. Six forms, each drawn twice: its sections sized by their share of the bars, and sized by their share of what a listener has to store when a bar counts only if it is recognised from 1 bar of context. Returns are drawn pale with a dashed edge. twelve-bar blues: returns take 67 per cent of the clock and 13 per cent of the storage; thirty-two-bar AABA: returns take 25 per cent of the clock and 5 per cent of the storage; rondo, ABACA: returns take 40 per cent of the clock and 10 per cent of the storage; verse and chorus: returns take 50 per cent of the clock and 13 per cent of the storage; two eight-bar phrases: returns take 0 per cent of the clock and 0 per cent of the storage; a four-bar ostinato: returns take 88 per cent of the clock and 0 per cent of the storage. Form and structure

A return is shorter than its first hearing

A rondo's refrain takes three fifths of the clock and a verse-and-chorus song is balanced to the bar. Count instead the bars a listener could not have predicted when they arrived, and the returns shrink to between a tenth and a quarter of what is kept — so a song equal by the clock is between three and seven times heavier in its first half. A coder that learns repeats one bar at a time says the halves are equal, and the two memories disagree by more than any proportion a listener could confuse.

A final chord stands above the impression for a fraction of a second. How far a final chord at the tutti's own level stands above the listener's running impression at the instant it is released, against how long it lasts, for four ways of arriving at it. Straight out of the tutti it stands above nothing at any length; after 1.2 s of silence the impression is 8.1 dB down, and the chord stands highest, 5.14 dB, when it lasts 54 ms; after 3.5 s of silence the impression is 23.7 dB down, and the chord stands highest, 14.42 dB, when it lasts 28 ms; after a 6 dB diminuendo the impression is 5.0 dB down, and the chord stands highest, 3.18 dB, when it lasts 42 ms. Every curve is level again by half a second, because the impression's attack of 99 ms catches the note's attack of 22 ms, so a held chord is released at the impression's level whatever preceded it. Form and structure

A final chord stands out for a twentieth of a second

A general pause drives a listener's running impression down, and the final chord that follows is supposed to cash the fall in. It cashes in at most two thirds of it. The note's own loudness rises with a 22-millisecond constant and the impression with a 99-millisecond one, so after a bar of silence the chord stands furthest above the impression 54 milliseconds in, by 5.1 of the 8.1 decibels the silence bought, and after 206 milliseconds the two are within a phon of each other. A short stamp spends most of its life standing out; a chord held a second and a half spends a seventh of it.

Four bars read by where the chords change, barline by barline. A constructed passage of four bars of eight quavers, its barline at the first slot and its chords C, F, Em, Dm. Notes: slot 1 C, slot 3 E, slot 4 G, slot 5 E, slot 6 C, slot 9 C, slot 10 C, slot 11 F, slot 13 A, slot 16 C, slot 17 B, slot 19 B, slot 20 E, slot 21 E, slot 24 G, slot 25 F, slot 29 F. For each of the eight places the barline could fall: as written metre, z 3.97, chords, z 2.00, change, z 4.30, metre + change, z 8.27; 1 quaver late metre, z -2.45, chords, z 2.14, change, z -0.87, metre + change, z -3.32; 2 quavers late metre, z -1.38, chords, z 2.15, change, z -0.54, metre + change, z -1.93; 3 quavers late metre, z -0.31, chords, z -0.27, change, z -2.64, metre + change, z -2.95; 4 quavers late metre, z 3.97, chords, z -1.63, change, z -4.30, metre + change, z -0.33; 5 quavers late metre, z -2.45, chords, z 1.27, change, z 0.87, metre + change, z -1.58; 6 quavers late metre, z -1.38, chords, z 1.18, change, z 0.54, metre + change, z -0.84; 7 quavers late metre, z -0.31, chords, z 1.05, change, z 2.64, metre + change, z 2.32. Best metre, z: as written and 4 late. Best chords, z: 2 late. Best change, z: as written. Best metre + change, z: as written. Harmony and voice leading

The chords mark the barline by changing there

Read bar by bar, the chords stood barely above chance at the barline and broke the metre's half-bar tie two times in three at best. Read instead by where they change — how different the chords are across a candidate's barlines against how different they are across the middle of its bars — the same notes break the tie right on 81 to 96 per cent of passages, and added to the metre they find the barline on up to 89 per cent against 61. The weakness was the question the old reading asked, not the harmony.

Level does not dilute the register's roughness, it multiplies it. The mean roughness of the I – vi – IV – V – I arrivals at 4 registers, each relative to the register as written, read three ways. Level-free, the bass is 8.6 times rougher than the treble. With every note at 70 dB it is 8.6 times, the same factor, because one level rescales every pair alike. With each chord played at the level that makes it as loud as the written register's chords — 82.8 dB −2 octaves, 75.5 dB −1 octave, 70.0 dB as written, 66.8 dB +1 octave — the bass is 343 times rougher than the treble, because roughness grows with the square of the pressure and the bass needs more of it to be heard at the same loudness. Harmony and voice leading

A rough arrival is rough because of its spacing

The pair the expectation essays report for every chord — how surprising it was, how rough its voicing is — has no level in it. Putting level back in answers the question it left open, and not the way it was framed. At one written dynamic the arrivals keep their order from 40 to 90 dB at three registers of four, and the bass stays 8.6 times rougher than the treble. Made equally loud, the bass has to be played 12.8 dB harder, and it is 343 times rougher: level does not explain the register's roughness away, it multiplies it.

The thirds' difference-tone line, note by note, against the dynamic. Each note of the line the difference tone f₂ − f₁ draws under a scale in just thirds, as its level above the higher of the threshold of hearing and the primaries' masking, against the level of the primaries. The product sits 50 dB below the primaries at 60 dB and grows twice as fast as they do. C2 above the limit from 73.5 dB; A1 above the limit from 76 dB; C2 above the limit from 73.5 dB; F2 above the limit from 70 dB; G2 above the limit from 68.5 dB; F2 above the limit from 70 dB; G2 above the limit from 68.5 dB; C3 above the limit from 66 dB. Intervals and chords

A combination-tone bass needs a forte

A scale in just thirds draws a diatonic bass line through its difference tones, and in sixths the cubic product draws one. Given the two published level laws, with their constants swept, the thirds' bass is not heard at all below primaries of about 66 dB and is heard whole only from 71 to 81. The cubic products are a different kind of object: the primaries mask them decibel for decibel as they rise, so no dynamic changes whether they are heard. Most of the thirds' inner line never is, and the sixths' bass needs a forte and a gentle law.

A listener who knows every metre recognises none of them inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps from the downbeat before every other one of the 1820 arrangements of twos and threes has been contradicted by the stream, on onsets alone, with the long beats accented and with the downbeat accented, against the 16 steps a present of 3.5 seconds holds. onsets alone: recognised within the present for 0 of 1820; long-beat accent: recognised within the present for 0 of 1820; downbeat accent: recognised within the present for 85 of 1820. The dashed line is the present. Rhythm and metre

Knowing every metre is slower than knowing none

A long aksak bar cannot be told from its shorter cuts by induction before one step into its last beat, and no accent carried by the notes moves that floor. The obvious escape is a listener who knows the repertoire and recognises the metre instead. Recognition among all 1,820 arrangements of twos and threes never beats the floor, is never quicker than induction, and is slower for half the metres: a nine induced in 9 steps is recognised in 27. What breaks the floor is a small repertoire that leaves out the metre's own longest cut — with the cut known, no repertoire of any size does.

With a fading memory the slowest order of the gaps 1 1 2 2 2 2 2 is not the one a perfect memory finds. For each of the 3 cyclic orders of the gaps 1, 1, 2, 2, 2, 2, 2 in 12 steps, the bits of position still unknown once listening has settled, against how many steps a listener's memory of a step takes to halve, with a mismatch costing 6. 2 2 2 1 2 1 2: 12 → 4e-9, 6 → 4e-5, 4 → 0.007, 3 → 0.070, 2 → 0.540, 1.5 → 1.130, 1 → 1.786; 2 2 2 2 1 1 2: 12 → 2e-5, 6 → 0.035, 4 → 0.335, 3 → 0.795, 2 → 1.405, 1.5 → 1.694, 1 → 2.005; 2 2 1 2 2 1 2 (the standard bell pattern): 12 → 2e-5, 6 → 0.027, 4 → 0.231, 3 → 0.535, 2 → 1.028, 1.5 → 1.368, 1 → 1.819. With perfect memory the slowest to locate is the standard bell pattern; at a half-life of 6 steps the highest floor is the order 2 2 2 2 1 1 2, and at 1 it is the order 2 2 2 2 1 1 2. Rhythm and metre

The bell pattern is slowest only to a perfect memory

Among the orders of its own gaps, a named timeline is usually both the most even and the slowest to locate — for a listener who never forgets. Give the listener a memory that halves and the result comes apart. Of six timelines slowest among their orders with perfect memory, only the fume-fume stays slowest for every forgetting listener, and the standard bell pattern, which is the fume-fume with onsets and rests exchanged and settles at exactly the same floors, is second of its three orders for every memory of half its cycle or less. The census ranking survives better, and in fourteen of twenty-one censuses it was the arithmetic of a pattern that repeats.

Come in part-way with the downbeat accented, and no bar of sixteen units or more is recognised inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps a listener who knows every arrangement of twos and threes needs to recognise the metre and where its bar begins, with the downbeat accented: coming in at a sample of steps inside the bar, against hearing it from its written downbeat. On onsets alone, or with the long beats accented, a metre entered part-way is never told from its rotations. 9: from inside the bar 10 to 17, 16 of 20 inside the present; from the downbeat 10 to 10; 10: from inside the bar 11 to 19, 12 of 20 inside the present; from the downbeat 11 to 11; 11: from inside the bar 12 to 21, 9 of 18 inside the present; from the downbeat 12 to 12; 12: from inside the bar 13 to 23, 8 of 24 inside the present; from the downbeat 13 to 13; 13: from inside the bar 14 to 25, 8 of 28 inside the present; from the downbeat 14 to 14; 14: from inside the bar 15 to 27, 4 of 28 inside the present; from the downbeat 15 to 15; 15: from inside the bar 16 to 29, 4 of 32 inside the present; from the downbeat 16 to 16; 16: from inside the bar 17 to 31, 0 of 32 inside the present; from the downbeat 17 to 17; 17: from inside the bar 18 to 32, 0 of 36 inside the present; from the downbeat 18 to 18; 18: from inside the bar 19 to 32, 0 of 36 inside the present; from the downbeat 19 to 19; 19: from inside the bar 20 to 37, 0 of 40 inside the present; from the downbeat 20 to 20; 20: from inside the bar 21 to 35, 0 of 40 inside the present; from the downbeat 21 to 21; 21: from inside the bar 22 to 40, 0 of 44 inside the present; from the downbeat 22 to 22; 22: from inside the bar 23 to 39, 0 of 44 inside the present; from the downbeat 23 to 23; 23: from inside the bar 24 to 44, 0 of 48 inside the present; from the downbeat 23 to 24; 24: from inside the bar 23 to 39, 0 of 48 inside the present; from the downbeat 23 to 25; 25: from inside the bar 22 to 44, 0 of 52 inside the present; from the downbeat 23 to 25. In all, 61 of 590 entries are recognised within the 16 steps of a 3.5-second present. Rhythm and metre

A dancer who comes in late needs the downbeat marked

Every window for recognising an aksak metre so far started at its written downbeat. A dancer joining a dance already going has not heard the downbeat, and the arithmetic of that is blunt: a metre entered part-way is, onset for onset, each of its own rotations heard from their downbeats, and the rotations are metres too — 2+2+3 and 3+2+2 are counted differently. So on onsets, and with the long beats accented, no metre is ever told from its rotations. Only an accented downbeat tells them apart, and with it a listener who knows thirty metres recognises 54 per cent of them inside the present from a random entry, against 1 per cent without.

Every product of every pair of partials is a harmonic of one absent note. A 4 : 5 interval on 261.6 and 327.0 hertz, just, each note carrying 6 partials at one over n, with the fundamental at 80 decibels. Every pair of partials makes its own difference tone, and there are 19 distinct frequencies among them. All of them are exact multiples of 65.41 hertz — the note neither instrument is playing — because partial j of a p·f₀ note and partial k of a q·f₀ note differ by (kq − jp)·f₀ whatever j and k are. The heavy line is what each product has to clear — the threshold of hearing at its own frequency, or the masked threshold the two primaries cast there, whichever is higher. 11 of the 19 get through. Intervals and chords

Three harmonics of the bass arrive before the bass

Every product priced until now was between two pure tones, and nothing that plays thirds is pure. Give each note a spectrum and the ear receives every pair of partials — and for a just interval p:q every one of their products is an exact multiple of the same absent fundamental. That crowd lands where the threshold of hearing is tens of decibels cheaper, so it names the bass at 71 decibels where the component at the bass's own frequency needs 74, and at 75 against 85 an octave lower. Tempered, the crowd still forms and names a note seventy cents flat.

The ghost bass drops when the passage gets louder. The note the whole crowd of products names, as a multiple of the fundamental the interval implies, against how loudly the interval is played. a major third: 2.9999999999999996 times the fundamental below 70 decibels and 1 times above it, a drop of 19 semitones; a minor third: 4 times the fundamental below 62 decibels and 2 times above it, a drop of 12 semitones; a fourth: 2 times the fundamental below 72 decibels and 1 times above it, a drop of 12 semitones. Softly, only the cubic products clear their thresholds, and they are an exact series on (2p − q) times the fundamental with no gaps in it. Loudly, the difference tones fill in the low harmonics, no template on the higher note can explain them, and the fit falls. Nothing about the interval has changed; the listener is simply being given a different subset of the same harmonic series. Intervals and chords

The ghost bass drops a twelfth at a forte

Both crowds arrive at once and every member of both is a multiple of the same absent fundamental, so a listener is never given a choice between them — only a different subset of one harmonic series at every dynamic. Softly, the subset is an exact gapless series on three times the fundamental. Loudly, the difference tones fill in the low harmonics and no template on the higher note survives them. Between 62 and 72 decibels, depending on the interval, the note the crowd names falls by an octave or a twelfth, and the two qualities of third cross at different levels.

A count is least reliable at both ends and best in the middle. How precisely a listener knows the length of a section they are counting, against how many units long it is, at three kinds of timing judgement. A slip — one unit miscounted, at 2% a unit — accumulates as a random walk, so its relative cost FALLS as the section lengthens. A lapse — the count lost altogether, at 1% a unit — compounds, so the chance of still having the count falls geometrically and a long enough section is certain to lose it. A listener who has lost the count is back to timing, so the two failures mix into a floor. Against a Weber fraction of 7.5% the count is worth most at 21 units, where it is 1.7 times finer than timing, and falls back under a quarter better by 98 units; Against a Weber fraction of 15% the count is worth most at 10 units, where it is 2.4 times finer than timing, and falls back under a quarter better by 101 units; Against a Weber fraction of 38% the count is worth most at 4 units, where it is 3.7 times finer than timing, and falls back under a quarter better by 102 units. The length at which it stops being worth much is nearly the same in all three, because it is set by the lapse rate alone. Form and structure

A count is not an estimate

Both established routes to a proportion are estimates — a duration timed, blurred by a Weber fraction, and a duration stored, biased by what was new. A listener who has induced a hypermetre has a third, and it is exact until it fails. It fails two ways that pull opposite: a slip miscounts one unit and its relative cost falls as the section lengthens, while a lapse loses the count entirely and its chance compounds. The mixture has a floor at about four units, where counting is 3.7 times finer than timing, and it is worth almost nothing past a hundred.

Given the bar in octaves, the cue gets the degree back. How often the reading names the right scale degree, against how much the bass cue is worth, under three rules for what a bass note rewards. the bass is the root: 68 per cent at best; the bass is some chord tone: 67 per cent at best; the bar in octaves: 92 per cent at best; the bass is the root, on a root line: 89 per cent at best. The published cue rewards the triad rooted on the bass, which is right only when the chord is in root position; rewarding any triad containing the bass is right always and rewards three chords a bar instead of one. The third rule is not a bass cue at all — given the bar voiced in octaves a listener knows which three pitch classes are sounding, and that names the chord outright. It reaches 92 per cent on a real line against the published cue's 68 on the same line, and the dashed curve is what that cue manages on a line made entirely of roots — 89 per cent. Scales and modes

Given the bar in octaves, the degree comes back

A bass note is a bare pitch class in the usual figure, and a register in any realisation anybody plays. Voice each bar in octaves and a listener knows which three pitch classes are sounding and which is at the bottom, which names the chord outright — and the rule that uses it reads the right scale degree in 92 per cent of bars on a real bass line, against 68 for the published cue on the same line and 89 for that cue on a line made entirely of roots. The question was whether the register recovers the 89. It recovers it and passes it.

A displaced map is displaced by the same amount everywhere. How far a listener's heard direction is displaced, in units of the smallest angular change they could detect at that azimuth, for four constant offsets added to every interaural delay. Each curve is flat. an offset of 5 microseconds is worth 0.33 just-noticeable steps at every azimuth; an offset of 10 microseconds is worth 0.67 just-noticeable steps at every azimuth, and past 88° hands the listener a delay their own head cannot produce; an offset of 20 microseconds is worth 1.33 just-noticeable steps at every azimuth, and past 86° hands the listener a delay their own head cannot produce; an offset of 40 microseconds is worth 2.67 just-noticeable steps at every azimuth, and past 82° hands the listener a delay their own head cannot produce. The reason is exact: differentiating Woodworth's curve gives a slope proportional to (1 + cos θ), so the angular displacement a fixed offset produces carries a factor of 1/(1 + cos θ) — and so does the smallest detectable angle, so the ratio has no azimuth in it. That is the opposite of a wrong head radius, whose displacement is zero on the median plane and grows toward the side. Perception and the listener

The error that moves straight ahead

The essay before this one found that a listener whose internal head is the wrong size makes no error at all on the median plane, and has to look hard to the side to catch it. Every head drawn here has its ears at equal radii, which makes the delay curve odd and every error a factor — and a factor cannot move a zero. Real heads are not symmetric. A constant offset of twenty microseconds displaces a listener's straight ahead by two and a quarter degrees, and it displaces every other direction by the same number of just-noticeable steps, exactly.

With 4 of six required at the end, the best schedule hands parts over. Six parts entering, leaving and re-entering a five-chord passage over four sounding players, in the walk through all 64 sets of sounding parts that makes the least audible entrance as audible as possible, with no memory of the chord before and at least 4 of the six sounding at the last chord. I, open: clarinet on G4 enters at 0.6 dB; vi, close below: oboe on E4 enters at 0.3 dB, and clarinet leaves; IV, close above: brass on E3 enters at -0.7 dB, voice on C5 enters at 2.2 dB, and oboe leaves; V, bracketing: violin on G5 enters at 7.2 dB; I, hollow: flue pipe on C6 enters at 4.0 dB. The least audible entrance is -0.72 dB against -2.65 for the best schedule in which nobody leaves; the walk has 2 exits and 6 entrances. Form and structure

An exit is worth nothing until the tutti is given up

Six parts entering a five-chord passage have a best schedule when each enters once and stays, and letting parts leave and come back was supposed to improve it. Searched over every set of sounding parts at every chord, it improves it by exactly nothing, with or without the chord before still masking — as long as all six must be playing at the end. Let one part be missing from the final chord and the weakest entrance gains 1.4 decibels; let two be missing and it gains 1.9, by a relay in which the parts with least room come in, are heard for one chord, and give way.

Checkpoints sharpen the middle of a form and leave its top vague. A piece of 480 seconds divided 7 times, with how many proportions between 1 : 1 and 3 : 1 a listener tells apart at each level: timed, counted in 2-second units, and counted with a second count of 16-second phrases that can mend a lapse in the first. 480 s: 2.1 timed, 2.2 counted, 3.9 with 0 per cent of lapses shared and 2.5 with 50 per cent of lapses shared; 240 s: 2.1 timed, 2.5 counted, 7.1 with 0 per cent of lapses shared and 3.1 with 50 per cent of lapses shared; 120 s: 2.1 timed, 3.1 counted, 13.2 with 0 per cent of lapses shared and 4.0 with 50 per cent of lapses shared; 60 s: 2.1 timed, 4.1 counted, 19.4 with 0 per cent of lapses shared and 5.4 with 50 per cent of lapses shared; 30 s: 2.1 timed, 5.4 counted, 18.1 with 0 per cent of lapses shared and 7.2 with 50 per cent of lapses shared; 15 s: 5.2 timed, 12.2 counted, 12.2 with 0 per cent of lapses shared and 12.2 with 50 per cent of lapses shared; 7.5 s: 5.2 timed, 10.3 counted, 10.3 with 0 per cent of lapses shared and 10.3 with 50 per cent of lapses shared. With the two counts failing independently, the level of 60 seconds goes from 4.1 to 19.4, and the whole piece only from 2.2 to 3.9. Form and structure

Checkpoints sharpen the middle of a form, not its top

A listener who counts bars loses the count somewhere in a long section and is thrown back on timing the whole of it. A listener who also counts phrases can mend the lapse at the last phrase. If the two counts fail independently, the level a minute long goes from four distinguishable proportions to nineteen; the whole eight-minute piece goes only from two to four, because thirty phrases are long enough to lose a count as well. And if a fifth of lapses take both counts at once, three quarters of the gain is gone.

The content errs slow and the bass errs fast. Passages of eight bars built at three harmonic rhythms, each with a bass that states every new chord's root and moves to another chord tone on a beat 50 per cent of the time. Two readings are asked the chord rate: one from how much the pitch-class content changes across a grid, one from how completely the bass's moves land on it. At two chords a bar the content reading names the rate 63 per cent of the time, too fast 0 and too slow 37; the bass reading 100, 0 and 0. At one chord a bar the content reading names the rate 67 per cent of the time, too fast 0 and too slow 33; the bass reading 18, 80 and 2. At a chord every two bars the content reading names the rate 98 per cent of the time, too fast 2 and too slow 0; the bass reading 0, 100 and 0. The two readings miss in opposite directions and at opposite ends of the tempo range: the content reading at fast harmonic rhythms, by naming a multiple, and the bass reading at slow ones, by naming its own arpeggiation. Harmony and voice leading

The bass errs fast where the content errs slow

Asked how often the chords change, a reading built on pitch-class content names a slower multiple and never a faster rate. Give the passage a bass that states each new root and moves between chord tones inside a chord, and a reading built on the bass's moves errs the other way: it names a faster grid and never a slower one. At two chords a bar the bass is right every time; at a chord every two bars it is never right. Six ways of combining the two readings each trade one end of the range for the other.

Put back beside its notes, the crowd names the bass at every dynamic. A just major third on complex tones, drawn on one axis of harmonic numbers of the fundamental its ratio implies: the partials of the two played notes, and the products of those partials that clear threshold, at 55 and 80 dB. At 55 dB the products alone name 3 times the fundamental with 0 empty slots; the products and the notes together name 1 times it with 4 empty slots, and the notes alone name it with 9. At 80 dB the products alone name 1 times the fundamental with 0 empty slots; the products and the notes together name 1 times it with 0 empty slots, and the notes alone name it with 9. The soft reading on a higher note exists only when the loud notes are set aside. Taken together, the products do not decide which fundamental is named; they decide how many holes its template has. Intervals and chords

The played notes already name the ghost bass

The products of a just third's partials, fitted on their own, name a note a twelfth above the bass when the interval is soft and drop to the bass when it is loud. Put the two played notes back beside them and the drop disappears: the notes and their products name the bass at every dynamic, because the notes' own partials are harmonics of it already. What the dynamic changes is not which note is implied but how complete its harmonic series is — nine holes from the notes alone, four when soft, none when loud.

The ceiling is thirteen bars, and every one of them has a name. Every bar of the six schemes read by the key-finder at a key cost of 3 and a bass weight of 3, with the chord in every bar given exactly — no segmentation is involved. 175 of 188 bars have both the key and the degree right, 93.1 per cent. The misses: 7 in the thirty-two-bar song's bridge, where the chain III7 is read in E major, III7 is read in E major, VI7 is read in E major, VI7 is read in D major, II7 is read in D major, II7 is read in D major, V7 is read in D major; 4 bars of the rondo's A minor episode read as C major; 2 next to a change of key; and 0 of any other kind. Scales and modes

The ceiling is thirteen bars with names

The key-finder's two cues stop buying anything at about 93 per cent of scheme bars read right, and the obvious suspect was the chord segmentation feeding it. The reading was never given a segmentation: every bar arrives with its true chord. What the ceiling is made of can be listed instead, and it is thirteen bars of 188 — seven in a bridge of secondary dominants, four in a minor episode whose chords C major also owns, and two at the edges of a modulation. No weight of any cue moves one of them.

A listener who hears only the landmarks recognises the metre sooner. The share of metres recognised within the 16 steps of a 3.5-second present after coming in at a random step, against how many metres the listener knows, for five listeners: every onset with nothing marked, every onset with the long beats accented, only the downbeats and long beats with nothing marked, every onset with the downbeat accented, and only the downbeats and long beats with the downbeat marked. Every onset, nothing marked: 5 known, 49% within the present, 3% never; 10 known, 19% within the present, 8% never; 30 known, 1% within the present, 17% never; 100 known, 0% within the present, 34% never. Every onset, long beats accented: 5 known, 68% within the present, 3% never; 10 known, 39% within the present, 8% never; 30 known, 6% within the present, 17% never; 100 known, 0% within the present, 34% never. Landmarks only, nothing marked: 5 known, 69% within the present, 0% never; 10 known, 61% within the present, 3% never; 30 known, 26% within the present, 3% never; 100 known, 9% within the present, 9% never. Every onset, downbeat accented: 5 known, 88% within the present, 0% never; 10 known, 74% within the present, 0% never; 30 known, 54% within the present, 0% never; 100 known, 14% within the present, 0% never. Landmarks only, downbeat marked: 5 known, 91% within the present, 0% never; 10 known, 82% within the present, 0% never; 30 known, 56% within the present, 0% never; 100 known, 23% within the present, 0% never. Rhythm and metre

A late dancer needs the landmarks, not the rhythm

A listener who joins an additive-metre dance part-way recognises it far more reliably with the downbeat accented. Strip the stream down to its landmarks — the onsets that begin a bar or a long beat, with every other onset removed — and the listener does as well or better: knowing a hundred metres, 23 per cent are recognised within the present against 14 with every onset. Unmarked, the landmarks still beat every onset with the long beats accented. Neither half does it alone; what identifies a metre from a late entry is where its long beats sit relative to its bar.

Against a pulse, the bell pattern is the quickest of its orders to place. The bits of position a listener is still missing, averaged over the first cycle heard, for each cyclic order of the gaps 1 1 2 2 2 2 2 in 12 steps, heard alone, against a pulse every three steps and against a pulse every four, at perfect memory, half-life 3 steps, half-life 1.5 steps. 2 2 2 1 2 1 2: alone 1.08, 1.23, 1.73; against a pulse every 3 steps 0.69, 0.75, 0.96; against a pulse every 4 steps 0.63, 0.67, 0.85. 2 2 2 2 1 1 2: alone 1.22, 1.63, 2.02; against a pulse every 3 steps 0.60, 0.73, 0.92; against a pulse every 4 steps 0.83, 1.07, 1.36. 2 2 1 2 2 1 2 (the standard bell pattern): alone 1.25, 1.45, 1.82; against a pulse every 3 steps 0.54, 0.56, 0.67; against a pulse every 4 steps 0.55, 0.57, 0.68. Alone, the bell pattern is not the quickest order to place at any memory. Against either pulse it is the quickest at every memory. Rhythm and metre

Against a pulse the bell pattern is the easiest to place

Heard alone, the standard bell pattern is not the quickest order of its own gaps to place in its cycle, for a listener with any memory. Heard against a pulse every three steps or every four — which is how anyone hears it — it is the quickest, at every memory and at every alignment of pulse and bell, and by a wide margin: at a memory of a quarter of the cycle, 0.56 bits unplaced over the first cycle against 0.73 for either rival against a pulse in threes. Six of eight named timelines do the same. A timeline's order of gaps looks chosen for how it sits against the beat, not for how it sounds alone.

Counted over what arrives, the balanced bass is not the roughest register. The mean roughness of the I – vi – IV – V – I arrivals with each chord played as loud as the written register's, relative to the written register, counted over every partial and over the partials that stand above what the rest of the chord masks. Every partial: 70 −2 octaves, 7.48 −1 octave, 1.00 as written, 0.20 +1 octave. Delivered partials only: 6e-9 −2 octaves, 2.83 −1 octave, 1.00 as written, 0.19 +1 octave. Over every partial the lowest register is 343 times rougher than the highest; over what arrives it is the smoothest of the four, and the roughest is −1 octave, 2.8 times the written register. Perception and the listener

A bass chord low enough to balance has already hidden its tenor

Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.

Holding the bass through one mid-bar change in six finds the barline five times in six. Passages of eight bars at two chords a bar, with the bass arpeggiating on 50 per cent of its beats, and each barline reading's share of passages it places correctly, against the share of mid-bar chord changes voiced over the bass already sounding. 0% held (convention strength 0): metre then bass 50%, bass alone 13%, metre alone 50%, chord changes alone 10%; 4% held (convention strength 0.1): metre then bass 62%, bass alone 34%, metre alone 50%, chord changes alone 10%; 6% held (convention strength 0.2): metre then bass 68%, bass alone 44%, metre alone 50%, chord changes alone 15%; 9% held (convention strength 0.3): metre then bass 75%, bass alone 57%, metre alone 50%, chord changes alone 18%; 15% held (convention strength 0.4): metre then bass 82%, bass alone 68%, metre alone 50%, chord changes alone 13%; 16% held (convention strength 0.5): metre then bass 85%, bass alone 74%, metre alone 50%, chord changes alone 14%; 20% held (convention strength 0.6): metre then bass 88%, bass alone 79%, metre alone 50%, chord changes alone 11%; 26% held (convention strength 0.7): metre then bass 92%, bass alone 86%, metre alone 50%, chord changes alone 13%; 27% held (convention strength 0.8): metre then bass 92%, bass alone 86%, metre alone 50%, chord changes alone 13%; 28% held (convention strength 0.9): metre then bass 95%, bass alone 92%, metre alone 50%, chord changes alone 10%; 33% held (convention strength 1): metre then bass 96%, bass alone 93%, metre alone 50%, chord changes alone 10%. The metre ties the barline with the half-bar and the chord changes are at chance, since the chords change at both; the bass's holds are the only evidence that separates them. Harmony and voice leading

A bass that holds through a change marks the barline

At two chords a bar the chords change on the barline and on the half-bar alike, so a reading of where they change is at chance, and the metre ties the two. The bass has one more piece of evidence: a change inside the bar can be voiced over the note already sounding, and a change on the barline is voiced over its root. Hold the bass through one mid-bar change in six and the metre and bass together place the barline in 85 per cent of eight-bar passages; one in three, 97. The convention cannot be stronger than that, and the chord changes, asked first, only get in the way.

All themes