Theme

The limits of the listener

The quietest audible tone, the smallest audible difference, the fastest series that can still be a beat, the delay at which one sound becomes two. Each is a measured boundary with a published source and a spread between listeners, and each one bounds a claim made somewhere else on this site.
Every point on one curve sounds equally loud. The equal-loudness contours of ISO 226:2003, evaluated from the standard's own parameters. The lowest curve is the threshold of hearing. Because the curves are not parallel — they crowd together in the bass and spread apart in the middle — the same change in decibels is a different change in loudness at every frequency, and a spectrum that was balanced at one level is not balanced at another. Perception and the listener

The quietest thing audible, and why the volume knob is a tone control

A decibel is a fact about air. A phon is a fact about a listener, and the two do not line up — the map between them bends with frequency, and it bends differently at every level. One consequence is that turning a piece of music down does not turn all of it down equally, and the amount by which it does not is a number.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward. Perception and the listener

One sound hides another, and it hides upward

A tone can be made completely inaudible by a second tone that is nowhere near it in frequency, and the region it disappears into is lopsided. Masking spreads up the spectrum and barely down it, and the reach grows with the masker's level — so which line in an arrangement vanishes is a prediction, not a matter of taste.

The threshold before, during and after a burst. The level a brief probe needs in order to be heard, plotted against when it happens relative to a 70 dB burst that occupies the shaded band. To the right is forward masking, which decays over about 200 milliseconds. To the left is backward masking: the threshold is raised for a probe that has already finished before the masker begins. Perception and the listener

A sound hides what came before it

Masking does not stop when the masker does. A loud sound raises the threshold for about a fifth of a second after it ends, which is unremarkable, and for several milliseconds before it begins, which is not. The auditory present is a window rather than an instant, and inside the window the order of events is not the order they arrived.

A 7-semitone sequence at 120 ms a tone. Tones drawn as pitch against time, one bar per tone. The events are the same in both readings of this pattern; what changes is whether a listener assigns them to one line that leaps back and forth or to two lines that each stay put. Nothing in the drawing decides which, and nothing in the sound does either. Perception and the listener

The ear builds objects, and sometimes offers a choice

What arrives at an ear is one pressure signal. What a listener gets is a set of separate things — a violin, a voice, a car outside. The assignment is a construction, and the clearest evidence is that it can be flipped by changing nothing but the speed: one sequence of tones is a single line when slow and two lines when fast, with a wide region in between where the listener may choose.

A source 45° off centre, and the path difference it makes. A head from above with a source to one side. The near ear is reached first; the far ear's path runs round the head, and the difference between the two is 13.1 centimetres, which at 343 metres a second is 381 microseconds. That number, and the level difference the head's shadow produces, are the whole of what the ear has to work with. Perception and the listener

Two ears, and the whole of the difference is 655 microseconds

Direction is computed from two numbers — when a sound reaches each ear and how loud it is at each — and which of the two is usable is decided by the wavelength against the width of a head. The changeover frequency is not a design choice. It falls out of 343 metres a second and 17.5 centimetres, and it is why the mechanism of hearing where something is changes halfway up the piano.

How late a reflection has to be before it is an echo. What a single reflection does to the sound it follows, against its delay, on a logarithmic axis. Under a millisecond the two combine into one image that is pulled towards the earlier source. From there out to a few tens of milliseconds the reflection is not heard as a separate event at all and does not move the image — it only changes the timbre. Past the echo threshold it becomes a second sound, and the threshold is five times later for speech than for a click. Perception and the listener

The first wavefront wins

A room sends a hundred copies of every note to a listener from a hundred directions, and the listener hears one note in one place. The mechanism that does it is brutal and simple: for the first few tens of milliseconds after a sound arrives, everything that follows is denied a vote on where it came from — even when it is louder than the original.

The smallest audible difference, and what has to clear it. The difference limen for frequency, converted from Wier, Jesteadt and Green's 1977 fit into cents, against the intervals and commas the rest of these essays argue about. Anything drawn below the curve is a quantity nobody can hear as a change of pitch; anything well above it is a quantity a listener can be asked about. The limen is for pure tones, successive, with trained listeners — the most favourable case there is, and therefore the right one to test a claim against. Perception and the listener

How small a difference is audible

Every essay here about tuning has assumed a listener who can hear the difference between two systems. The assumption has a number: about five cents in the middle of the range. It clears the two commas fourfold and it does not clear the schisma at all, which sorts the whole subject of tuning into the part that is about music and the part that is about arithmetic.

Three octaves, and none of them is 2:1. How far above an exact doubling the upper note of an octave is set, against frequency. The listener's octave is measured with pure tones, which have no partials to beat against each other, so nothing about a stiff string can account for it. The piano's stretch is a different quantity with a different cause, and the two are drawn together only so that the difference is visible. Perception and the listener

The octave that is not two to one

The octave is the one interval nobody argues about: two to one, exact, in every tradition that has one. Asked to set an octave by ear, listeners set it wide — and they do it with pure tones, which have no partials to beat against each other. Whatever is stretching the octave, it is not the stiffness of a piano string.

How much bow force is allowed, and where. Schelleng's diagram. The lower bound is the least force that will trigger a slip on every pass of the corner and goes as one over beta squared; the upper bound is the most the string will take before it sticks for more than a period and goes as one over beta. At beta = 0.09 the usable range spans a factor of 9.0; at 0.03, near the bridge, it is 3.0, and at 0.2, over the fingerboard, 20.0. The window closes in proportion to beta, so the difficulty of playing near the bridge is a slope on this picture rather than a matter of opinion. Instruments and their design

How much bow is allowed

Too little force and the corner fails to trigger a slip on every pass; too much and the string sticks for more than a period. Both bounds depend on where the bow is, and they depend on it differently — one as the square of the distance from the bridge and one linearly — so the window between them closes in proportion as the bow approaches the bridge. Sul ponticello is difficult by a power law.

A phrase is a number of seconds, and the bars follow the tempo. Phrase durations for 1, 2, 4, 8, 16-bar phrases at seven tempos, on a logarithmic seconds axis, with the 2 to 8 second window shaded. The window is a property of the listener and does not move; which bar count falls inside it is decided entirely by the tempo. Form and structure

A phrase is a number of seconds

Musical phrases are described in bars, and four is the number everybody names. But the constraint that fixes a phrase is a property of the listener and is measured in seconds, so the bar count is whatever the tempo makes it. Across seven ordinary tempos the bar count that lands inside the window moves by a factor of eight, while the window itself does not move at all.

The same induction, one level up. Bar-level onsets from thirty-two-bar AABA — a bar is marked where a section or a key begins — scored against hypermetres of 2, 3, 4, 6, 8 bars with the identical function the beat-level figures use. The best-fitting period is 8 bars, which at 108 beats a minute lasts 17.8 seconds. Form and structure

The bar above the bar

A four-bar group is a bar whose beats are bars. That is not an analogy — it is the same computation, and the same metre-induction model produces one when it is handed bars instead of beats, unchanged. What decides where the hierarchy of levels stops is not in the arithmetic at all, and it is a number the phrase essay already measured.

Nine commas, and the line under which none of them matters. Each named comma at its true size in cents, against the difference limen of 3.87 cents at 500 Hz — the smallest change of frequency a listener can detect, computed from the same formula the perception essays use. One comma falls below it, and it is the only gap in the subject that no tuning system has to do anything about. Pitch and tuning

A comma under the threshold

Eight pure fifths taken downwards arrive at a major third 1.95 cents flat of a perfect 5:4. That gap has a name and a history, and it is the only one in the subject nobody has ever had to hide — it is under the smallest difference a listener can detect, which makes a chain of untempered fifths a source of almost-just thirds and makes one eighteenth-century tuning nearly free.

Cycles of roughness completed in 125 milliseconds. For each interval and each register, how many cycles of its own beating fit inside a note 125 milliseconds long. Cells at or above 4 cycles are the ones a listener has time to hear as rough; the pattern runs the opposite way from roughness itself, which is largest in the bass. A short dissonance low down is the case where the two disagree. Intervals and chords

A dissonance has to last

Roughness is a fluctuation, and a fluctuation needs cycles. A minor second at the bottom of a cello fluctuates thirty-three times a second, so a semiquaver holds four of them and a demisemiquaver holds two — which turns the counterpoint rule that a dissonance may pass if it is short into a number, and puts that number at about seventy milliseconds through most of the range, once the pairs beating too slowly to be roughness at all are taken out of the average.

Where each room stops being a set of resonances. The Schroeder frequency of 6 rooms — a carpeted bedroom at 217 Hz, a domestic living room at 183 Hz, a rehearsal room at 120 Hz, a jazz club at 68 Hz, a shoebox concert hall at 21 Hz, a gothic cathedral at 36 Hz — marked on a logarithmic frequency axis with the ranges of 4 instruments underneath. Below the mark a room is a handful of separable modes and a note's loudness depends on where the listener is standing; above it the modes overlap and the room is described by one decay time. Timbre and acoustics

Where a room stops being a room

A small room has frequencies it supports and frequencies it will not, and a hall has a reverberation time. Those are two separate accounts and they are descriptions of the same building at different frequencies. The crossover is one formula, and in a bedroom it lands at about two hundred hertz — in the middle of the bass register, and above nothing at all in a concert hall.

Direct and reverberant sound in a shoebox concert hall. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 5.5 metres in a room of 18700 cubic metres with a 2-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument. Perception and the listener

How far away the room takes over

Direct sound falls six decibels every time the distance doubles and the reverberant field does not fall at all, so the two cross once. In a concert hall the crossing is at about five and a half metres, which is nearer than nearly every seat — so almost everybody in almost every hall is hearing the building more than the players, and the number that says so is built from two quantities already computed — a room's reverberation time and an instrument's directivity.

The period, as the piece goes by. The strongest lag of thirty-two-bar AABA computed on only the bars heard so far, against how many bars that is. The final answer is 4 bars; it is revised 6 times on the way, and is not reached for the last time until bar 29 of 32, which is 91 per cent of the way through and 64 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have. Form and structure

The form a first hearing cannot have

Every figure so far was computed with the whole piece in hand. Run the same methods over only the bars already heard and one of the two methods survives intact — the boundary operator turns out to be causal at a fixed delay of a few bars — while the other collapses. The period of a piece is not knowable until the piece is nearly over, and in two of the six schemes here not until its last bar.

The ranking is settled either side of one narrow band. Remembered repetition — each bar's best match to an earlier bar, discounted by exp(−Δt/τ) with Δt in seconds — for 6 schemes at 108 beats a minute, against the decay constant τ on a logarithmic axis. The order of the schemes changes only between 8 and 13 seconds; outside that band it is fixed, so an estimate of τ wrong by any amount that stays outside it leaves the ranking alone. Perception and the listener

A return has to be remembered

A stripe four bars off the diagonal and a stripe twenty-four bars off it are the same ink and are not the same experience. Convert the lag axis to seconds, discount every comparison by how long ago it was, and the ranking of these six schemes by how repetitive they are changes — and the decay constant and the tempo turn out to enter the arithmetic as one number rather than two.

Four voices, placed by the arithmetic. The rules in force are: no parallel octaves; no parallel fifths; no voice crossing; no gap over an octave above the tenor; the leading note is not doubled; the leading note resolves, outer voices; no augmented melodic interval. The four parts are drawn lowest to highest — bass, tenor, alto, soprano. I – vi – ii – V – I in C major, realised in four voices by the cheapest set of voicings obeying 7 rules, at 16 semitones of motion in all. Every gap between adjacent voices narrower than the fission boundary of 5.2 semitones is marked, and below that boundary two parts cannot be heard as two however hard a listener tries. Perception and the listener

A voice is a stream, and the ear decides which

Seven earlier essays have assigned voices to notes. Whether a listener follows the assignment is a separate question with laboratory numbers attached, and the numbers are unkind to it: two parts closer than about five semitones cannot be heard as two at any speed, a third of the gaps in the cheapest four-part writing are inside that limit, and in a third of chord changes the ear's own rule for continuing a line does not recover the parts as written.

Which harmonics of a 200 Hz note arrive one to a filter. One row per harmonic of a 200.0 Hz tone. The bar is the ear's analysis band at that harmonic on the equivalent rectangular bandwidth model, and the two small marks either side are the neighbouring harmonics. A harmonic is counted as resolved when the spacing to its neighbours, 200.0 Hz, exceeds that bandwidth, and 8 of 12 are. The third to fifth harmonics are picked out because published measurements put the pitch's dominance region there; nothing in this drawing derives that. Intervals and chords

Which harmonics carry the pitch

A missing fundamental is inferred from a pattern, and a pattern has to be legible before it can be matched. Counting how many harmonics of a note land in separate auditory filters prices the inference — and the answer at the bottom of a bass guitar's range is none of them.

A 200-a-second click train, correlated with itself. The autocorrelation of a click train at 200 a second, smoothed by the ring of an auditory filter centred at 4000 Hz — an equivalent rectangular bandwidth of 456 Hz, so a ring of 2.2 ms. The regular train peaks at 5.0 ms, one period. With each click displaced by a standard deviation of 20 per cent of the period — 1.00 ms — the peak's contrast against the surrounding lags falls from 0.41 to 0.12. The average rate and the long-term spectrum are unchanged by the jitter; only the timing is. Perception and the listener

A pitch with nothing to match

Filter a click train into a band where no partial is separable from its neighbours and it still has a pitch at its repetition rate. Displace each click by a fraction of a millisecond, leaving the average rate and the long-term spectrum exactly where they were, and the pitch goes. The mechanism is reading the timing — which bounds the account endorsed here from the start.

Roughness and loudness of one interval against level. A minor third on C4 evaluated at every level from 20 to 100 dB SPL per note, with both quantities drawn relative to their own value at 60 dB. Roughness is the Plomp–Levelt sum over the partials that are above ISO 226's threshold at that level; loudness is the same partials in sones. Between 50 and 95 dB the interval grows 3.2 × 10⁴ times rougher and 24.9 times louder, so roughness grows like loudness raised to the power 3.2. Intervals and chords

The same chord is harsher when it is louder

Every roughness number so far was computed at a level nobody stated. Roughness is the product of two partial amplitudes, so it is quadratic in pressure, while loudness is compressive — which makes a minor third at middle C thirty-two thousand times rougher at fortissimo than at pianissimo and only twenty-five times louder. A chord has no single consonance to report.

One pattern, four metres. The same 16-step onset pattern read under 4 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are downbeat on step 1 0, downbeat on step 2 -12, downbeat on step 3 -12, downbeat on step 4 0, so downbeat on step 1 and downbeat on step 4 tie and the model does not choose. Nothing about the sound differs between these readings; the bar line is supplied by the listener. Perception and the listener

The beat that is never sounded

A listener who has heard four bars of a groove and then hears two bars with the downbeats taken out does not move the downbeat. This site's rule set does, every time, on every pattern tried — and the direction it moves in says exactly what kind of model would be needed instead.

A tracker following a tempo that will not stay still. The tracker's error as a fraction of a beat, against beat number, for 3 rates of tempo change with a period-correction gain of 0.2. It settles at 0.025 of a beat behind at 0.5 per cent a beat, 0.092 of a beat behind at 2.0 per cent a beat, 0.206 of a beat behind at 5.0 per cent a beat. It does not lose the beat; it lags, by very nearly the rate divided by the period-correction gain, and the lag reaches a quarter of a beat at 6.3 per cent a beat — at which point the tracker is nearer the wrong onset than the right one. Form and structure

A metre has to be able to change its mind

Replace the scoring function with a phase-corrected oscillator and the two failures reported earlier separate. The beat now survives two silent bars, drifting 23 milliseconds of a 560-millisecond beat. But it does not lose a moving tempo so much as lag behind it, by the rate over the correction gain — and a cadential ritardando that halves the tempo in eight beats is faster than the model can follow.

The swing ratio, against the categories it passes through. The same swing curve read against the boundaries between duration categories rather than against notated values. A category's centre is a simple ratio — 1:1, 2:1, 3:1 — and the boundary between two of them is the midpoint, which is arithmetic. The curve crosses 2 of them: out of 2:1 and into 1:1 at 240 beats a minute, out of 3:1 and into 2:1 at 171 beats a minute. So the same notated figure is, by the categorical criterion, a different rhythm at each end of an ordinary tempo range, and the notation says triplet feel throughout. Perception and the listener

How late is a different note

A deviation of thirty milliseconds is expression and a deviation of two hundred is a wrong note, so there is an edge. The edges in time are arithmetic — the midpoints between the simple ratios — and the swing ratio crosses two of them as the tempo rises, at 171 and at 240 beats a minute, while the notation says triplet feel throughout.

Roughness across an octave. Sensory dissonance between two complex tones as the upper one is swept through an octave, computed by summing the roughness between every pair of their partials. Nothing here is placed by hand. The wells this spectrum produces sit on 4/3, on 3/2, on 5/3, found by scanning the curve rather than by marking them. And the tempered minor third sits 198 cents below the nearest well, and the tempered major third sits 98 cents below the nearest well, which is a real feature of this model and not a defect of the drawing. Intervals and chords

The third the model has no opinion about

A neutral third — halfway between major and minor, and a scale degree in most of the music between Morocco and Iran — is nothing at all on an ordinary roughness curve. It becomes a well at 347 cents, which is exactly 11:9, only for a spectrum whose ninth and eleventh partials are at least three times their natural strength. No instrument has that spectrum.

Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement. Perception and the listener

Three answers to how finely a pitch can be heard

Two notes one after the other are told apart at about four cents at A440. Whether a melodic interval is in tune is a judgement an order of magnitude coarser. And two notes held together are heard to beat at a third of a cent, because the question is answered by counting rather than by hearing pitch at all. Every equal division ever built sits between the coarsest and the finest.

A note sung at 440 Hz, drawn in cents. Deviation from the notated pitch against time, for a vibrato of ±71 cents at 6.0 cycles a second — Seashore's and Prame's measured values, which agree. The note is 142 cents wide, and the shaded bands across the middle are the syntonic comma at 21.5 cents, the Pythagorean comma at 23.5 cents, the difference limen at 440 Hz at 4.0 cents. The pitch a listener reports is near the middle of the excursion rather than at either edge — and not exactly at the middle either: because cents are logarithmic and frequency is not, the mean frequency of this trace sits 0.73 cents above the centre line. Pitch and tuning

A note that is never at its pitch

Eleven essays in this field argue about differences of one to twenty-four cents. An ordinary operatic vibrato is a hundred and forty cents wide and completes six excursions a second, so every one of those distinctions fits inside a single sung note several times over — and the beat rate a tuner would null passes through zero eleven times a second.

The same intervals, played higher and higher. Sensory roughness for four fixed intervals as the pair is transposed up five octaves, computed from the same Plomp–Levelt model as the dissonance curve. Every one of them falls as it rises, and the small intervals fall furthest — so how consonant an interval is depends on where it is played, not only on what it is. Intervals and chords

A chord is a register

The same three pitch classes are five times rougher at the bottom of a piano than in the middle, and the arrangement that minimises the roughness over a low bass turns out to be the bass's own fifth and sixth partials. The orchestration rule about low thirds is not a convention. It falls out of the width of a critical band, exactly.

3 ways to fill the same fourth. Each row is a tetrachord: a span of 498 cents — a pure fourth — with two notes inside it, drawn in cents from the lower bound. The bounding notes never move, which is what makes the tetrachord a unit; everything that varies is interior. diatonic (ditonic) puts them at 90 and 294 cents, giving steps of 90, 204, 204; intense chromatic puts them at 81 and 231 cents, giving steps of 81, 151, 267; enharmonic puts them at 63 and 112 cents, giving steps of 63, 49, 386. Against the twelve equal steps below, 1 of 3 land within twenty-five cents of a semitone at every degree; the worst miss is 37 cents, which is a quarter of a semitone and has no name on a keyboard. Nothing in this construction is a subset of an equal division, so nothing the census here does applies to any of it. Scales and modes

A scale built downward from a fourth

A tetrachord is a bounded span with two free notes inside it, which is a different generator from a chain of fifths and a different object from a subset of an equal division. At the maqam tradition's own quarter-tone resolution the construction admits thirty-six fillings of the fourth and 1,296 octaves built from them; the census here can see six of the first and 2.8 per cent of the second, and one of the octaves it cannot see is an ordinary mode of a living tradition.

Ode to Joy as a path. Ode to Joy plotted as 30 notes against the 8 scale degrees it uses, one column per note. Its largest melodic interval is 2 semitones and it spans 7; the mean absolute step is 1.24 semitones. Beethoven, Ninth Symphony, finale, 1824 — the theme as first stated, eight bars. Scales and modes

A melody is a walk, not a set

Nine essays here are about which seven of the twelve a scale takes, and every one of them describes a set. A tune is not a set; it is a path across one, and the path is nearly all small steps. That is not a matter of taste. Above about eight notes a second the ear stops being able to hold a large interval and a small one in the same line, and at sixteen the choice disappears altogether — so a fast passage is scalar because a fast passage that leaps is two pieces of music.

How much of the rule a walk with no rule reproduces. Post-skip reversal in 20,000-note random walks with no melodic knowledge of any kind. An unbounded walk reverses after 50.0 per cent of leaps, which is the chance rate and is the check that the measurement is right. Confining it to 12 semitones raises that to 61.3 per cent. Reaching the 70 per cent that corpus studies report needs a central tendency of 0.95 — an almost deterministic pull back toward the middle at the edges of the range. A wall is not enough; there has to be a spring. Form and structure

The leap that pays itself back

Every melody textbook teaches that a leap should be followed by a step in the opposite direction, and every corpus that has been counted agrees — around seven leaps in ten are answered that way. A random walk with two walls, no memory of the leap and no rule of any kind reverses after 61 per cent of them, and the residue is not a rule either. What is left when the walls are accounted for is a prediction the rule does not make, and it is the prediction that decides between them.

The gap that decides whether two clocks are two. The closest approach of the two streams in each ratio, at 100 to the minute in 4-time — a bar of 2.4 seconds. It is the bar divided by the product of the two numbers, which is an identity and is checked here against the measured minimum. 9:5 is the last ratio whose onsets are securely separate at this tempo; past it the two streams' events fall inside the window in which the ear cannot put two onsets in order, and what is heard is one irregular pattern rather than two clocks. The bound has a product in it, which is why 3:2 and 4:3 are everywhere and 11:7 is a notation. Rhythm and metre

The ratio that stops being two

Eight earlier essays have taken two clocks at a rational ratio to be a thing a listener can hold. There is a ratio past which it is not, and the bound is not where anyone would look for it: the cycle is exactly one bar long at every ratio, so the length of the pattern separates nothing. What separates them is the closest the two streams ever come, which is one part in their least common multiple — 400 milliseconds for three against two, and seventeen for thirteen against eleven, which is not two events at all.

The same steps, counted by the clock. The step distribution of Ode to Joy and Twinkle, twinkle counted two ways: once per interval, which is what every earlier figure did, and once weighted by how long the note it leaves is held. The two disagree because a tune's long notes are not distributed evenly over its interval sizes — in Ode to Joy the 2-semitone step is 55.2 per cent of the moves and 60.0 per cent of the time. Which of the two a claim about melodic motion means has never been stated here, and the answer matters most exactly where a tune slows down, which is at the ends of its phrases. Rhythm and metre

The note that has a length

Every melodic figure so far reads a table where each note is a pair — a pitch and a duration — and throws the second number away. A step between two minims and a step between two quavers have been one event in every histogram it has drawn. Weighting the same statistics by time moves the step distribution by up to seven points, changes forty-four of a hundred and one contour signs, and turns up an off-by-one in the one figure that did use the durations: it took the length of the note arrived at where the time between two onsets is the length of the note left.

What a contour costs to remember. A melody of n notes over 8 degrees carries 3 bits a note. Its contour carries fewer, and fewer than the number of distinct contours suggests, because the contours are not equally likely: at 6 notes there are 243 of them but the entropy is 6.59 bits, an effective alphabet of 96. Each further note adds 1.28 bits of contour against three of melody, so the shape keeps a stable 37 per cent of what is there however long the tune. Perception and the listener

The part of the tune that is kept

Contour survives transposition, retuning, a change of instrument and a doubling of every interval, and the usual explanation is that it is what a listener retains. That can be counted rather than assumed. A six-note melody over eight degrees carries eighteen bits; its contour carries 6.59 — not the 7.92 the number of distinct shapes suggests, because the shapes are wildly unequal — and the effective alphabet is ninety-six out of two hundred and forty-three. Each further note adds 1.28 bits of shape against three of melody, and at about nine notes a contour is specific enough to pick one tune out of a thousand.

Where a twelfth comes from. The range of a walk with no walls, against how many notes it runs for, at three settings of the one parameter it has. The parameter is fitted to the post-skip reversal rate and to nothing else; the range is then read off. With no central tendency at all the walk passes two octaves by 60 notes and keeps going. At the setting that reproduces 70 per cent reversal — κ = 0.78 — the range is 11.9 semitones at thirty notes and 17.9 at a hundred and twenty. It grows logarithmically, so over the whole plausible length of a tune it sits between an octave and a fifteenth, and a twelfth is the middle of that. The three tunes carried here are marked and all three fall below the curve. Form and structure

The twelfth, and where it comes from

Melodies occupy about an octave and a fifth, and an earlier essay set out to explain that by the singer's register break and found that it does not: the chest mechanism alone spans two octaves and a semitone. The answer is in a parameter the essay on leaps fitted and then put down. A walk with no walls whose central tendency reproduces the post-skip reversal rate has a range that grows logarithmically — six semitones at eight notes, twelve at thirty, eighteen at a hundred and twenty — so across every length a tune plausibly has, the span is between an octave and a fifteenth.

The series has three tops. Where the harmonic series stops, asked three ways, at four fundamentals. Consecutive partials stop being separately resolvable around partial 8 — that one depends on the fundamental, since a critical band is a fixed width in hertz and the spacing is not. They stop being a semitone apart at partial 17 at every fundamental, because the ratio (n+1)/n does not know what n is measured in. And they stop being distinguishable in pitch at all between partials 34 and 140. Every claim about how far up the series something happens is a claim about which of these three was meant. Intervals and chords

The series has three tops

How far up the harmonic series can an ear go? The question has three answers and they are an order of magnitude apart. Consecutive partials stop being separately resolvable somewhere around the eighth, and where exactly depends on the fundamental. They stop being a semitone apart at the seventeenth, at every fundamental, because the ratio does not know what it is measured in. And they stop being distinguishable in pitch at all between the thirty-fourth and the hundred and fortieth. Every claim about where the series runs out is a claim about which of the three was meant.

note length against the onsets. Every candidate metre's fit to a 16-step pattern with 4 onsets, plotted against how strongly note length is weighted. At a gain of zero the scoring is the onset-only one every earlier model used, and the winner is step 1 and step 4. At a gain of 0.05 the answer becomes step 4. 2 candidates are exactly flat — step 2 and step 3 have no onset on any strong position, so there is no credit for the cue to multiply and no weighting of it can move the line. Form and structure

What the onsets left out

Eight essays induce a metre from a list of ones and zeros, and every failure they recorded was argued about as a failure of the rules. Two of the three are not. Note length is already in that list and the scoring throws it away: put it back and the son clave's two-way tie resolves to the notated downbeat. But the groove with its beat removed cannot be repaired by any cue at any strength, and the reason is arithmetic rather than empirical — the true phase has no onset on any of its strong positions, so there is no credit for a cue to multiply and its line is exactly flat.

A model that cannot say two keys does not say it is unsure. I – IV – V – I played in C major and in a second major key at the same time, with the second key moved round the circle of fifths. For each separation: the correlation the standard key-finder gives its single best answer, and the correlation reached by the best PAIR of key profiles — a hypothesis the finder does not have. The pair recovers both keys that are sounding at every separation, 7 of 7. The single answer names neither of them at 4 of the 7, and its confidence does not fall when it is wrong: at four steps apart it reports E minor at r = 0.886, against 0.959 for the same progression in one key. Harmony and voice leading

Two keys at once

Every key figure so far assumes one key is sounding, and the standard key-finder has no value it can return that means two. Play one progression in C and the same progression a major third away at the same time, and the model does not report uncertainty: it reports E minor, at a correlation of 0.886, against 0.959 for the same progression in one key. It is as confident as it ever is, and neither key it names is being played. Give it the missing hypothesis — pairs of key profiles rather than single ones — and it recovers both keys at every separation, all seven of seven.

How many notes an octave can hold, asked twice. The share of trials on which a category is named correctly, against how many equal categories the octave is cut into, for a listener whose internal estimate carries 11 cents of noise — the logistic scale this site's identification figures already use, which is the thirty-cent transition the studies report. At 95 per cent accuracy the ceiling is 6 categories, and seven scores 94.9 per cent — on the line. The other ceiling is resolution: 151 to 356 difference limens fit in an octave depending on register, which is a factor of forty larger. Every system marked below sits between the two, and the marks separate: the number of degrees a mode uses clears the criterion, and the size of the gamut it chooses them from does not. Intervals and chords

How many boxes an octave holds

An identification model of pitch is usually handed twelve categories, and nothing ever asked how many an octave can hold. There are two answers and they are a factor of forty apart. Resolution allows between 151 and 356 — a listener can tell that many pitches apart in a direct comparison. Naming one of them without a comparison is a different faculty and it runs out at six or seven, which is where every mode in every tradition compared here sits. Turkish theory names fifty-three commas to the octave and a makam uses seven of them, and the gap between those two numbers is the whole of the argument.

What a spread costs a chord. a major triad of 3 notes lasting 600 ms each, with the onsets spread by up to 320 ms. The upper line is the share of each note's length during which every note is sounding; the lower is the chord's roughness weighted by that share, since roughness is a property of two partials sounding at the same time. At a spread of 30 ms — the asynchrony at which a mistimed partial stops belonging to its note — the chord is still 89 per cent simultaneous. It stops being simultaneous at all at 300 ms, which is where the last note arrives after the first has finished. Intervals and chords

The chord that is not played at once

Every chord until now starts its notes at the same instant, and no figure ever set the asynchrony to anything else. A spread chord is not a defective simultaneity: at forty milliseconds a triad of half-second notes is still eighty-seven per cent simultaneous, so it carries almost all its roughness, and it stops being a chord at all only when the last note arrives after the first has finished. What none of this can explain is the one thing every keyboard player knows — that a chord is rolled upward. The masking asymmetry that ought to explain it is 2.4 decibels at a close voicing, which is not enough.

How long a note has to be before its pitch is worth arguing about. The smallest audible frequency difference at 440 Hz, against how long the note lasts. The flat line is the steady-tone difference limen of 4.0 cents that every tuning argument on this site rests on. The falling line is the bound a finite duration imposes on its own frequency, 1/2T in cents, which no listener can beat. They cross at 486 milliseconds: below that the note is the limit and above it the listener is. A tenth of a second gives 19.6 cents and a quarter gives 7.9, against the commas drawn across the figure. Perception and the listener

How long a note has to be

Every difference limen quoted so far is for a tone that lasts as long as the listener needs, and no note in music does. A tone of duration T occupies a band about 1/2T wide whatever the ear does with it, so at 440 hertz the quoted five-cent limen is the right number only for notes longer than 486 milliseconds. A tenth of a second gives 19.6 cents, which does not clear the syntonic comma. Most of the tuning arguments in this collection are about a quantity that only exists in long notes, and the essays that made them said so about the listener and not about the note.

500 Hz in one ear, 504 in the other. Two tones 4 hertz apart, one to each ear. They never meet in the air, so neither eardrum sees any modulation at all and there is no acoustic beat to hear. What changes is the phase between the ears, which advances a whole cycle every 250 milliseconds — and the direction that phase implies sweeps with it, drawn here as azimuth against time. The sweep is clipped at the edges, because the implied delay leaves the range a head can produce. A head 17.5 cm across gives at most 656 microseconds, so the phase stops naming a direction above 762 Hz. Perception and the listener

The beat that is not in the air

Every sound this site synthesises reaches both ears identically, and that is the assumption none of its figures ever varied. Put 500 hertz in one ear and 504 in the other and nothing sums anywhere: each eardrum sees a steady sinusoid with no modulation on it at all. A listener still hears a four-per-second beat, which means the arithmetic is being done behind the ears rather than in the room. And it stops working above about a kilohertz — not where phase locking gives out at five, but where a head 17.5 centimetres across stops being able to name a direction, which is 762 hertz.

3 tones, one power, and the interval between them. 3 tones of fixed total power, spread symmetrically about 440 hertz, drawn against the interval between neighbours. Piled on one pitch they are one sound of that power; separated by more than a critical band — 4.5 semitones here — they are 3 sounds whose loudnesses add, and the same power reaches 2.08 times the loudness at 5 semitones. The two lines are two models of the same rule and they disagree about how abrupt the change is, not about where it goes. Perception and the listener

A chord is not as loud as its notes

The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.

How finely a fifth can be heard, against how fast it goes by. The smallest audible mistuning of a melodic fifth above 440 hertz, against how long each of its two notes lasts. The lower solid line is one note's own limen; the upper one is the interval's, larger because two independent errors add in quadrature. A tenth of a second gives 23.5 cents against the note's own 19.6, the floor for long notes is 5.4, and the syntonic comma is not cleared until each note lasts 111 milliseconds. The dotted line at 1.31 cents is the same interval heard as a simultaneity, where partials 3 and 2 coincide at 1320 hertz and one beat every 2 seconds can be counted there. Intervals and chords

An interval is two errors

Two notes in succession are two pitch estimates, and what a listener judges is their difference — so a melodic interval is heard less finely than either of the notes in it. At a semiquaver the limen is nineteen cents, which lands inside a published range long quoted from the literature and never derived. Sounded together instead of one after the other, the same interval is judged eighteen times more finely.

Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 2.8 cents at ten to one and 3.6 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it. Perception and the listener

The boundary that barely moves

Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.

A duration category has a tempo range of its own. Each simple ratio's short note is the beat divided by one more than the ratio, so at a high enough tempo it falls under the fastest interval that can be a beat at all — 100 milliseconds. Each bar here runs from the slowest tempo at which the ratio's long note still belongs to a beat to the fastest at which its short note is still a note: 1:1 ends at 300 bpm, 2:1 ends at 200 bpm, 3:1 ends at 150 bpm, 4:1 ends at 120 bpm. The line is this site's swing curve, and where it crosses a category boundary the category it is leaving has already ceased to exist. Rhythm and metre

The short note is sitting on the floor

Swing is modelled here as a short note of constant absolute length, which was measured from drummers and left as a fitted parameter. The tempo window's fast edge — the shortest interval a series of events can be a beat at — was measured from listeners tapping. Both are a hundred milliseconds, and if that is not a coincidence then the swing ratio has no free parameter in it at all: the short note is not held constant, it is resting on the floor.

How much earlier an accent is heard, by mechanism. An accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller. Perception and the listener

Playing louder is playing earlier

An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.

Crescendo, and what the impression does. A crescendo of 20 dB over 8 seconds, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.02 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up. Form and structure

Loud is relative, and it comes down slowly

The account of loudness had a model of a moment and the account of closure asked it for a model of a form. The published one exists and its content is a pair of numbers that are not the same: a listener's running impression of how loud the music is rises to meet a step in a fifth of a second and takes seven seconds to come back down. A twenty-decibel crescendo spread over eight seconds therefore buys almost no contrast at all, and the same twenty decibels taken as a step buys a factor of two.

The distribution that can be measured is not the one the player has. A player aiming at a short note of 100 milliseconds with a standard deviation of 15, against a floor at 100. The pale curve is what the player is doing and the heavy one is what can be recorded, because the 50 per cent of the parent below the floor arrives at the floor instead. The measured mean is 112.0 milliseconds rather than 100 and the measured standard deviation is 9.0 rather than 15. At 160 beats per minute that turns an intended swing ratio of 2.75 into a measured 2.35. Rhythm and metre

A quantity resting against a wall

The short note of a swung pair was found sitting exactly on the fast edge of the tempo window. If that edge is a floor rather than a fitted number, then every swing statistic computed until now was computed on a censored sample — and a censored sample has a mean that is 0.80 standard deviations too high, a spread that is 40 per cent too low, and a correction gain that can come out twice what the players actually have.

How much correlation it would take to matter. The limen of a 7-semitone interval at a note length of 0.25 seconds, against the correlation between the two notes' errors. The independent model at the left gives 9.44 cents. Halving that needs a correlation of 0.75; a fifth off it needs 0.31. The curve is √(1 − ρ) and nothing else, so the correlation required for a stated improvement is arithmetic — which turns the question from “does a key help?” into “by how much, and here is the number it must reach”. Intervals and chords

How much an anchor would have to be worth

Two pitch errors added in quadrature assume an independence nobody measured — a listener inside a key hears a note as a scale degree, and a shared reference is exactly a correlated error. Turning the dial is not evidence. What is evidence is that the dial is not free: a shared error cancels out of a difference completely, so a listener's single-note limen and their interval limen give the two components with nothing left over, and halving the interval limen needs a correlation of exactly 0.75.

Every result so far, against the listener's own noise. Three findings drawn against the one parameter all of them assume: how finely the listener resolves a pitch. At 11 cents — a trained listener, and the value every earlier essay used — the octave holds 6 nameable categories, twelve equal ones are named right 91 per cent of the time, and a 20-to-one expectation moves a boundary by 3.6 cents. At 35 cents it is 2 categories, 72 per cent, and 37 cents. The capacity falls roughly as one over sigma and the shift rises as its square, so the three curves separate rather than moving together. Scales and modes

The listener the model was never run for

Five earlier essays rest on one number — how finely a listener resolves a pitch — and every one of them used a trained listener's eleven cents. The model's dependence on it is not gentle: the capacity goes as its reciprocal and the expectation shift as its square, so an untrained listener at thirty-five cents has two nameable categories per octave rather than six, and a foreign tuning system is not mis-transcribed by them but absorbed.

Twinkle, twinkle, phrased at the level each tempo selects. The number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 5 boundaries, drawn as the flat line; the model matches it best at 160 beats per minute, where the present holds 8.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter. Form and structure

The level the tempo chooses

The boundary detector has a scale parameter and an earlier essay left it free, ending with the sentence that names this one: the scale is in notes and the psychological present is in seconds. The psychological present is two to eight seconds, a tempo converts one to the other, and the level a listener reads then stops being a parameter at all — which is a prediction with teeth, because the same tune at two tempos should be phrased differently at levels the arithmetic names in advance.

The heard moment against the pitch, on an instrument whose own attack is 8 ms. A note cannot establish an amplitude in less than 4 of its own cycles, so the attack has a floor of 4 periods — 145 milliseconds at A0 and 1.9 at C7. Below A4 the floor is longer than the instrument's own attack and the pitch decides the heard moment; above it the instrument does. The lag runs from 46.0 milliseconds at the bottom to 2.5 at the top, a spread of 44 milliseconds that no player can play their way out of. Rhythm and metre

A low note cannot start on time

Three earlier essays have held the pitch at one value. A note cannot establish an amplitude in less than a few of its own cycles, so the attack has a floor that rises as the pitch falls — 146 milliseconds at the bottom of a piano and three at the top. On an instrument whose action takes eight milliseconds everywhere, that is a forty-three millisecond spread across the keyboard from the period alone, and no player can do anything about it.

How far a wall moves the mean, in spreads. The bias a floor puts into an observed mean, in units of the parent's own spread, against how far the parent sits above the floor. With the mean exactly on the wall the truncated bias is 0.797 spreads and the censored bias is 0.399 — half of it, exactly, because half the mass sits at one point and the other half is an upper half-normal. The earlier essay used the upper figure for a process that produces the lower one, so every bias it quoted is twice what a floor on execution actually causes. Rhythm and metre

The shape a wall leaves behind

A floor on execution biases every statistic computed on swing timing, and the bias was priced with Pearson's truncated-normal formulas. Those are the formulas for a sample with everything below the wall thrown away. A player who cannot execute a short gap does not throw the attempt away — it comes out at the floor. That is a censored sample, its bias is exactly half, and its skew is two thirds larger.

What the notes in between do to the anchor the interval is measured against. How finely a 7-semitone interval can be judged when its two notes are separated by other notes rather than by silence, under the two published accounts. Confirming material restates the key and refreshes the shared reference, so the correlation climbs from 0.5 toward a ceiling and the limen falls to 4.81 cents. Overwriting material competes for the same memory, so the correlation decays to 0.04 and the limen rises to 9.28. By 8 notes the two accounts differ by 4.5 cents, which is 47 per cent of the limen with no anchor at all — and no experiment here distinguishes them. Intervals and chords

The notes in between

Every figure until now is about two notes with nothing between them, and a melody is notes with other notes between them. Two published accounts of what the intervening material does predict opposite signs — one says the key is restated and the shared reference is refreshed, the other says each note competes for the same memory and it decays. By eight notes they differ by four and a half cents, which is nearly half the limen the interval would have with no anchor at all.

Every resolution claim here, against the number it rests on. How many equal steps of the octave can be named at 95 per cent accuracy, against the internal noise the model gives a listener. The laboratory value this collection quotes everywhere is 11 cents, which gives 6 nameable categories — the "about seven per octave" every claim here has been repeating. The laboratory measures it on isolated intervals and music never presents one, so the effective value inside a piece is smaller by an amount nobody has measured: at 4 cents it is 18, which is the chromatic scale, and the conclusion changes from "the ear has fewer boxes than the notation" to "it has exactly as many". Nothing here measures it. This is what it is worth if it moves. Scales and modes

The number every claim here has been quoting

Sigma is the internal noise a listener's pitch judgements carry, it is measured in a laboratory on isolated intervals, and music never presents an isolated interval. Every resolution claim here rests on the laboratory value. Sweep it and the headline finding moves: at eleven cents the ear has about six nameable categories per octave, and at six it has twelve — which is the chromatic scale, and turns 'the ear has fewer boxes than the notation' into 'it has exactly as many'.

Roughness and loudness do not rise together. One four-note chord, played at levels from 35 to 95 decibels, with both quantities drawn as multiples of what they are at the quietest. Roughness is quadratic in pressure, so 60 decibels multiply it by 1.0e+6. Loudness is compressive — about ten phons to a doubling of sones — so the same range multiplies it by 96. The gap between the two lines is the quantity: roughness per sone rises by a factor of 1.0e+4 between a pianissimo and a fortissimo of the same chord. Harmony and voice leading

The ranking survives the dynamic and the chord does not

Roughness is quadratic in pressure and loudness is compressive, so sixty decibels multiply a chord's roughness by a million and its loudness by ninety-six. Roughness per sone therefore rises ten thousandfold between a pianissimo and a fortissimo of the same four notes — and yet the ranking of which doubling is smoothest, over four hundred and eighty voicings, does not move by a single place.

Every arrival at one seat, against the delay at which it would be an echo. The echogram: each reflection at its delay after the direct sound and its level relative to it, in a 22 by 45 by 15 metre hall with 18 per cent absorption. The line is the published echo threshold for speech — 40 milliseconds at equal level and about 3.0 more for each decibel of attenuation — so anything to the RIGHT of it is late enough and loud enough to be heard separately. The once-reflected rear wall arrives at 210 milliseconds, 25 decibels down, against a threshold of 114 — well past it. Nothing here stands clear enough of its neighbours to be heard as an echo, and 181 arrivals are fused with the direct sound instead. Perception and the listener

An echo is prevented by the crowd around it

The echogram says when every reflection arrives and how loud it is; the published echo threshold says when a reflection that late and that quiet is heard separately. Put one against the other and the rear wall of every hall anybody builds is past the threshold — a 45-metre hall puts it 210 milliseconds late and 25 decibels down against a threshold of 114. It is not heard as an echo, and what saves it is not the geometry. It is everything else arriving at the same time.

Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together. Perception and the listener

The partials that do not die together

The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.

A vibrato flattens the dissonance curve. Each interval twice: hollow is its roughness computed at the two notes' nominal frequencies, filled is the average of its roughness over a vibrato cycle of 50 cents at 6 hertz. They are not the same number, because roughness is a curved function of the frequency difference and the average of a curve is not the curve of the average. The largest effect is at octave, where the moving average is 19.3 times the still value — an interval sitting in a deep narrow minimum is smeared out of it. The smallest is at major seventh, where it is 0.95: an interval near a maximum is smeared out of that too. Vibrato pushes every interval toward the middle, and what it takes away from the consonances is much more than what it takes away from the dissonances. Intervals and chords

A roughness with a rate of its own

Every roughness figure so far computes one number for a steady spectrum. Evaluate the same sum at every instant of a vibrato and there are three numbers instead — a mean, a depth and a rate — and the mean is not the roughness of the mean frequency. On an octave it is nineteen times it, because an octave sits in a deep narrow minimum and a vibrato smears it out of one.

The same interval, started on each of the twelve. An interval of 7 semitones started on each pitch class of a major key, against how strongly the key specifies its two notes — the mean of the probe-tone profile at each. The interval account says the listener encodes a distance, so the key cannot enter and the prediction is a horizontal line at 5.4 cents. The degree account says the listener refers each note to the key, so its precision on a note falls as the key's specification of that note weakens; scaled to agree at the most stable start, it rises from 5.4 cents on C to 8.5 on E♭. Every earlier figure measures a quantity the second account says is not being formed at all. Intervals and chords

The quantity a rival account says is not there

Two earlier essays measure how much two notes' errors are correlated through a shared anchor, and price what that correlation would be worth. There is a rival account in which a listener refers each note to a key and never forms the distance at all — under which the correlation is not small, it is a description of something that is not happening. The two accounts agree on almost everything and disagree on one manipulation, and the manipulation costs an afternoon.

The tempo turns, and almost nothing moves. The earlier arrival reading — what is sounding at the final chord over what the listener has been hearing — swept over bar lengths from 0.5 to 5 seconds, which is 480 down to 48 beats a minute, at 3 closing lengths. Every curve is nearly flat. Across a tenfold change of tempo one gesture's reading moves by a factor of 1.201 and the other's by 1.098, while the gap between the two gestures — which is what that essay was measuring — is 1.228. The expectation was that the tempo would decide the answer, on the grounds that a two-second bar against a two-second release is a comparable pair. The premise is wrong in a way the sweep makes obvious: the thing being compared with the release is not a bar, it is the WHOLE ENDING, which is 2 to 8 bars long and is therefore far longer than the release at every tempo anybody plays. The running impression has caught up with the closing texture before the final chord arrives, at 0.5 seconds a bar and at 5, and what is left is the last bar's own jump. Form and structure

The parameter that did not decide the answer

An earlier essay on closure ended by naming the tempo as the thing every number in it was resting on, and said it was the kind of parameter that had caused trouble before by turning out to decide the answer. Turned across a tenfold range at a closing gesture of fixed length it moves the reading by four per cent, against a twenty-three per cent gap between the gestures it is distinguishing. The parameter beside it in the same figure — how many bars the gesture occupies — moves it by twenty, and nobody had named that one at all.

Which notes of a scored chord have to be played early. Four parts of one chord, each with its own instrument, its own pitch and its own dynamic, and the perceptual centre that comes out of all three. piano, sforzando on E1: an attack family of 8 milliseconds against a pitch floor of 97, so the pitch is what limits it, shortened by the dynamic to 65, heard 20.5 after it starts and needing to be played 12.0 early; flute, quiet on A5: an attack family of 60 milliseconds against a pitch floor of 5, so the instrument is, shortened by the dynamic to 69, heard 21.8 after it starts and needing to be played 13.2 early; violin, mezzo forte on E4: an attack family of 90 milliseconds against a pitch floor of 12, so the instrument is, shortened by the dynamic to 90, heard 28.5 after it starts and needing to be played 19.9 early; trumpet, forte on A3: an attack family of 30 milliseconds against a pitch floor of 18, so the instrument is, shortened by the dynamic to 27, heard 8.6 after it starts and needing to be played 0.0 early. The spread is 19.9 milliseconds, which is well above the two or three a listener resolves, so a conductor asking for these four to sound together is asking for four different physical onsets. Rhythm and metre

Which notes have to be played early

There are three separate contributions to one quantity — the instrument's attack family, the dynamic it is played at, and the note's own period — and every figure so far varies one and holds the others. Added together for a real scoring they do not add: a sforzando low piano note is pitch-limited to a hundred-millisecond attack and the sforzando shortens it back to sixty-five, so flattening the dynamics makes the ensemble's spread larger rather than smaller.

The violin's attack times, string by string. Every playable cell of the earlier map, with its bow-force window turned into a time. A player aiming at the geometric centre of the window has to wait until the accelerating bow's maximum force rises to that value, which is a fraction one over the square root of the window width of the way through the bow's ramp — so a narrow window is a late note and the exponent is a half. The latest cell is C♯4 on the G3 string at 21.9 milliseconds, against 4.1 at C6 on the A4: a factor of 5.3 in time out of a factor of 29 in window width. The darkest cell is the same hardest place, unchanged — the map is the same map under a monotone change of units, and what is new is that the units are milliseconds, which a player and a listener both have access to. Rhythm and metre

The hardest place is also the latest

A bow-force window is a ratio of forces, which nobody can hear. An attack time is milliseconds, which a player and a listener both have. The map of the violin's windows becomes a map of its attack times under a change of units, and the exponent turns out to be a half — so a window twenty-nine times narrower is only five times later.

How far out of tune an interval has to be before its beat is usable. An earlier essay produced a modulation index for every interval on a stated timbre; whether a fluctuation of that index at that rate can be detected is a published function of both. Running one against the other turns a table of decibels into a window in cents. The bar is the mistuning over which the beat is both deep enough to notice and at a rate a tuner can use — not so slow that a beat takes half a minute to complete, not so fast that it has stopped being a beat. the major third gives the widest window, 1.3 to 60.0 cents, and the minor sixth the narrowest, 0.8 to 38.9. The ordering is the opposite of the dip's: the octave has the shallowest dip on a string spectrum and the widest usable window, because its coincidence sits at a low partial and a given mistuning therefore produces a slower beat. Depth and rate pull opposite ways, and it is the rate that decides. Intervals and chords

The beat a tuner can actually use

There is a modulation index for every interval on every timbre, and the published threshold for detecting a fluctuation is a function of exactly that and its rate. Running one against the other turns a table of decibels into a window in cents — and reverses the ordering, because the interval with the shallowest dip has the widest window.

The same roughness, before and after the window it has to be heard through. The instantaneous roughness of an interval under a vibrato, and the same quantity after a running average of 59 milliseconds — the time a dissonance has to last to be heard as one, which is 4 cycles of this interval's own 68-hertz fluctuation rather than a number chosen for the figure. The mean is identical to every digit, 0.1487 against 0.1487, because a running average cannot change an average — so the earlier Jensen factor of 1.0 survives the window untouched and its prediction that the window would shrink it is wrong. What the window destroys is the depth: 0.30 of the mean becomes 0.23, which is 77 per cent. The roughness a vibrato adds is heard; the fact that it is moving is mostly not. Intervals and chords

The mean survives the window

A roughness that moves has a mean, a depth and a rate — all three of which a listener could only have through a temporal window. Applying the window already to hand settles which of the three survives, and the answer refutes the guess: a running average cannot change an average, so the octave's factor of nineteen stands and the movement is what goes.

Where the sound is, and how wide. An earlier essay sorted a hall's arrivals into echoes and everything else. Everything else is not nothing: a reflection too early to be heard as a separate event still moves the apparent source, widens it and colours it, and all three come out of the same list of times, levels and angles. At this seat the direct sound arrives from 26.6 degrees off the front and the image is heard 8.5 degrees left of it, pulled by the near side wall. The apparent width is 33 degrees, from a lateral energy fraction of 0.23. And the strongest early reflection arrives 0.6 milliseconds behind off 1× the floor, which is a comb filter with notches every 1608 hertz and 25 decibels deep. The trading ratio and the discount on late arrivals are stipulated rather than measured, so the degrees are ordinal: what the figure claims is the direction and the shape, not the number. Perception and the listener

A position and a width

Sorting a hall's arrivals into echoes and everything else settled that fusion is not a yes or a no. Everything else is not nothing: a reflection too early to be heard separately still moves the apparent source, widens it and colours it. All three come out of the same list of times, levels and angles, and none of them needed a new input.

What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged. Perception and the listener

The exchange rate nobody has

There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.

A loud chord is a smaller chord. The share of a voicing's partials that stand above what the rest of it masks, and the share of its computed roughness that is between partials a listener actually has, from 30 decibels to 100. Both fall: 83 per cent of the partials survive at 30 decibels and 38 at 100, and the roughness share goes from 88 per cent to 67. The direction is the upward spread of masking, which grows faster than linearly with level: a loud partial masks a band above itself much wider than a quiet one does, so the chord's own top disappears into its own bottom. Two earlier essays are drawn at one level, and this is what they were holding. Harmony and voice leading

A loud chord is a smaller chord

Two earlier essays hold the level fixed, and the level decides how much of a chord a listener is given. At thirty decibels twenty of a triad's twenty-four partials stand above what the rest of it masks; at a hundred, nine do. Every roughness figure until now counts partials that are in the score, and a partial the chord masks is not a partial the listener has.

A ritardando does not spend the diminuendo. The arrival reading under a deceleration into the ending, from no ritardando at all to a final tempo 30 per cent of the starting one — which stretches the closing bars from 12.0 seconds to 20.2. The expectation was that it would matter: a ritardando lengthens exactly the bars the gesture is happening in, so a diminuendo that would have been absorbed at a steady tempo gets more of the smoother's own time to be absorbed in. It moves the reading by 0.00 per cent. Every line here is flat to within the thickness of the line, which is the second time a tempo parameter has been swept here and found to do nothing. Form and structure

The reading was a step response

Sweeping the tempo found it did not decide the answer. This one sweeps the deceleration across a factor of three and finds a null to five figures, and then sweeps the length of the closing gesture across a factor of forty-eight and finds it moves the reading by eight per cent — but not as a function of seconds. Sorted by seconds the twelve runs scatter; sorted by how many bars the instruction covers they fall into three tight groups. One sentence explains the null and the not-null together.

A written dynamic is an instruction to the listener's impression. Every earlier scoring holds one chord still. A passage is a succession, and the running impression of loudness carries a chord into the one after it, so what a marking asks for and what playing the marking produces are different things. Here is a five-chord passage with a written shape. Playing each chord at its own written loudness gives the running impression 2.4, 3.0, 4.2, 5.6, 4.0 sones against the 2.4, 3.0, 4.2, 5.6, 2.0 that were asked for — right until the last chord, where it misses by 2.0. Solving for levels that make the impression arrive at the marking does not fix it: the last chord's target is I, two parts, and it is unreachable — the correction runs to silence and the impression still sits 1.1 sones above. A subito piano after a full chord is not a level a player can produce. It is a rate of change, and the smoother's two-second release is what refuses it. Form and structure

A subito piano is a rate, not a level

All three earlier essays score one chord held still. An orchestration is a succession, and the running impression carries a chord into the one after it — so a written dynamic is an instruction to the listener's impression rather than to the instantaneous sound, and there are markings that cannot be produced at all. The correction runs to silence and the impression still sits above the target.

The same tune read at six widths of the psychological present. A later essay made the detector's smoothing width a function of position, which removed its last free parameter but one — and the one it cannot remove is the width of the psychological present, because that is a fact about listeners rather than a choice. So the honest object is not a reading but a family of them, one per width. A short present finds 6 boundaries and a long one finds 2, and the family agrees on 0 of them. The fixed-width control, at its own best width, scores 0.67 against the adaptive readings' 0.67, 0.75, 0.33, 0.33, 0.40, 0.40 — so the adaptation does not win, which is what that essay reported too. What the family adds is the ordering: a boundary in every row is a different claim from one in a single row, and a single reading has no way to say so. Rhythm and metre

A family of readings

Removing the detector's free parameter, and then its constant tempo, cost persistence both times — the property that made its boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than one, indexed by the width of the psychological present, which is the one parameter that cannot be removed, because it is a fact about listeners.

An ensemble finding an asynchrony nobody told it about. An earlier essay produced a map of required leads — which notes of a scoring have to be played early, and by how much — and nothing tells the players those numbers, because they are a property of the instruments' attacks rather than of the music. So an ensemble has to find them, and the mechanism is already here: each player hears sounds rather than onsets and moves their next onset toward the mean of the others'. The spread of arrival times starts at 20 milliseconds and settles at 4, crossing 5 milliseconds after 5 beats — about 1.3 bars of four. The leads it converges on match that map to within 0.2 milliseconds, which is what makes this a convergence rather than a coincidence: the fixed point of players listening to each other is every player leading by their own attack. Rhythm and metre

How many bars an ensemble needs

The map of required leads is something nobody tells the players, because the leads are a property of the instruments' attacks. So an ensemble has to find them, and the mechanism is the one the microtiming essays describe: each player hears sounds rather than onsets and moves toward the others. It converges on the map to within a fifth of a millisecond, in five beats, and there is a best correction gain.

How much of the reading comes from what has not happened yet. Every margin reported earlier is two-sided: the best path through a key at a bar is the best score into it plus the best score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line, because the dynamic program already had both halves separately. The mean margin falls from 3.90 bits with hindsight to 1.79 without it, so 54 per cent of this passage's certainty is retrospective. The two passes never disagree about which key is best here, so the hindsight buys confidence rather than a different answer. This is the quantity every earlier essay has assumed and none has measured. Scales and modes

How much of the reading arrives late

Every margin reported earlier is two-sided: the best path through a key at a bar is the score into it plus the score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line. On a thirty-two-bar song it removes more than half the certainty, and on a passage built to be ambiguous it changes the key named at nine bars out of eleven.

Every standard rastral size against the two bounds a reader imposes. Print the notes larger and the eye-hand span stops fitting inside one fixation, so the reader has to saccade ahead faster than the eye can move. Print them smaller and a notehead stops subtending enough angle to be identified. Both bounds come from the reader and neither from the music. At 100 beats a minute with 2 notes to the beat, the acuity bound sits at 1.63 millimetres and the saccade bound at 7.5 — so the saccade rate is nowhere near binding and acuity is doing all the work, which is the opposite of what the eye-hand span suggests. 5 of the 9 standard rastrals clear the acuity bound: rastral 4 and larger. Those are exactly the sizes used for parts, and the ones below are used for study scores — which are read at a desk rather than played from at a stand, and a shorter viewing distance moves the bound with them. Scales and modes

The page is read by an eye

A sight-reader's eye sits a fixed number of notes ahead of the sounding one and a fixation takes in a fixed number of millimetres, and the spacing rule converts between them. Two bounds follow, from the reader rather than from the music — and the one everybody would expect to bind does not. The saccade rate has enormous headroom at any playable tempo, and what decides is acuity.

The violin is the one where the bow's width catches up. Three separate limits on how near the bridge a bow can go, for the four bowed instruments. The ribbon's near edge reaches the bridge at 5, 6, 6, 8 millimetres; the ribbon covers half the corner's bridge-side excursion at 10, 11, 12, 15; and Schelleng's force window narrows to a factor of 3 at 10, 12, 21, 32. On the viola, cello and double bass the force window binds first, by a margin that widens to a factor of two on the bass. On the violin the ribbon binds first, at 10 millimetres against 9.8. A bow's hair is as wide as a hand can control and a string is as long as its pitch requires, so the ratio between them is a fact about the violin rather than about bowing. Instruments and their design

The bow is not a point either

Seven earlier essays set the bowing point to a number between 0.02 and 0.3 and drew it as a point. A violin bow's hair is ten millimetres wide on a string of three hundred and twenty-five, so near the bridge the ribbon is wider than its own distance from it. In the spectrum that turns out to be worth nothing. In the geometry it is the reason sul ponticello has a floor, and the violin is the one instrument of the four where it arrives before the force does.

The period is still there, and it is wider. The autocorrelation of a 12-partial complex on 220 hertz, drawn twice: steady, and averaged over one cycle of a 71-cent vibrato. A vibrato moves every partial by the same number of cents, so the complex is exactly harmonic at every instant and nothing is mistuned — what moves is the period the extractor is looking for. The peak survives. It loses 6 per cent of its height above the surrounding lags and gains 11 per cent in width, because the vibrato swings the period by 0.37 milliseconds against a peak 0.90 wide. Its maximum also moves, to 2.4 cents sharp of the still tone's, which is a prediction with a sign in it. Instruments and their design

The pitch that does not wobble

Three earlier essays have treated a vibrato as a modulation of roughness. The reason singers use one is what it does to the note, and there is an extractor here that turns a set of partials into a pitch and has never been asked what it does with partials that will not hold still. The period survives, at a cost that rises with the extent — and the practice stops within a hair of where the cost becomes total.

Settling and being heard are not the same quantity. Across, how long the instrument takes to reach its steady amplitude, computed from its own physics — a resonance's Q, a bow's capture, an exciter's contact. Up, how long after its physical onset a listener places the note, computed from the measured shape of its envelope. Five instruments both accounts hold. The diagonal is where they would agree and nothing is on it. The ratio between them runs from 0.10 to 6.6, a factor of 69, and it sorts perfectly by mechanism: about 6.6 for a struck or plucked string, 2.2 for a bowed one, and about 0.15 for a wind. Ordering the five by each measure changes the place of 3 of them, and the one that moves furthest is the violin — third slowest to settle and the last to be heard. Rhythm and metre

A note starts twice

One account computes how long an instrument takes to settle, from its own physics. Another computes how long after its onset a listener places a note, from the shape of its envelope. Both come out in milliseconds and neither has ever been shown the other. Paired on the five instruments they share, the ratio between them spans a factor of sixty-nine and sorts perfectly by mechanism — and the violin is third slowest to settle and the last to be heard.

One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics. Form and structure

The dissonance arrives and the dynamic does not

A scoring decides two things at once and both of them have to be integrated by a listener before they exist. The loudness smoother's release is two seconds and the roughness window is thirty-seven milliseconds, and that ratio of fifty decides which of the two survives at the pace music is actually played. Nothing anybody performs is fast enough to blur a dissonance, and a great deal of it is fast enough to average a dynamic.

The hall, as two numbers a listener has. Every reflection at this seat, placed by the interaural delay it produces rather than by the direction it came from. Time runs down; the dot's size is its energy. The direct sound is at 0 microseconds and the reverberation spreads over the whole available range, with a root-mean-square width of 322 against a geometric maximum of 656. That is the position and width computed earlier, in the units a listener has instead of the vectors used until now. 10 of the 56 reflections arrive from behind and carry 9 per cent of the energy — and they are drawn where they are because the interaural delay of a reflection from 120 degrees is identical to one from 60. Perception and the listener

The hall through a head

Every direction computed so far is a vector from a seat to an image source, and a listener has no vectors. Run the echogram through the head computed six essays ago and two things happen: the position and width become microseconds, and a third of the room disappears — because both cues fold at ninety degrees and a reflection from behind is identical to one in front.

Nothing at all until fifteen decibels, and then it depends on the tempo. The fraction of a line's partials that stay above threshold, over how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement — forward masking leaves a threshold at most ten decibels below the masker, and a note's own partials mask each other from the same components at full level. The boundary is between twelve and eighteen decibels, and beyond it the loss grows with the tempo: at 280 to the crotchet and 36 decibels of contrast, 26 per cent of the line's partials are gone. Fifteen decibels is about the gap between a forte and a piano. Perception and the listener

An equal note cannot be masked

Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.

A listener has a quarter of an analyst's confidence and the same answer. The mean margin between a passage's best two key readings, against how many bars a listener's memory of the evidence takes to halve. The two-sided reading — the one that uses bars that have not happened yet — sits at 3.90; the forward pass with perfect recall at 1.79; a forward pass whose evidence halves every 6.6 bars at 1.00. The dots' size is how often that reading names the same key as the two-sided one: 100 per cent at perfect recall and 81 at a one-bar half-life. So forgetting costs a great deal of confidence and very little accuracy — the key is robust and the certainty is not. Harmony and voice leading

The listener who forgets

Setting an analyst's reading of a key against a listener's measures what arrives late. Both passes assume perfect recall of their own half — which is as wrong going forward as knowing the future is going back. Put a decay on the forward pass and a listener with a memory of a few bars keeps a quarter of the confidence and nine tenths of the answers.

The number nobody has moves the size and not the order. The mean total surprise per chord change, against how much the two surprises share. At zero they are independent and the total is their sum; at one they are the same event and the total is the larger of the two. The mean falls by a factor of 1.53 across that whole range, which is the size of the thing a corpus would settle. The ordering of the events by total surprise does not move at all until the very end: 5 of the 6 correlations swept give exactly the ordering independence gives, and only perfect dependence changes it, by 3 places out of 7. The most surprising event in the passage is the same one at every correlation. So the corpus three separate accounts have recorded wanting would change what this figure reports and not what it concludes. Harmony and voice leading

Two surprises and one event

A chord change is surprising twice over — in which chord it is, and in when it comes — and a listener meets one event. Adding two surprises needs to know how much they share, which is a fact about a repertoire nobody has. Sweeping it instead: the total moves by half, the ordering does not move at all, and the most surprising moment in a passage is the same one whatever the answer turns out to be.

The same eight notes are four times as much to read. How many bits each note of a line carries, taken as minus the log of the probability of the interval that reached it, under the distribution of melodic steps measured over the tunes used throughout. A scale costs 1.76 bits a note and a wide leaps costs 7.02 — a factor of 4.0 at the same number of notes on the page. Every quantity computed until now counts notes, and the page cannot tell these apart: eight quavers are eight quavers of horizontal space whichever line they spell. Scales and modes

A reader does not read notes

Eleven earlier essays count notes, and the page cannot tell one line of eight quavers from another. A reader can: a scale of eight is one object where eight leaps are eight. Measured against the melodic interval distribution, the same eight notes are four times as much to read — and the eye–hand span, the best-measured quantity in the reading literature, is four notes of a tune and one of a leaping line.

A rest is a diminuendo, and a long one. How far a listener's running impression of loudness falls during a silence, converted into the diminuendo that would have taken it the same distance. Half a second of nothing is worth 3.2 decibels, a second and a bit is worth 8.1, and two and a half seconds is worth 17. The marked line is three and a half seconds, which is where a gap starts to be heard as an ending rather than as a pause: at that length the reference has fallen by 24 decibels, which is more than a fortissimo to a pianissimo. A tempo swept over a factor of ten, a deceleration over a factor of three and a gesture length over a factor of forty-eight all returned the same reading to four significant figures. This one moves it by twenty-four decibels. Form and structure

A rest is a diminuendo

Three parameters swept over factors of ten, three and forty-eight returned the same reading to four significant figures. The one manipulation left unswept moves it by twenty-four decibels: a silence. A listener's running impression decays at the loudness smoother's two-second release, so a general pause is a diminuendo nobody wrote, and at the length that makes a gap an ending it is worth more than any marking a composer has.

The passage that separates them, and a listener cannot hear it. A scoring changes at the halfway bar, and the two maps of required leads differ by 18.1 milliseconds at their widest. An ensemble that has internalised the map applies the new one on the first note of it and its spread never leaves zero. An ensemble that is listening to each other has to re-converge: its spread jumps to 11.8 milliseconds and takes 3 beats to get back under 5. The dashed line is twenty milliseconds, which is what a listener notices — and the disagreement never reaches it. So the two accounts are separable on a recording and very nearly not separable by ear, which is why nobody has noticed the distinction and why the measurement is worth making. Rhythm and metre

The passage that separates two players

An ensemble that has learnt where the asynchronies are applies them; one that is listening discovers them. In steady state the two are identical, which is why nobody has separated them. Change the scoring mid-phrase and they are not: one ensemble is wrong by twelve milliseconds for three beats and the other is not wrong at all — and twelve milliseconds is under what a listener notices and far above what a microphone resolves.

A turn of 0.99 degrees tells front from back. A source 45 degrees off centre and its mirror image 135 degrees off, which produce the same interaural delay and are therefore the same signal to a listener who does not move. As the head turns the two predictions separate: the front source's delay falls and the rear source's rises, because the fold at ninety degrees puts them on opposite branches of the same curve. They differ by the 15-microsecond threshold after 0.99 degrees of turn — which is exactly half the 1.97 degrees a source would have to move for the same listener to notice it moving, and it is half for a reason: a turn displaces the two hypotheses from each other by twice what it displaces either of them from where it started. Perception and the listener

The turn is half the angle

A stationary head cannot tell a sound in front from the same sound behind, and an earlier essay said so at length. The turn that breaks the confusion is 0.84 degrees — exactly half the angle a source would have to move for the same listener to notice it moving, and half for a reason. In a hall the same turn does something else: the source swings at 8.9 microseconds a degree and the room swings at 2.7, so a listener who moves is separating the soloist from the reverberation as well as the front from the back.

The same geometry at four sizes of head. Woodworth's interaural delay against direction, for 4 head radii from 5.8 to 9.8 centimetres. The whole range runs from 431 microseconds for a newborn to 735 for a large adult, and it scales exactly with the radius because the delay is (r/c)(θ + sin θ) and r is a multiplier. The detection threshold does not scale with the listener, so the number of distinguishable delays across the whole range falls from 98 to 57: a smaller head has the same directions in front of it and a shorter ruler to measure them with. Perception and the listener

A smaller head in the same hall

Ten earlier essays draw one head. Every parameter belonging to the room has been varied by some figure and the one belonging to the listener never has, and it is the only one whose change the detection threshold does not follow: a six-year-old in the same seat receives the same fifty-four reflections at the same instants and reads them onto an axis with seventy distinguishable positions instead of eighty-seven. The speed of sound, swept over every temperature a hall is ever at, changes nothing at all — and the reason it cannot is the reason head size can.

Eleven partials is one partial too many. What fraction of a spectrum the harmonicity census finds fused, against how many partials it is asked to census. At ten a perfect harmonic series fuses 10 of 10 and the fundamental it finds is the right one. At eleven it fuses 5 of 11 and the fundamental jumps to exactly 2.00 — the octave above. The cause is the cap the census carries for a reason established earlier: without it a bell fuses perfectly at a fundamental nobody could hear, so the search refuses any fundamental more than about ten harmonics below the top partial. At eleven partials the first thing that cap excludes is the series' own fundamental, and the census then takes the octave and calls every odd partial inharmonic. So the number of partials and the cap are the same number, and nothing had ever said so, because every earlier figure censuses ten. Perception and the listener

Eleven partials is one too many

Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.

How long a bar of unequal beats can be, at a present of 3.5 seconds. Two quantities against the number of subdivisions in the bar. The bars are how many genuinely distinct unequal metres that length admits — groupings of twos and threes, up to rotation, discarding any that repeats a shorter grouping — and they run from one at 5 to 28 at 23. The line is how good a beat the best of those metres can manage once the whole bar is required to fit inside a psychological present of 3.5 seconds. It is flat at 0.865 up to a bar of 16 units, which is where the bar at the best subdivision first overruns the present, and falls after it: 17 at 0.830, 18 at 0.797, 19 at 0.765, 20 at 0.735. The supply of metres is still growing where the quality has begun to fall, so the lengths a tradition can use are a bounded prefix of an unbounded list. Rhythm and metre

How long a limping bar can be

The bound on an unequal beat turned out to be arithmetic, and the tempo window was left with only the tempo to decide. It decides nothing: every metre built from twos and threes gets the same answer, because the window is asked a yes-or-no question. Graded instead, an unequal beat costs 0.135 of the window's own preference at every bar length — and the constraint that does depend on length is the one nobody applied, that the whole bar has to fit inside the psychological present. At the subdivision that suits both beats best, a bar of sixteen units just fits and a bar of seventeen does not, which is where the supply of distinct metres has only started to grow.

Where a cycle of 16 at 16, 8, 4 outruns the listener's memory. The residual uncertainty a listener is left with once the evidence has stopped accumulating, against how long one turn of a 16-step cycle takes. The listener's memory of a step halves after 3.5 seconds throughout; what changes is how many steps that is. At a cycle of 1.6 seconds it is 35 steps and every design reaches certainty, which is the regime a clave is played in. At 60 seconds it is 0.93 steps and none of them does: layers at 16, 8, 4 settles at 1.89 bits, son clave settles at 2.27 bits, the bossa-nova pattern settles at 2.38 bits, the best single line of 7 settles at 1.85 bits. That is the range a gong cycle occupies, and it is the design that wins there. Rhythm and metre

The cycle that outruns the memory

A timeline and a colotomy were compared at equal strokes and the comparison had no clock in it. A memory span is a number of seconds and a cycle is a number of steps, so the two only meet through a tempo — and at a clave's two seconds a listener's memory covers twenty-eight steps and forgets nothing, while at a gong cycle's forty it covers 1.4 and forgets almost everything. The single line is the better locator up to twenty-three seconds a cycle and the layered code is better after it, which is very close to where each is actually used.

The bias a wall puts in a mean, against the parent it came from. The bias a wall puts into an observed mean, in spreads, for each of eight standardised parents, with the wall exactly on the parent's mean. The censored values run from 0.354 to 0.433, a range of 0.079; the truncated from 0.582 to 1.000, a range of 0.418. The formula now used is the one that hardly depends on a distribution nobody has measured, and the formula it replaced is the one that depends on it a great deal. Rhythm and metre

The parent nobody measured

Every figure drawn behind a wall so far assumed a normal parent, and the assumption turns out to matter in exactly the wrong place. The censored bias is half the parent's mean absolute deviation — a theorem, not a coincidence — so it lands between 0.35 and 0.43 spreads for every distribution tried, and the truncated one runs from 0.58 to 1.00. But the third moment proposed as the test moves 2.9 across parents against 0.65 between the two rules, and a censored sample from a slightly left-skewed parent has a skew of 1.007 where a truncated normal has 0.995. The statistic that does work is a count of ties.

What a slow start costs a trumpet, in its own settling time. The amplitude of a trumpet's A4 resonance — the 4th impedance peak, of Q 37, time constant 27.1 milliseconds — driven from rest by a pressure that rises over 10, 40, 80, 140 milliseconds, against the step every earlier figure has assumed. The step reaches 90 per cent of its final amplitude in 62.5 milliseconds. A ramp of 140 takes 157.8, which is 95.3 more — and that excess is 68 per cent of the ramp's own length. Across the whole range a player works in the excess is a little over half the ramp: 0.52, 0.56, 0.61, 0.68 at 10, 40, 80, 140 milliseconds. So the tongued attack is not something added to the note. It is the step, which is what has been computed all along, and what has a price is its absence. Rhythm and metre

What the tongue actually removes

The question this essay was written against asked for an impulse: a tongued attack, a martelé stroke and a struck key all deliver one before the steady drive begins. The arithmetic refuses the framing. A tongue release does carry energy at the note's own frequency, and it is worth 0.17 milliseconds on a trumpet against a settling time of 62 — capped at about 1.13 over the resonance's Q. What articulation is worth is the ramp it removes, and that is about half the ramp's own length: 44 milliseconds, and very nearly the same 44 on every wind instrument in the collection.

The census with the criterion moved under it. Every instrument's speaking time in milliseconds, against the fraction of the steady amplitude counted as speaking. The criterion is in the wind instruments alone: a resonance takes −ln(1−p)·Q/(πf) to reach a fraction p, so those lines rise across the whole picture, while a bow's capture and an exciter's contact contain no criterion at all and are flat. Every earlier figure sits at 0.9, where the wind instruments are the slowest things in the collection by a factor of 20.5. At 0.05 they are the fastest: a violin's G3 string is the slowest at 15.5 milliseconds and a trumpet takes 1.4. The three clusters cross at a criterion between 0.18 and 0.39, which is inside the range the perceptual measurements work in — their three named criteria are 15 decibels below peak, 6 decibels below peak, and ninety per cent — and the settling figures have only ever used the third of them. Perception and the listener

Read at two different heights

Ten placements of these figures, one value: the settling criterion is nine tenths in every one of them, and nothing is measured behind it. It is a multiplicative constant only inside the mechanism that has it — a bow's capture and an exciter's contact contain no criterion at all — so moving it rescales one of three clusters against two that stand still. The most-quoted number here, a factor of sixty-nine between the instrument's account and the listener's, is 5.9 at the criterion the listener's own measurements use, and the ordering an earlier essay was written about does not exist below a fifth.

An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none. Perception and the listener

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

Articulation is worth 2.6 phons, and nobody counts it. A passage of notes at 80 decibels, 2 to the beat, at 7 tempi and 3 articulations, scored against the same level held continuously. The variable is the fraction of each inter-onset interval that is sounding — 0.95 is a legato, 0.4 a staccato — and the vertical axis is what that costs the passage's running loudness in phons. Nothing here is anybody playing harder or softer. At 40 to the beat the span from legato to staccato is 1.47 phons; at 200 it is 2.61, because a staccato note there lasts 60 milliseconds and no longer reaches its own loudness either. Perception and the listener

A staccato is a dynamic mark

Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.

The top voice arrives whole and the bottom one arrives as a sine. 4 parts sounding together, each of 8 partials, with every partial tested against the summed masked threshold of every component in the texture. A filled mark is a partial the listener receives and an open one is a partial the part would have had alone and does not have here. The bass at C3 keeps 1 of 8, the tenor at C4 keeps 2 of 8, the alto at E4 keeps 5 of 8, the soprano at G5 keeps 8 of 8, every one of them at 70 decibels. Every part is at the same level and the difference is entirely where each one sits: masking spreads upward, so the part at the top of the texture has nothing above it to be masked by and the part at the bottom has everything. Perception and the listener

The listener is given the top voice, and the bass as a sine

Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.

The roughest chord on the page is at C1 and the roughest one heard is at A♭2. One close triad at 70 decibels through 17 registers, with its roughness computed twice: over every partial in the score, and over only the partials that stand above what the chord itself masks. The written curve rises all the way down and its maximum is the lowest register drawn, C1, which is the low-interval rule as it has always been computed here. The delivered curve turns over at A♭2 and falls to nothing below A♭1: a close triad down there is not rough, because it is not arriving as a chord — 1 of its 24 partials survives at C1 and there is almost nothing left to beat against anything. Perception and the listener

A low chord stops being rough by stopping being a chord

Every count of audible partials until now is of a chord at middle C, and the three registers it did compare span C3 to C5 — a third of the range a chord is written in. Move the same triad down and the count collapses: 79 per cent of its partials arrive at E3 and 4 per cent at C1. So the roughest chord on the page is the lowest one and the roughest chord a listener receives is at G2, and where that maximum sits moves nearly two octaves with the dynamic.

What the chord before takes out of the chord after. Five chords at 1.2 seconds each, with the roughness each one has on its own — its simultaneous masking and the threshold of hearing already applied — and the roughness it actually has once the chord in front of it has raised the threshold. Four of the five are untouched. The fifth, i, two parts, follows the only step in this passage that falls more than fifteen decibels, and it arrives into a hole: it is entirely below threshold for its first 13 milliseconds and takes 240 to get all of itself back. Masking can only remove partials, so it can only lower a roughness — and the chords it can reach are the ones a written dynamic has just made quiet, which are already the smooth ones. Across the passage the dissonance contrast goes from 3021 to 3113: the mask widens it by 3.0 per cent rather than eating it. Form and structure

A soft chord has to fade in

Forward masking sits between the two integration times already in play — two hundred milliseconds against a thirty-seven millisecond roughness window and a two-second loudness release — and it was owed as the term that might eat the dissonance contrast. It does not. It widens it, by three per cent at a chorale's pace and fifty-nine at four chords a second, because it can only ever remove partials and it can only reach the chord a dynamic has already made quiet. What it does instead is stranger: one chord in the passage is entirely inaudible for its first twelve milliseconds and takes a quarter of a second to arrive whole.

Which bars the key is decided by, and which bars it is believed on. Every bar of a 32-bar scheme removed in turn, with what its absence costs. The column is how far the passage's mean margin falls without that bar — how much of the model's certainty it supplies. The dot is how many bars are then read as a different key — how much of the answer it supplies. The 18 bars an analysis would point at — a section opening or closing, a dominant, the chord a dominant resolves to — average 0.079 bits of certainty and 0.50 bars moved; the 14 ordinary bars average 0.047 and 0.07. So the structural bars carry 1.7 times as much certainty as the ordinary ones, and 7.0 times of the answer. Those are different quantities, and the second is the one an analysis is about: an ordinary bar can carry a great deal of a passage's certainty and none of its reading. Harmony and voice leading

The bars a key is made of

A discount that treats every bar alike is the wrong shape for a memory, so take each bar away in turn and see what it was worth. The bars an analysis points at carry seven times as much of the answer as the ordinary ones and slightly less of the certainty — and a memory built only of them reads the passage worse than a memory with no structure in it at all.

A third of the timing surprise is paid for by nothing happening. Every beat of a bar of 4/4 at 1 chord change a bar, with what it costs a listener in expectation. The lower block is the arrival — the chance the chord changes here times what that change costs to be surprised by. The upper block is the hold, which is paid when the chord does not change and is charged at every beat rather than at every event. Over the bar the arrivals come to 2.61 bits and the non-arrivals to 1.29, so the term never spent before is 33 per cent of the total rather than most of it. The two together are exactly the binary entropy of each beat's own hazard, which is why the bar's whole timing bill is 3.90 bits and cannot be raised by rearranging where the changes fall. Harmony and voice leading

The surprise of nothing happening

A hazard charges a listener twice — once when the chord changes and once, quietly, at every beat it does not. The second term was expected to dominate, because there are more beats than changes. It is a third of the bill at one chord a bar and never reaches a half at any rate a metre survives, because the cost of a beat where nothing happens is second-order small.

Of 8 beats at A3, 3 can be attended to. Every member of the beat family a 2:1 mistuned by 6.0 cents makes on a real string at A3, placed by how separable its coincidence is from the partials beside it — in auditory filter widths, across — and by how fast it beats, up. The shaded region is the set a listener can receive: wider than one filter, and between 0.4 and 15 fluctuations a second. 4 of 8 members fall inside it, and once rates within a factor of 2 are counted as one modulation channel there are 3. The members that fail do so for two different reasons: the low ones sit under the roughness ceiling but their coincidences are buried in a filter that holds three partials, and the high ones are resolved and far too fast. Intervals and chords

Three beats at most, and only in the middle of the keyboard

A mistuned octave on a real piano makes eight beats at once, and a listener attending to one of them is doing something that has a threshold. Two thresholds, in fact — a rate and a place — and once both are applied the eight become four at A3, one at A1 and one at A5. Every interval a tuner sets goes to zero at both ends of the compass and peaks at eight countable beats in the octave the bearing is laid in.

6 step patterns, 2 category counts, and only the count matters. Each scale drawn as its own steps across the octave, with the share of trials a listener whose internal noise is 11 cents names correctly — and, beside it, the same figure for a scale of equally spaced degrees of the same size. The two agree to three decimal places on every row, though the narrowest step here is 100 cents on the diatonic major, tempered and the widest is 300. Naming is lost at boundaries, an n-degree scale has n of them wherever they are put, and each costs the same as long as no category is narrow enough for the noise to carry an estimate clean across it — 9.1 standard deviations, at the worst here. Scales and modes

A boundary costs the same wherever it is put

Every capacity figure drawn until now cuts the octave into equal categories, and no scale in the world is equal. Putting the real step patterns through the same model returns exactly the same accuracy to four decimal places — because naming is lost at boundaries, an n-degree scale has n of them wherever they are, and each costs the mean absolute value of the noise. That turns the most-quoted result about hearing from a search into one line: 1504 times one minus the criterion, over sigma.

How much of A4's pitch error a key could possibly remove. The largest correlation two notes' pitch errors can have at A4, against how long each note lasts. A note's error has two parts and only one of them is the listener's: the steady-tone limen of 4.04 cents, which context might reduce, and the bound a note of length T puts on its own frequency, which context cannot touch. Taking them in quadrature, the shareable fraction is the curve. At a quarter-second note it is 0.21, so the one-half priced earlier is not available at all until each note lasts 486 milliseconds — which is exactly the crossover found by a different route, because a correlation of a half is the two parts being equal. The step is the convention used here, which takes the larger of the two rather than their sum and therefore says the shareable fraction below the crossover is zero. Intervals and chords

The part of the error a key cannot touch

Four earlier essays turn one dial — the correlation between two notes' pitch errors — and apply it to the whole of a note's limen. Half of that limen is not the listener's: a note of finite length does not carry its frequency more finely than 1/2T, and no context can put information into a signal that is not there. So the correlation has a ceiling, it is 0.21 at a quarter-second note at A4 and 0.07 at A2, and the figure that prices a correlation of one half is drawn where one half is unavailable.

The top of a struck note's series falls while the note lasts. The highest partial still above the threshold of hearing, against time, for a note struck at 80 decibels on a fundamental of 130.8 hertz with a 1/n spectrum and a loss rising as the partial number to the power 1. It starts at partial 152 and it is falling from the first millisecond. The horizontal lines are the three tops computed from frequency alone, all of which assume a note that never ends: the note's own top drops past the difference-limen top after 0.03 seconds, past the semitone top after 0.32, and past the resolvable top after 0.64. After that the series is shorter than the ear could have resolved, and what stops it is the clock. Timbre and acoustics

The top that falls while the note lasts

Four earlier essays have asked where the harmonic series stops, and all four answered with a number computed for a tone that never ends. A struck C3 has 152 audible partials at the strike and eight after two-thirds of a second, so the ear's own resolution limit governs the first ten per cent of the note and the decay governs the rest. Playing ten decibels louder buys the ear a further tenth of a second, and doubling its reign would take fifty-eight.

A struck note's two ends are the same for every loss law. The partial levels of a string spectrum struck at 80 decibels on 130.8 hertz, and what is left of it when the fundamental itself falls under the threshold of hearing, for three laws relating a partial's decay rate to its number. The left panel is every one of them: a loss law cannot change the spectrum at the instant of the strike, because no time has passed. The other three are every one of them too: whatever the law, the note ends with nothing above the threshold. So both ends of the slide are shared, and everything that distinguishes an exponent of 0.5 from an exponent of 1 from an exponent of 2 is in the middle. Timbre and acoustics

The middle nobody could have guessed

A struck note has no steady state, only a slide from one spectrum to another — so the question is what the middle carries that the ends do not. The answer is exact rather than statistical: every loss law in the family leaves the strike with the same spectrum and ends in the same silence, so both endpoints carry precisely nothing about which of them it is. The whole difference is 41.3 decibels, and it peaks 0.38 seconds in, seven per cent of the way through the note.

There is less to collapse the higher the note is. The same string model struck at 80 decibels on nine fundamentals, an octave apart. The heavy line is how far the spectral centroid falls between the strike and the note's death, in semitones — the whole of the colour drain, in the unit a musician has for pitch. It runs from 25.2 semitones at A0 to 6.6 at C8, a factor of 3.8, and it falls monotonically. The reason is the dashed line, which is how many partials exist at all: 727 under the audio ceiling at A0 and 4 at C8. From C7 upward a struck note has fewer partials than the ear could have resolved — 9 against 10 — so there is nothing left for a collapse to take away. Timbre and acoustics

The collapse belongs to the bass

Every envelope drawn until now has no pitch in it: the decay model counts partials rather than measuring them in hertz. Put a fundamental in and the collapse it describes shrinks monotonically up the keyboard, from 25.2 semitones of colour drain at A0 to 6.6 at C8, because a partial has to fit under twenty kilohertz to exist and the top note has four of them. Above C6 a struck note has fewer partials than the ear could have resolved, and there is nothing left for a collapse to take away.

One attack time, three shapes, 48 ms of disagreement. Three amplitude envelopes with the same 90-millisecond attack time, which is the only quantity the published tables report. a resonator from a step rises as one minus a decaying exponential; an excitation ramping rises in a straight line; a ramp through a resonator is a raised cosine. Each is normalised so that its own 10-to-90 per cent rise takes exactly 90 milliseconds, so all three are the same measurement. The horizontal rules are the three criteria a heard moment is read off. At 6 dB below peak the three shapes put the heard moment at 28.5, 56.4, 76.3 milliseconds — a spread of 48, on one attack time, from a property nothing in the table records. Perception and the listener

An attack time is not an attack

Eight earlier essays read a heard moment off an envelope, and every one of them used the same envelope shape without saying so: the source table records one curve for all nine of its families, and the map's own arithmetic does not carry the parameter at all. A published attack time fixes a ten-to-ninety time and nothing else. Under the two other shapes the same measurement admits, every millisecond computed so far doubles — and the constructed passage called inaudible earlier becomes three times a listener's threshold.

A general pause of a bar is worth 4.9 decibels in a hall and 8.1 in silence. What a rest is worth as a diminuendo, in rooms with different reverberation times. The upper curve is the earlier figure, which assumed the sound stops when the players do; each curve below it lets the hall go on sounding, falling sixty decibels in its own reverberation time until it reaches the background. At the bar of silence a general pause usually is — about 1.2 seconds — the dry value is 8.1 decibels, shoebox concert hall keeps 60 per cent of it and gothic cathedral keeps 22. At 3.5 seconds, where a gap stops being a pause and becomes an ending, the same hall keeps 86 per cent. A long silence outlives any hall's tail and a short one does not, which is why the room costs the device most at exactly the length a composer writes it. Form and structure

A rest needs a dry room

The twenty-four decibels a general pause is worth assume the sound stops when the players do. Put a hall under it and a bar of silence keeps 60 per cent of its value in a shoebox concert hall and 22 per cent in a cathedral, while a three-and-a-half-second one keeps 86 and 45 — because a hall's tail has a length and a rest either outlives it or does not. A written bar of silence is worth half its dry value at 2.76 seconds of reverberation, which falls between the concert hall and the stone church.

Two registers 50 microseconds apart on one string. The partials of a 355-millimetre string plucked at 32 and 50 millimetres from the nut, drawn twice: once with the two quills releasing together, once with the far one releasing 0.05 milliseconds later, which is 0.026 of this string's period. A delayed release turns partial n through 2·pi·f_n·dt, so the rotation is proportional to the partial number: the fundamental is turned 9 degrees and partial 24 is turned 226. Simultaneous, the pair is missing partials 17 and 20 — holes it digs for itself where the two combs are equal and opposite. Staggered, it is missing none of them: a rotation of anything at all takes two amplitudes out of opposition. The partials neither comb can excite at all — 7, 11, 22 — are filled either way, because where one comb is zero the sum is the other one whatever its phase. The fundamental's gain over the far register alone falls from 4.36 decibels to 4.34. Instruments and their design

The interval between two quills

Two jacks on one key are voiced separately and do not let go at the same instant. That interval turns each partial of the later pluck through an angle proportional to its number — so it leaves the fundamental alone and inverts the twentieth partial, and the holes the pair digs for itself vanish at a hundredth of a period. What a regulator can tolerate turns out to be one fixed fraction of a period at every pitch, which on a five-octave instrument is a factor of sixteen in milliseconds.

A head model a few millimetres out reads every azimuth but the front. The azimuth a listener reports against the azimuth a source is at, for internal head radii from 8.22 to 9.28 centimetres against a true radius of 8.75. The delay a source produces is (r/c)(θ + sin θ) and the listener inverts it with the radius they believe they have, so their answer solves θ̂ + sin θ̂ = (r/r̂)(θ + sin θ). Every curve passes exactly through the origin: on the median plane there is no delay and therefore no error, whatever the head model is. The error grows with azimuth and is largest at the side. An internal head 5.3 millimetres too small runs out of azimuth at 82 degrees: beyond that the world is delivering a delay larger than any its owner's model can produce, and every source out there collapses onto the side. Perception and the listener

Where a wrong head gives itself away

Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.

Scored the way these figures score it, a clarinet is the worst of the six. One close triad at 70 decibels through 17 registers, drawn once for each of the 6 spectra to hand. The score is the share of every partial written, which is the quantity the register figure published. pure 100 per cent at best, string 79 per cent at best, clarinet 54 per cent at best, reed 71 per cent at best, bell 52 per cent at best, organ 72 per cent at best. A clarinet's four even partials are twenty-eight decibels below its odd ones and are inaudible beside their own neighbours before any chord is built, so counting them in the denominator makes the spectrum that survives its own masking best look like the one that survives it worst. Perception and the listener

A clarinet keeps what a string loses

Every masker, probe, chord, line and texture until now is eight partials falling as 1/n, and it was not even an option a placement could pass. Sweeping the six spectra to hand says the clarinet is the worst of them — 54 per cent of itself at best against a string's 79 — and that answer is an artefact of the score. Counted against what each note keeps on its own, the clarinet keeps 100 per cent where the string keeps 79, because its components stand a twelfth apart rather than an octave. The missing parameter was the spectrum; the second missing parameter was the denominator.

Thirty cents out of tune is heard as 8 on a 125 ms note and 25 on a 1 s one. How far out of tune a note sounds against how far out of tune it is, at 4 note lengths, at 440 hertz. The key is treated as a prior over pitch: a mixture of Gaussians on the twelve scale degrees, weighted by Krumhansl and Kessler's probe-tone profile and given the width the degree account already uses. The likelihood is the note's own effective limen, which for a short note is the Fourier bound 1/2T. The estimate is the posterior mean, and the shrinkage toward a prior is one line of arithmetic. At 1 s a thirty-cent mistuning is heard as 25.0 cents and at 125 ms as 7.8. Every curve turns back up near the middle of the semitone, because past there the nearest degree is the other one and the pull reverses. The buttons sound at A4, which is the pitch the figure is computed at, at the shortest note length it draws. Intervals and chords

A short note is heard more in tune than it is

Four earlier essays treat a key as something that reduces the noise in a pitch judgement. Treat it instead as a prior and the prediction changes kind: not a smaller error but a systematic bias, pulling a short note toward the nearest scale degree by an amount the Fourier bound sets. Thirty cents out of tune on an eighth-of-a-second note is heard as eight. And the part the debt got wrong is the part that matters — the bias does not vanish on a long note. It stops at 17 per cent at A4 and at 48 per cent at A2, because the likelihood's width has a floor that no duration removes.

Eighty-one chords the expectation model cannot tell apart. Every voicing of a dominant seventh on G inside the three octaves above its own root, placed by how rough it is and how far its outer voices are apart, and coloured by which member of the chord is at the bottom. The roughness runs from 0.269 to 1.449, a factor of 5.4, computed from each voicing's own spectrum under Plomp and Levelt's roughness model. The identity surprise the expectation model assigns is 3.51 bits for every one of the 81, because it is a function of a scale degree and its predecessor and there is no register anywhere in it. What separates them is spacing rather than inversion: roughness falls as the outer voices spread apart, correlating -0.42 with the span, and is indifferent to which member of the chord is at the bottom at 0.02. The seventh in the bass is not what makes a chord rough; a fourth and a third packed together at the bottom of the range is. Harmony and voice leading

Eighty-one chords, one number

A dominant seventh has eighty-one arrangements inside three octaves and their roughness spans a factor of five and a half. The tonal-expectation model gives every one of them the same 3.51 bits, because its states are scale degrees and there is no register anywhere in them. Conditioning the surprise on the voicing costs no corpus — and the arithmetic says the conditioning belongs beside the probability rather than inside it, for three reasons that can each be computed.

Both notes of a fifth, and the coincidences going out. The highest partial still above the threshold of hearing for each note of a fifth struck at 80 decibels on 130.8 hertz, against time, with a 1/n spectrum and a loss rising as the partial number to the power 1. Both curves come down from off the top of the frame — the lower note starts with 152 partials and the upper with 102, because twenty kilohertz is a ceiling in frequency and not in partial number. The rings are the interval's partial coincidences at the moment they stop existing — successive multiples of one ratio, which go out from the top down: 12:8 at 0.47 s, then 9:6 at 0.64 s, then 6:4 at 1.04 s, then 3:2 at 2.14 s. The lowest, 3:2, is the last, and after it the two notes have no partial in common that either of them can still supply. Timbre and acoustics

A fifth on a piano is not a fifth a second later

Nine essays draw one note in silence, and a listener is given a texture. Put two struck notes an interval apart and consonance comes apart into two quantities that had agreed while nothing moved: the partial coincidence that names the interval outlives the strike in exactly the order common-practice theory ranks its intervals, and the roughness that scores it reorders itself inside a third of a second, with the fifth overtaken by four intervals every treatise calls harsher.

Where a bar of 25 and a bar of 9 first disagree. The additive metre 2+2+2+3+2+2+2+3+2+2+3 — 25 units, an onset at the head of every group — with the accents it predicts drawn above the accents predicted by reading it as a repeating bar of 9, which is the cut of it that agrees longest. The two rows are identical for 24 consecutive steps and differ for the first time at step 25, where the shorter reading expects an accent and the metre does not supply one. Nothing before that step distinguishes the two hypotheses, so a listener who has not heard 25 consecutive steps has no evidence either way — whatever they are disposed to hear. Rhythm and metre

A twenty-five is a nine until its last unit

Every account of long additive metres says they are heard as groups of shorter ones, and the metre-induction model had never been pointed at the claim. Pointed at it, the model does not prefer the group — it prefers the long bar outright, and would go on preferring it more the longer anybody listened. What it cannot do is start: the evidence that separates a bar of twenty-five from a bar of nine does not exist until the whole bar has been heard, and at the tempo an unequal metre is best played at the psychological present holds sixteen units.

A bar of silence is worth 13.1 decibels written last and 6.0 written first. Four closing gestures, drawn against how many seconds of silence each contains. All four hold the same 12-part texture, write the same 6.0-decibel diminuendo over 4 seconds, and differ only in what order the diminuendo and the silence are written in. The quantity is how far the listener's running impression has fallen when the final chord arrives, in decibels of equivalent diminuendo. Written with the silence last, a general pause of a bar is worth 13.06 decibels and one of three and a half seconds is worth 28.3. Written with the silence first, both are worth 6.00 — exactly the diminuendo's own depth, because the music resuming after the silence puts the reference back at its own level. The silence written first is worth less than the silence written with no diminuendo at all, which reads 8.12 at a bar: a diminuendo placed after a general pause takes 2.12 decibels away and adds nothing. Form and structure

A general pause is spent by the note after it

Whether a composer should write the pause before the diminuendo or after it looks like a question about how big the ensemble is. It is not. Forty decibels of ensemble are worth one decibel of silence, and the order is worth seven — because a running impression rises twenty times faster than it falls, so half a general pause is spent by ninety-seven milliseconds of sound.

A note on the downbeat costs 1.89 bits and one on the offbeat 5.70. What each position in a bar of 4/4 asks of a reader, by two routes. The solid bar counts where the notes of this collection's own three tunes actually fall — 28, 2, 26, 4, 28, 0, 16, 0 notes at the 8 positions — and takes minus the log of the frequency. The rule across each bar is the same quantity from the stated metrical weights, 1, 0.15, 0.5, 0.15, 0.85, 0.15, 0.5, 0.15, normalised and logged the same way. Nothing makes the two agree. They put the eight positions in the same order, and they price the tunes' own rhythm a fifth of a bit apart — while differing by more than a whole bit about the quaver after the downbeat, which two notes in a hundred and four ever use. The two positions these tunes never touch at all are drawn at the floor, which is the same floor the melodic measure gives an interval nobody plays. A weight was always a probability waiting to be read as one. Rhythm and metre

Where the note is costs more than which note it is

Thirteen earlier essays measure a page, and the three that price a reader price only its pitches — every line they measure is a run of equal notes. A metrical weight normalised by its own sum is a probability, and minus its logarithm is bits — the same substitution made earlier for intervals. Measured over the tunes used throughout it comes out at 2.23 bits a note against the pitches' 1.89, so the larger half of a reader's load is where the note is.

Every beat in a family is the same depth, and none is near the threshold. The 8 members of the beat family a octave mistuned by 6.0 cents makes at A3 on a real string, each placed at the rate it beats and at the modulation index a listener's filter delivers there. The rising curve is the published detection threshold for amplitude modulation, which is flat at 0.03 below about fifty fluctuations a second and rises above it. The flat dashed line at 0.667 is the depth the pair has in isolation, and it is the same for every member: on a spectrum falling as one over n, the k-th member pairs partials 2k and 1k, whose ratio is 2.00 whatever k is. What the filled points show is the smaller effect that does depend on the member — the partials on either side of the coincidence leak into the same filter, add level without adding fluctuation, and dilute the index from 0.662 to 0.405. Inside the countable rate window the narrowest margin over the threshold is a factor of 16.7, on member 4. The depth criterion removes nothing. Intervals and chords

Every member of a beat family is the same depth

A mistuned octave's beats have been counted on their rate and their place, with the depth recorded as the thing left out and a prediction that the shallow upper members would take the count from four to two. The depth turns out not to fall at all: on any power-law spectrum every member of a family has exactly the modulation index its interval's own ratio gives, at every register, on every wire. The count does fall to two, and the thing that takes it there is the criterion that essay was already using.

Three clocks receive one entrance, and they do not agree about when. An oboe joining 5 players already sounding, at time zero, with each of the listener's three readings drawn as its own share of the change it eventually makes. The roughness window is 49 milliseconds wide and has half the change at 25; the short-term loudness smoother has half at 15; the long-term one, whose release is the two seconds an earlier essay is about, has half at 90. The two-second release is on the wrong side of the smoother to hide an entrance. Its attack is 99 milliseconds, so an entrance is received promptly and it is a departure that is not. Form and structure

The release is on the wrong side

Whether the loudness model's two-second release makes an entrance inaudible has the answer no, for a reason the question did not anticipate. The smoother is asymmetric — ninety-nine milliseconds going up and two seconds coming down — so a rise is tracked twenty times faster than a fall, and an entrance is received promptly by every one of a listener's three readings. The colour of it arrives first, at twenty-five milliseconds against ninety, and the reading that moves with the ensemble is the one nobody would have picked.

The same player, arriving and leaving, read as a share of the change. One oboe joining 5 players and the same oboe leaving them again, with both loudness readings drawn as the share of their own change that has arrived. The entrance is half received in 90 milliseconds and the exit in 1.43 seconds, a factor of 15.9. The roughness readings, drawn faintly, are 25 and 25 milliseconds and lie on top of each other. A score that writes a diminuendo under a departing part is not softening the exit. It is doing the smoother's release for it, on a clock the smoother would otherwise take two seconds over. Form and structure

A part that leaves is not a part that arrives

The same player, the same note, the same level, and the only difference is which way round it happens. A listener's loudness reading takes 1.43 seconds to receive half of a departure and 90 milliseconds to receive half of an arrival — a factor of sixteen with nothing asymmetric in the sound at all, since both readings integrate the same two states in the same order. The colour reading receives the two identically, because a window has no direction, so a departure is a change whose grain arrives at once and whose level takes most of two seconds.

The same interval, mistuned by the same amount, at each of its two ends. A C to G in the major key, played 25 cents wrong, with the departure carried by the lower note, split between the two, and carried by the upper note. All three are the same interval size; what differs is which note is off the scale. The share of the departure that survives into what a listener hears is 32 per cent when the lower note carries it and 56 when the upper does. The middle bar is the mean of the other two to within a hundredth, so the averaging is linear and the asymmetry is the whole of the effect. Two things produce it: the prior is 9.0 cents wide at the C and 10.0 at the G, and the likelihood is 13.2 cents wide at the lower pitch and 8.8 at the higher. Intervals and chords

An interval is two posteriors subtracted

Treating a key as a prior over one note predicts that an interval's pull is not the single-note pull doubled, because the two degrees are not equally weighted. Half of that is wrong: splitting a mistuning between the two notes gives exactly the mean of what each end gives alone, to a thousandth, at every one of the twenty-one intervals in the scale. What is not the mean is which end carries it — and the pull turns out to be largest not on the shortest notes but on notes of about an eighth of a second, where the likelihood is a quarter of a semitone wide.

Which intervals in a key can be mistuned invisibly, and from which end. Every interval between two degrees of the major scale, ranked by how differently its two ends treat a 25-cent departure on notes of 0.25 seconds. The C to B is the most lopsided, at 47 points: a mistuning on its upper note reaches the listener nearly 2.5 times as strongly as the same mistuning on its lower one. The D to E is the most even, at -0. A negative bar is an interval whose LOWER note is the one that carries a mistuning into the listener, which happens whenever the lower degree is the less specified of the two. No account of interval perception predicts a table like this, because an interval is usually treated as one quantity rather than as a difference of two estimates. Intervals and chords

Which end the mistuning is on

Twenty-one intervals in the major scale, each with two ends, and the same twenty-five cents reaches a listener at anywhere between 32 and 79 per cent of its size depending on which of the two notes carries it. The most lopsided is the tonic to the leading note, where a departure on the upper note arrives two and a half times as strongly as the same departure on the lower. Two mechanisms produce it and they can be separated by one flag: two thirds of the asymmetry is register and one third is the key.

The setting a listener reports is not the interval they preferred. Five published preferences for an interval size, and the value a listener would have to play in a key context for that preference to be what they hear. The key pulls a heard interval back toward the scale, so a listener adjusting until it sounds right has to overshoot — and the reported setting, which is what they played, exaggerates the preference. The corrections run from 1.6 cents to 12.8 on notes of 0.25 seconds. The largest is nearly a syntonic comma, on a preference of thirteen cents, which is to say that the correction is bigger than the effect it is a correction to. Intervals and chords

The setting is not the preference

Every published number for an interval listeners prefer — the pure third a quartet is said to find, the raised leading note, the harmonic seventh — is a value somebody adjusted until it sounded right. A listener adjusting inside a key is adjusting through the posterior computed just before, so the value they stopped at is not the one they preferred: it is the one whose heard size equals it. On quarter-second notes the correction for a pure major third is 12.8 cents, which is nearly a syntonic comma and is larger than the 13.7-cent preference it corrects.

Two ways to match a tuning note, and they are not the same size. How finely one player can put their A on another's, at 440 hertz on a note of 2 seconds, by each of the two criteria available. Judging one pitch is good to 4.0 cents; comparing two of them adds two errors in quadrature and is good to 5.7. Nulling the beat between them is a different operation altogether — a mistuning of 2.0 cents makes a beat of 0.50 hertz, which shows one full cycle inside the note — and it is 2.9 times finer. The beat criterion is available only when the two tones sound together and share a partial. A player tuning to a note that has already stopped has the coarse one, and so does a singer with nothing to beat against. Pitch and tuning

An orchestra is given a note

Every earlier essay on pitch standards draws a standard as a number an ensemble is at, and no ensemble is at a pitch. It is handed one, by one player, on one note, and everything else is matched to it by ear — so a standard reaches an orchestra through a limen nobody quotes. Matching by comparing two pitches is good to 5.7 cents at A; nulling the beat between them is good to 2.0, and which of the two is available depends on how long the oboe holds the note. The crossover is at about seven tenths of a second, which is shorter than an oboe's A and longer than a plucked one.

The ensemble agrees with itself more and more, about a pitch that is moving. Runs of 16 players each correcting toward the mean of their neighbours, with no term anywhere pulling them back to the note they were given, over 480 corrections. The shaded band is the root-mean-square displacement across all eight runs — the envelope a random walk has — and it grows from 0.84 cents a quarter of the way through to 2.01 at the end, which is the square-root growth a random walk has. Four individual runs are drawn inside it and the furthest of the eight over the top, ending at 4.16 cents. Meanwhile the spread AMONG the players falls from 2.8 cents to 0.5. A consensus with no anchor cannot hold a pitch, and it also cannot lose one quickly: a movement's worth of corrections is a few cents rather than the semitone unaccompanied choirs are said to fall by. Pitch and tuning

A consensus with nothing to hold it

Once the oboe has stopped, no reference is left in the room. Each player corrects toward what they hear around them, which is other players correcting toward them — and a consensus dynamic has a fixed point at every common value, so it pulls the ensemble together and nothing pulls it anywhere in particular. Simulated, the players' spread falls from 2.8 cents to 0.5 while the ensemble as a whole random-walks. The size is the result and it is small: two or three cents over a movement, which is a tenth of what unaccompanied choirs are said to lose.

One noise everywhere, and the noise each boundary actually gets. The seven boundaries of the tempered diatonic scale, with the noise the earlier model gives each of them — 11 cents, the same everywhere — and the noise the harmonicity model gives them instead, which runs from 11.0 cents to 35.6. The error rate is the sum of these over the octave rather than seven copies of one, so it rises from 5.1 per cent to 11.8. And the moment the boundaries differ, where they are put matters: a scale that moved its degrees would move its boundaries onto different intervals and would pay a different sum. That is the earlier null broken by the assumption its own last section named. Scales and modes

A boundary beside a fifth

A closed form established earlier says an n-category division costs n·σ·√(2/π) whatever the widths are, so a boundary costs the same wherever it is put and the step pattern cannot matter. Its own last section named the assumption that produces the null: one σ, applied to every boundary in the octave. Let σ follow how securely each interval is held and the formula becomes a sum over boundaries rather than n copies of one — the diatonic's error rises from 5.1 per cent to 11.8, and where the degrees are put matters again.

The best seven of the twelve is a scale nobody has ever used. All 462 ways of choosing seven of the twelve semitones with the tonic fixed, ranked by the identification error the harmonicity model gives them. The best is C C♯ F♯ G A♭ B♭ B at 8.8 per cent and the worst is 13.4; the diatonic major sits at rank 376, in the worse fifth of the ranking, at 11.8. The optimum is a cluster of semitones around the tonic and around the fifth, and the reason is visible in the criterion rather than in music: a boundary next to the unison or the fifth is a boundary with very little noise on it, so the cheapest way to satisfy this measure is to crowd the degrees where the model says the ear is sharpest. A criterion whose optimum is a scale nobody plays is a criterion that is not what scales are chosen for, and the useful reading of this drawing is that rather than its winner. Scales and modes

The best seven of the twelve

Once the noise is allowed to differ from boundary to boundary, a scale can be chosen to minimise identification error — and the choice is a search over four hundred and sixty-two sets rather than an argument. Run, it returns a cluster of semitones around the tonic and around the fifth, and puts the diatonic major at rank 376 of 462, in the worse fifth of the ranking. A criterion whose optimum is a scale nobody has ever played is a criterion that is not what scales are chosen for, and the reason it fails is legible in the model rather than in the music.

Each tradition's own steps against the equal division of the same size. Six scales, each drawn against the equal division into the same number of degrees, under both models of how securely an interval is held. On the harmonicity model the tempered diatonic is 11 per cent better than seven equal steps; on the profile model the same comparison is 0.9 per cent, which is nothing. The two models disagree about the one comparison anybody would want the measure for, and only one of them is free of circularity: the probe-tone profile was measured on listeners raised inside the diatonic tradition, so using it to explain why the diatonic is well chosen assumes the answer. The harmonicity model assumes only that a simple ratio is easier to hold than a complicated one. Scales and modes

The unequal scale that is easier to name

The whole of the earlier result was that a scale's step pattern cannot matter, and every tradition it drew agreed with the equal division of its own size to three decimal places. With one noise per boundary the comparison is live again, and the tempered diatonic beats seven equal steps by eleven per cent — while the pentatonics gain nothing and the maqam scales gain two tenths of one per cent. Only one of the two security models produces the effect, and it is the one that was not measured on listeners raised inside the tradition it is being used to explain.

A struck octave becomes countable a second after the strike, or never. The number of separable beats a mistuned octave on A3 delivers, second by second after both notes are struck at 80 decibels, with every partial dying at its own rate (a 12-second fundamental, losses rising as frequency to the power 0.7). Read with the filter broadened by the level of the whole note, which is how the level-dependent count was first computed, the count is zero at every instant: the partials fall below audibility before the filter has narrowed enough to separate them. Read with the filter broadened by the level inside itself, which is what the published parameterisation was fitted against, the count is 0 at the strike, reaches 2 at 1.0 s and falls to nothing at 3.3 s. Pitch and tuning

Counted in the decay, or not at all

A mistuned octave struck hard delivers no countable beat at the strike, and the reconciliation offered for that was that a tuner listens to the decay. Computed through a real decay it fails on its own terms: the partials fall silent before the filter has narrowed enough to separate them. It succeeds only when the filter is broadened by the level inside it, which is what the published parameterisation was fitted against — and then the window opens at a twelfth of the note's life and shuts at a quarter.

A tempered interval moves its difference tone several times further than itself. For every interval inside the octave tuned to twelve equal steps, how far the difference tone f₂ − f₁ lands from where the just interval would put it, in cents, with the interval's own departure from just drawn as the thin bar beside it. minor second −198.0 (the interval −11.7); major second −35.5 (the interval −3.9); minor third −96.0 (the interval −15.6); major third +67.4 (the interval +13.7); fourth +7.8 (the interval +2.0); fifth −5.9 (the interval −2.0); minor sixth −36.7 (the interval −13.7); major sixth +38.8 (the interval +15.6); minor seventh −39.8 (the interval −17.6); major seventh +25.0 (the interval +11.7). The major third's product is +67.4 cents out and the minor third's −96.0, and the largest error is the minor second's, at −198: a product moves p/(p − q) times as far as the interval p:q that made it. Intervals and chords

The third sound magnifies cents, not hertz

Tartini's third sound is said to be a few cents off on a tempered interval. It is sixty-seven cents off on a major third and ninety-six on a minor third, because a difference tone moves p/(p − q) times as many cents as the interval p:q that made it. In hertz it moves exactly as far as the note that moved, and no further — so what the magnifier is worth is the ear's finer resolution at the low frequency where the product lands, which is a factor of two for a long note and nothing at all for a short one.

A tempered fifth is a beat a second on the violin and one in four and a half seconds on the cello. How fast each open fifth of a string quartet beats when it is narrowed by 1.955 cents, the narrowing that meets an equal-tempered keyboard. The beat is the lower string's third partial against the upper string's second, so it is proportional to the lower string's frequency. Cello C2–G2: 0.22 a second, one beat every 4.5 seconds; cello G2–D3: 0.33 a second, one beat every 3.0 seconds; cello D3–A3: 0.50 a second, one beat every 2.0 seconds; viola C3–G3: 0.44 a second, one beat every 2.3 seconds; viola G3–D4: 0.66 a second, one beat every 1.5 seconds; violin G3–D4: 0.66 a second, one beat every 1.5 seconds; violin D4–A4: 1.00 a second, one beat every 1.0 seconds; violin A4–E5: 1.49 a second, one beat every 0.7 seconds. The slowest, the cello's C2–G2, is 6.7 times slower than the violin's A4–E5. Pitch and tuning

The cello cannot hear its own tempering

Narrowing a quartet's fifths to meet a piano is one number, 1.96 cents a fifth, and it is a different beat on every string: once every two thirds of a second on the violin's A–E and once every four and a half seconds on the cello's C–G. Set by ear for two seconds a fifth, the violin's E lands within two thirds of a cent and the cello's C within 5.7 — which is as large as the Pythagorean error the tempering was meant to remove. The string whose tuning is most wrong is the string whose tuning is least certain, and a cellist tuning down the chain cannot tell pure from tempered.

An inversion lasts as long as its outer sixth. The six three-note voicings of a major and a minor triad, each over a bass of C3 struck at 80 decibels, with how long the strongest partial coincidence of each of its three intervals survives the strike. The interval that goes first is marked, and its time is how long the chord keeps the evidence of all its intervals at once. major, root position: major third 5:4 1.26 s, minor third 6:5 1.03 s, fifth 3:2 2.14 s; the chord 1.03 s. major, sixth chord: minor third 6:5 1.04 s, fourth 4:3 1.62 s, minor sixth 8:5 0.74 s; the chord 0.74 s. major, six-four: fourth 4:3 1.60 s, major third 5:4 1.27 s, major sixth 5:3 1.26 s; the chord 1.26 s. minor, root position: minor third 6:5 1.04 s, major third 5:4 1.27 s, fifth 3:2 2.14 s; the chord 1.04 s. minor, sixth chord: major third 5:4 1.26 s, fourth 4:3 1.63 s, major sixth 5:3 1.26 s; the chord 1.26 s. minor, six-four: fourth 4:3 1.60 s, minor third 6:5 1.03 s, minor sixth 8:5 0.74 s; the chord 0.74 s. The major six-four lasts longest and the major sixth chord shortest; every root position is held to its minor third's life. Timbre and acoustics

An inversion lasts as long as its outer sixth

A struck interval keeps the partial coincidence that names it for a time set by its ratio, and a chord is three intervals at once. Voiced over one bass and struck on a piano, a triad keeps the evidence of all three only as long as its weakest one lasts, and for an inversion that is the sixth on the outside: a major sixth lasts as long as a major third, a minor sixth dies first. So the major six-four and the minor sixth chord are the most durable voicings of their triads and the major sixth chord and the minor six-four the least — and unlike a dyad, a triad's inversions keep their order by roughness through almost the whole decay.

An accent moves the longest separable bar from 16 units to 18, and no further. For every bar length from 9 to 25 units, over all 1820 arrangements of twos and threes that are not a repeat of a shorter bar, the fewest and the most consecutive steps before the whole bar beats every shorter cut of it. 9: onsets alone 9 to 14, long beat predicted 7 to 12, downbeat predicted 7 to 8; 10: onsets alone 10 to 13, long beat predicted 8 to 11, downbeat predicted 8 to 9; 11: onsets alone 11 to 18, long beat predicted 9 to 16, downbeat predicted 9 to 10; 12: onsets alone 12 to 17, long beat predicted 10 to 15, downbeat predicted 10 to 11; 13: onsets alone 13 to 22, long beat predicted 11 to 20, downbeat predicted 11 to 12; 14: onsets alone 14 to 23, long beat predicted 12 to 21, downbeat predicted 12 to 13; 15: onsets alone 15 to 26, long beat predicted 13 to 24, downbeat predicted 13 to 14; 16: onsets alone 16 to 25, long beat predicted 14 to 23, downbeat predicted 14 to 15; 17: onsets alone 17 to 30, long beat predicted 15 to 28, downbeat predicted 15 to 16; 18: onsets alone 18 to 29, long beat predicted 16 to 27, downbeat predicted 16 to 17; 19: onsets alone 19 to 34, long beat predicted 17 to 32, downbeat predicted 17 to 18; 20: onsets alone 20 to 35, long beat predicted 18 to 33, downbeat predicted 18 to 19; 21: onsets alone 21 to 38, long beat predicted 19 to 36, downbeat predicted 19 to 20; 22: onsets alone 22 to 37, long beat predicted 20 to 35, downbeat predicted 20 to 21; 23: onsets alone 23 to 42, long beat predicted 21 to 40, downbeat predicted 21 to 22; 24: onsets alone 24 to 41, long beat predicted 22 to 39, downbeat predicted 22 to 23; 25: onsets alone 25 to 46, long beat predicted 23 to 44, downbeat predicted 23 to 24. A present of 3.5 seconds holds 16.0 steps at 218 milliseconds a step, so the longest bar some arrangement of which separates inside it is 16 units on onsets alone, 18 with the long beat predicted and 18 with the downbeat predicted. Rhythm and metre

The accent buys two units, however loud it is

A twenty-five cannot be told from a group of shorter bars on its onsets until more steps have gone by than a listener's present holds, and the obvious objection is that nobody plays an aksak bar as bare onsets: the long beat is louder, and the bar's first beat is marked. So how loud does an accent have to be? The question has a surprising answer. Loudness is not the variable. The existing accent cue changes nothing, and delays the answer where it changes anything. An accent that a reading has to predict works at any strength at all, and at no strength does more than a fixed amount: on the long beat it buys the two steps of a short beat, and on the downbeat it takes every arrangement to one floor — the bar less its last beat — which no cue carried by the notes can break. The longest bar that can be heard as one moves from sixteen units to eighteen.

Six named proportions, as blurred as the durations that make them. Six proportions between two parts of a piece — 1 : 1, 4 : 3, 3 : 2, golden section, 2 : 1, 3 : 1 — placed on one axis by the logarithm of the ratio of the longer part to the shorter, and drawn as bars one criterion wide (d′ = 1) for a listener timing both parts with a Weber fraction of 7%, 15%, 35%. Bars that overlap are proportions that listener cannot tell apart. At 7%, 4 of 5 neighbouring pairs stay apart; at 15%, 3 of 5 neighbouring pairs stay apart; at 35%, 0 of 5 neighbouring pairs stay apart. Form and structure

A proportion is only as fine as its two durations

Analyses of form measure proportions in bars and report them to three figures — a climax at 0.618, a section in the ratio 3 : 2. A listener has each part only as an estimate of how long it lasted, and a ratio of two estimates is blurred by both. Timed as well as anyone times a single second, eleven proportions fit between 1 : 1 and 3 : 1; timed from memory over minutes, two do. The golden section is told from 3 : 2 only below a Weber fraction of 5.4 per cent.

The same forms by the clock and by what is stored. Six forms, each drawn twice: its sections sized by their share of the bars, and sized by their share of what a listener has to store when a bar counts only if it is recognised from 1 bar of context. Returns are drawn pale with a dashed edge. twelve-bar blues: returns take 67 per cent of the clock and 13 per cent of the storage; thirty-two-bar AABA: returns take 25 per cent of the clock and 5 per cent of the storage; rondo, ABACA: returns take 40 per cent of the clock and 10 per cent of the storage; verse and chorus: returns take 50 per cent of the clock and 13 per cent of the storage; two eight-bar phrases: returns take 0 per cent of the clock and 0 per cent of the storage; a four-bar ostinato: returns take 88 per cent of the clock and 0 per cent of the storage. Form and structure

A return is shorter than its first hearing

A rondo's refrain takes three fifths of the clock and a verse-and-chorus song is balanced to the bar. Count instead the bars a listener could not have predicted when they arrived, and the returns shrink to between a tenth and a quarter of what is kept — so a song equal by the clock is between three and seven times heavier in its first half. A coder that learns repeats one bar at a time says the halves are equal, and the two memories disagree by more than any proportion a listener could confuse.

A final chord stands above the impression for a fraction of a second. How far a final chord at the tutti's own level stands above the listener's running impression at the instant it is released, against how long it lasts, for four ways of arriving at it. Straight out of the tutti it stands above nothing at any length; after 1.2 s of silence the impression is 8.1 dB down, and the chord stands highest, 5.14 dB, when it lasts 54 ms; after 3.5 s of silence the impression is 23.7 dB down, and the chord stands highest, 14.42 dB, when it lasts 28 ms; after a 6 dB diminuendo the impression is 5.0 dB down, and the chord stands highest, 3.18 dB, when it lasts 42 ms. Every curve is level again by half a second, because the impression's attack of 99 ms catches the note's attack of 22 ms, so a held chord is released at the impression's level whatever preceded it. Form and structure

A final chord stands out for a twentieth of a second

A general pause drives a listener's running impression down, and the final chord that follows is supposed to cash the fall in. It cashes in at most two thirds of it. The note's own loudness rises with a 22-millisecond constant and the impression with a 99-millisecond one, so after a bar of silence the chord stands furthest above the impression 54 milliseconds in, by 5.1 of the 8.1 decibels the silence bought, and after 206 milliseconds the two are within a phon of each other. A short stamp spends most of its life standing out; a chord held a second and a half spends a seventh of it.

Level does not dilute the register's roughness, it multiplies it. The mean roughness of the I – vi – IV – V – I arrivals at 4 registers, each relative to the register as written, read three ways. Level-free, the bass is 8.6 times rougher than the treble. With every note at 70 dB it is 8.6 times, the same factor, because one level rescales every pair alike. With each chord played at the level that makes it as loud as the written register's chords — 82.8 dB −2 octaves, 75.5 dB −1 octave, 70.0 dB as written, 66.8 dB +1 octave — the bass is 343 times rougher than the treble, because roughness grows with the square of the pressure and the bass needs more of it to be heard at the same loudness. Harmony and voice leading

A rough arrival is rough because of its spacing

The pair the expectation essays report for every chord — how surprising it was, how rough its voicing is — has no level in it. Putting level back in answers the question it left open, and not the way it was framed. At one written dynamic the arrivals keep their order from 40 to 90 dB at three registers of four, and the bass stays 8.6 times rougher than the treble. Made equally loud, the bass has to be played 12.8 dB harder, and it is 343 times rougher: level does not explain the register's roughness away, it multiplies it.

Opening a triad changes which interval goes first. The six voicings of a major and a minor triad over C3, each close and with its middle note raised an octave, struck at 80 decibels, with how long each keeps the coincidences of all three of its intervals and which interval goes first. major root position: close 1.03 s, held by its minor third 6:5; open 1.26 s, held by its major tenth 5:2. major sixth chord: close 0.74 s, held by its minor sixth 8:5; open 0.47 s, held by its minor tenth 12:5. major six-four: close 1.26 s, held by its major sixth 5:3; open 0.74 s, held by its eleventh 8:3. minor root position: close 1.04 s, held by its minor third 6:5; open 0.47 s, held by its minor tenth 12:5. minor sixth chord: close 1.26 s, held by its major third 5:4; open 1.26 s, held by its major sixth 5:3. minor six-four: close 0.74 s, held by its minor sixth 8:5; open 0.74 s, held by its minor sixth 8:5. Intervals and chords

An open triad lasts as long as its tenth

Close, every triad's weakest link is a third or a sixth. Raise its middle note an octave and the link becomes a compound interval, and compound intervals do not last alike: a major tenth, 5:2, lives as long as a major sixth, while a minor tenth, 12:5, needs the twelfth partial and lives 0.47 seconds. So the spacing orchestration manuals recommend for a major chord in the bass is the longest-lived voicing a struck triad has, and the same spacing halves the life of a minor chord — and the six-four's advantage reverses.

An unaccompanied quartet settles where its open strings put it. The average pitch of a quartet correcting toward itself over 480 corrections, in cents from the note it was given, averaged over 24 runs. With no pull from the open strings the ensemble random-walks, and the shaded band is how far: 3.7 cents root-mean-square by the end. With each open string pulling the notes that share its pitch class at a weight of 0.05, the ensemble settles at −0.97 cents in A major, against −1.01 from the open strings' weighted mean; −2.46 cents in C major, against −2.42 from the open strings' weighted mean; −2.53 cents in E♭ major, against −2.54 from the open strings' weighted mean. Pitch and tuning

An open string pulls the quartet flat

Once the tuning note has stopped, a quartet corrects toward itself and nothing holds its pitch. But four of its pitches do not move: the open strings, on a Pythagorean chain from C 5.9 cents flat to E 2.0 sharp, each ringing when a stopped note shares its pitch class. Give that sympathy a weight of a hundredth of a correction and it beats the random walk within a movement. The quartet settles flat in every major key — by 0.9 cents in E, 2.5 in C and 2.6 in A♭ — and in A♭ major the cellist's tuning scatter moves the whole ensemble by 1.7 cents.

A string that decays twice opens its count at once and shuts it early. The count of separable beats a mistuned octave on A3 delivers after both notes are struck at 80 decibels, with each filter read at the level inside it, under one exponential decay of 12 seconds and under two stages — a prompt sound of 1.5 seconds carrying all but the last 20 decibels, and an aftersound of 12 seconds. One exponential: open from 1.00 s to 3.30 s, 11.8 beats. Two stages: open from 0.15 s to 1.77 s, 8.8 beats — and the single exponential struck 20 decibels softer closes at 1.77 s. Pitch and tuning

A string that decays twice is counted early

A mistuned octave's beats were found countable only between a twelfth and a quarter of a note's life, on a note decaying once. A piano string decays twice, a fast prompt sound over a slow aftersound, and the prediction was that this would open the count sooner and close it later. It opens sooner — at a seventh of a second rather than a second — and closes exactly where a single decay struck twenty decibels softer closes, so at 80 dB it holds 8.8 beats instead of 11.8. The count now rises with the strike to 90 dB, and a tuner who strikes hard is right.

A listener who knows every metre recognises none of them inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps from the downbeat before every other one of the 1820 arrangements of twos and threes has been contradicted by the stream, on onsets alone, with the long beats accented and with the downbeat accented, against the 16 steps a present of 3.5 seconds holds. onsets alone: recognised within the present for 0 of 1820; long-beat accent: recognised within the present for 0 of 1820; downbeat accent: recognised within the present for 85 of 1820. The dashed line is the present. Rhythm and metre

Knowing every metre is slower than knowing none

A long aksak bar cannot be told from its shorter cuts by induction before one step into its last beat, and no accent carried by the notes moves that floor. The obvious escape is a listener who knows the repertoire and recognises the metre instead. Recognition among all 1,820 arrangements of twos and threes never beats the floor, is never quicker than induction, and is slower for half the metres: a nine induced in 9 steps is recognised in 27. What breaks the floor is a small repertoire that leaves out the metre's own longest cut — with the cut known, no repertoire of any size does.

With a fading memory the slowest order of the gaps 1 1 2 2 2 2 2 is not the one a perfect memory finds. For each of the 3 cyclic orders of the gaps 1, 1, 2, 2, 2, 2, 2 in 12 steps, the bits of position still unknown once listening has settled, against how many steps a listener's memory of a step takes to halve, with a mismatch costing 6. 2 2 2 1 2 1 2: 12 → 4e-9, 6 → 4e-5, 4 → 0.007, 3 → 0.070, 2 → 0.540, 1.5 → 1.130, 1 → 1.786; 2 2 2 2 1 1 2: 12 → 2e-5, 6 → 0.035, 4 → 0.335, 3 → 0.795, 2 → 1.405, 1.5 → 1.694, 1 → 2.005; 2 2 1 2 2 1 2 (the standard bell pattern): 12 → 2e-5, 6 → 0.027, 4 → 0.231, 3 → 0.535, 2 → 1.028, 1.5 → 1.368, 1 → 1.819. With perfect memory the slowest to locate is the standard bell pattern; at a half-life of 6 steps the highest floor is the order 2 2 2 2 1 1 2, and at 1 it is the order 2 2 2 2 1 1 2. Rhythm and metre

The bell pattern is slowest only to a perfect memory

Among the orders of its own gaps, a named timeline is usually both the most even and the slowest to locate — for a listener who never forgets. Give the listener a memory that halves and the result comes apart. Of six timelines slowest among their orders with perfect memory, only the fume-fume stays slowest for every forgetting listener, and the standard bell pattern, which is the fume-fume with onsets and rests exchanged and settles at exactly the same floors, is second of its three orders for every memory of half its cycle or less. The census ranking survives better, and in fourteen of twenty-one censuses it was the arithmetic of a pattern that repeats.

Come in part-way with the downbeat accented, and no bar of sixteen units or more is recognised inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps a listener who knows every arrangement of twos and threes needs to recognise the metre and where its bar begins, with the downbeat accented: coming in at a sample of steps inside the bar, against hearing it from its written downbeat. On onsets alone, or with the long beats accented, a metre entered part-way is never told from its rotations. 9: from inside the bar 10 to 17, 16 of 20 inside the present; from the downbeat 10 to 10; 10: from inside the bar 11 to 19, 12 of 20 inside the present; from the downbeat 11 to 11; 11: from inside the bar 12 to 21, 9 of 18 inside the present; from the downbeat 12 to 12; 12: from inside the bar 13 to 23, 8 of 24 inside the present; from the downbeat 13 to 13; 13: from inside the bar 14 to 25, 8 of 28 inside the present; from the downbeat 14 to 14; 14: from inside the bar 15 to 27, 4 of 28 inside the present; from the downbeat 15 to 15; 15: from inside the bar 16 to 29, 4 of 32 inside the present; from the downbeat 16 to 16; 16: from inside the bar 17 to 31, 0 of 32 inside the present; from the downbeat 17 to 17; 17: from inside the bar 18 to 32, 0 of 36 inside the present; from the downbeat 18 to 18; 18: from inside the bar 19 to 32, 0 of 36 inside the present; from the downbeat 19 to 19; 19: from inside the bar 20 to 37, 0 of 40 inside the present; from the downbeat 20 to 20; 20: from inside the bar 21 to 35, 0 of 40 inside the present; from the downbeat 21 to 21; 21: from inside the bar 22 to 40, 0 of 44 inside the present; from the downbeat 22 to 22; 22: from inside the bar 23 to 39, 0 of 44 inside the present; from the downbeat 23 to 23; 23: from inside the bar 24 to 44, 0 of 48 inside the present; from the downbeat 23 to 24; 24: from inside the bar 23 to 39, 0 of 48 inside the present; from the downbeat 23 to 25; 25: from inside the bar 22 to 44, 0 of 52 inside the present; from the downbeat 23 to 25. In all, 61 of 590 entries are recognised within the 16 steps of a 3.5-second present. Rhythm and metre

A dancer who comes in late needs the downbeat marked

Every window for recognising an aksak metre so far started at its written downbeat. A dancer joining a dance already going has not heard the downbeat, and the arithmetic of that is blunt: a metre entered part-way is, onset for onset, each of its own rotations heard from their downbeats, and the rotations are metres too — 2+2+3 and 3+2+2 are counted differently. So on onsets, and with the long beats accented, no metre is ever told from its rotations. Only an accented downbeat tells them apart, and with it a listener who knows thirty metres recognises 54 per cent of them inside the present from a random entry, against 1 per cent without.

From partial 3 the room is the slower of the two. Decay rates in nepers a second for each partial of a note on 130.8 hertz, in a concert hall. The rising curve is the string's own loss, which grows as the partial number to the power 1. The flat-ish curve is the room's, from its reverberation time at that partial's frequency. A reverberant field is the source convolved with the room, so a partial's tail falls at the SLOWER of the two — the heavy line — and the room keeps returning energy the string has stopped making. From partial 3, at 392 hertz, the room is in charge: 6 of the note's 8 partials are held up by the room rather than let go by the string. Those are exactly the partials the string was losing fastest, which is why the room does not merely lengthen the note. Timbre and acoustics

The room is the slower of the two

A reverberant field is the source convolved with the room, so a partial's tail falls at the slower of the two rates rather than at their sum — and the room is slower for exactly the partials the string is losing fastest. Half a note's colour is gone in 0.163 seconds in no room at all, 0.313 in a concert hall and 1.441 in a stone church. The destination is identical in all three, because a room cannot hold a partial up above the fundamental it is also holding. What a hall takes away is the rate, and the rate was the whole of the identity cue.

Both qualities reach the same ceiling. The longest-lived spacing of a major triad and of a minor one, over four basses, taken over every arrangement of the three pitch classes within 2 octaves. They are the same number at every bass — 1.16 seconds over C2, 1.26 seconds over C3, 1.26 seconds over C4, 1.41 seconds over C5 — and at each bass 2 major and 2 minor spacings are tied at it. The faint line is the worst a minor spacing can do, which is 3.5 times shorter. So the asymmetry found earlier is a fact about the minor tenth rather than about the minor triad: a minor chord has a spacing that avoids it, and that spacing is its first inversion, where the minor third between two of its notes appears as a major sixth instead. Intervals and chords

A minor triad can be spaced to last

Two spacings of each triad, drawn side by side, say that opening lengthens a major chord and halves a minor one. Drawn over every arrangement of the three pitch classes within two octaves of a fixed bass, the asymmetry disappears: a minor triad reaches 1.26 seconds over C3 and so does a major one, with two spacings of each tied at the top. The short-lived chords were never the minor ones. They were the ones with a minor tenth on the outside, and a minor triad has three spacings that avoid it.

The ghost bass drops when the passage gets louder. The note the whole crowd of products names, as a multiple of the fundamental the interval implies, against how loudly the interval is played. a major third: 2.9999999999999996 times the fundamental below 70 decibels and 1 times above it, a drop of 19 semitones; a minor third: 4 times the fundamental below 62 decibels and 2 times above it, a drop of 12 semitones; a fourth: 2 times the fundamental below 72 decibels and 1 times above it, a drop of 12 semitones. Softly, only the cubic products clear their thresholds, and they are an exact series on (2p − q) times the fundamental with no gaps in it. Loudly, the difference tones fill in the low harmonics, no template on the higher note can explain them, and the fit falls. Nothing about the interval has changed; the listener is simply being given a different subset of the same harmonic series. Intervals and chords

The ghost bass drops a twelfth at a forte

Both crowds arrive at once and every member of both is a multiple of the same absent fundamental, so a listener is never given a choice between them — only a different subset of one harmonic series at every dynamic. Softly, the subset is an exact gapless series on three times the fundamental. Loudly, the difference tones fill in the low harmonics and no template on the higher note survives them. Between 62 and 72 decibels, depending on the interval, the note the crowd names falls by an octave or a twelfth, and the two qualities of third cross at different levels.

A count is least reliable at both ends and best in the middle. How precisely a listener knows the length of a section they are counting, against how many units long it is, at three kinds of timing judgement. A slip — one unit miscounted, at 2% a unit — accumulates as a random walk, so its relative cost FALLS as the section lengthens. A lapse — the count lost altogether, at 1% a unit — compounds, so the chance of still having the count falls geometrically and a long enough section is certain to lose it. A listener who has lost the count is back to timing, so the two failures mix into a floor. Against a Weber fraction of 7.5% the count is worth most at 21 units, where it is 1.7 times finer than timing, and falls back under a quarter better by 98 units; Against a Weber fraction of 15% the count is worth most at 10 units, where it is 2.4 times finer than timing, and falls back under a quarter better by 101 units; Against a Weber fraction of 38% the count is worth most at 4 units, where it is 3.7 times finer than timing, and falls back under a quarter better by 102 units. The length at which it stops being worth much is nearly the same in all three, because it is set by the lapse rate alone. Form and structure

A count is not an estimate

Both established routes to a proportion are estimates — a duration timed, blurred by a Weber fraction, and a duration stored, biased by what was new. A listener who has induced a hypermetre has a third, and it is exact until it fails. It fails two ways that pull opposite: a slip miscounts one unit and its relative cost falls as the section lengthens, while a lapse loses the count entirely and its chance compounds. The mixture has a floor at about four units, where counting is 3.7 times finer than timing, and it is worth almost nothing past a hundred.

Timing blurs a whole form evenly; counting sharpens it downward. A piece of 480 seconds divided 7 times, each level half the length of the one above, with how many proportions between 1 : 1 and 3 : 1 a listener can tell apart at each. Timed, the answer is 2.1 at the top and 5.2 at the bottom, a spread of 2.4 — because a timing judgement's Weber fraction is a step function of duration and almost every level of a piece falls in one step of it. Counted, in units of 2 seconds, the answer runs 2.2 to 10.3, a spread of 5.5. At no level does timing separate 3 : 2 from the golden section. Form and structure

A form is sharp at the bottom and vague at the top

A movement is divided into sections, each into phrases, each into bars, and every level is a ratio of two estimates. Timed, the hierarchy is almost uniformly blunt — 2.1 distinguishable proportions at the top and 5.2 at the bottom, because a Weber fraction is a step function of duration and six of a piece's seven levels fall in one step of it. Counted, the same hierarchy runs from 2.2 to 12.2 and sharpens monotonically downward. At no level of either does timing separate 3 : 2 from the golden section.

The golden section and an equal division are one judgement. Where a boundary falls in a piece, as a share of its length, with the band a listener cannot tell from the golden section shaded. A stretch of minutes is judged with a Weber fraction of about 38%, so one criterion's worth of ratio spread around 0.618 covers everything between 0.492 and 0.730 — a quarter of the piece wide, and containing the halfway point. 1 : 1 and 4 : 3 and 3 : 2 and golden section and 2 : 1 are inside it. A claim that a climax falls at the golden section rather than at the middle is, at this resolution, not a claim about anything a listener could hear. Form and structure

A golden section is a coin toss with six coins

An analysis that reports a climax at 0.618 of a piece has not tested one prediction; it has looked at a piece with several defensible boundaries and reported whichever landed nearest. The rate at which that happens under no hypothesis is one line of arithmetic, and the tolerance it needs is not a number chosen on the page — it is the blur a listener's own timing puts on the judgement. Over a stretch of minutes that blur covers everything from 0.492 to 0.730 of the piece, which contains the halfway point, and six candidate boundaries produce a hit eighty per cent of the time.

The breath is the looser ceiling nearly everywhere. How long a trained singer can hold a phrase on one breath, across a compass and at four dynamics, against the 8-second ceiling the psychological present puts on the same phrase. The flow through the folds rises with pitch and with loudness, so the breath ceiling falls both ways: at 60 decibels it runs 32.6 seconds at the bottom of the compass to 21.2 at the top; at 70 decibels it runs 23.1 seconds at the bottom of the compass to 15.0 at the top; at 80 decibels it runs 16.4 seconds at the bottom of the compass to 10.6 at the top; at 90 decibels it runs 11.6 seconds at the bottom of the compass to 7.5 at the top. The shaded line is the listener's ceiling and it does not move. The breath binds only where the two lines cross — 1 of the 40 cells drawn, all of them loud and high. So the constraint everybody names when asked why a phrase is the length it is, is almost never the constraint that decides it. Form and structure

The ceiling everybody names is the loose one

Ask why phrases are the length they are and the answer given is the breath. It is arithmetic — usable lung volume over the air a note costs per second — and it comes out between fifteen and twenty-three seconds at a comfortable dynamic and between seven and twelve at a loud one. The ceiling the present moment imposes, the two-to-eight seconds inside which a stretch is heard as one thing rather than as a series, is two to three times tighter at almost every note and dynamic. A singer in an adagio is not running out of breath at the phrase end. They are running out of present.

Three G strings, and they are not one pitch. Each instrument's four open strings, at the Pythagorean position its own chain of fifths puts them, with the uncertainty its own tuning leaves drawn as a band. The A is given and carries no error; every other string is reached from it one fifth at a time, and a fifth set by ear is set by nulling a beat whose rate falls with frequency — so the error accumulates down the chain and is worst at the bottom. violin: G3 -3.9 ± 1.77, D4 -2.0 ± 0.98, A4 0.0 ± 0.00, E5 2.0 ± 0.44; viola: C3 -5.9 ± 2.83, G3 -3.9 ± 1.77, D4 -2.0 ± 0.98, A4 0.0 ± 0.00; cello: C2 -5.9 ± 2.83, G2 -3.9 ± 1.77, D3 -2.0 ± 0.98, A3 0.0 ± 0.00. The three G strings share a pitch class and are expected to sit 2.5 cents apart; the two C strings 4.0. Pitch and tuning

The quartet settles at two pitches, not four

The quartet's open strings have been treated as five fixed pitches on one chain, and they are not: the violin, the viola and the cello each tuned a G string by ear and the three are expected to sit two and a half cents apart. Giving each player their own strings, with their own scatter, and pulling each toward only their own, changes the ensemble's settled pitch by a hundredth of a cent. What it does change is systematic rather than random: a violin has an E string and no C, the lower instruments have a C and no E, so the quartet splits by section by a tenth of a cent in every key.

A cycle already known, against a cycle just arrived at. How many bits of uncertainty about position a listener has, against how long the cycle takes, for a single timeline and for a layered colotomy — each drawn twice, once as a listener arriving and once as a listener who has been hearing it long enough to settle. At 1.6 seconds a cycle the timeline goes 0.94 bits arriving and 0.00 settled, and the colotomy 0.71 and 0.00; At 16 seconds a cycle the timeline goes 1.10 bits arriving and 0.08 settled, and the colotomy 0.93 and 0.23; At 60 seconds a cycle the timeline goes 2.47 bits arriving and 2.27 settled, and the colotomy 2.14 and 1.89. The gap between each pair is what the repetitions are worth, and it narrows as the cycle slows. The two designs are drawn at their own step counts rather than at equal strokes, so the levels here are not the earlier ones and the gaps are. Rhythm and metre

Repetition buys least where it is needed most

Every locating figure so far is a listener arriving — the uncertainty averaged over the first cycle heard. Cyclic music comes round dozens of times, and the same model already carries the answer for a listener who has settled: a floor of uncertainty that nothing had read. At two seconds a cycle the repetitions close the whole gap. At sixty they close eight per cent for a single timeline and twelve for a layered code. A slow cycle is worse on the first hearing and gains less from the second, and the two disadvantages compound.

The change reading follows the chords, not the bar. How far above the other candidates the true barline stands, in standard units, for the reading that scores how much the pitch-class content changes at each candidate — at three harmonic rhythms. At 2 chords a bar the margin is 0.12 and the reading finds the barline 12 per cent of the time; At 1 chord a bar the margin is 1.66 and the reading finds the barline 42 per cent of the time; At a chord every two bars the margin is 0.47 and the reading finds the barline 27 per cent of the time, against a chance rate of 13 per cent. The passages read earlier all changed chord once a bar, which is the middle column and the only one where the reading has anything. Two chords a bar puts a change at the half-bar as well and the reading cannot tell the two apart; a chord every two bars leaves half the barlines with no change at all and the margin halves exactly. Harmony and voice leading

The change reading follows the chords, not the bar

Every passage read until now changes chord exactly at the barline, which is the one harmonic rhythm at which 'the chords change here' and 'the bar starts here' are the same sentence. Pull them apart and the reading goes with the chords: at one chord a bar it stands 1.52 standard units above the other candidates and finds the barline half the time, at two chords a bar it stands 0.01 above them and is at chance, and at a chord every two bars its margin is exactly half — because half the barlines then carry no change at all.

Asked for the rate, it answers a multiple of it. The change reading asked its own question — what period do the chords change at — over passages built at three harmonic rhythms, with its standardised score for each candidate period. Given 2 chords a bar it recovers the rate 33 per cent of the time and answers too slow 65; Given 1 chord a bar it recovers the rate 58 per cent of the time and answers too slow 38; Given a chord every two bars it recovers the rate 93 per cent of the time and answers too slow 0. It never errs fast in the way it errs slow, and the reason is structural: a chord change every four slots also produces a change at every eighth slot, so a slower grid inherits a faster rate's evidence and a faster grid cannot inherit a slower one's. That ambiguity is why the reading looked like a barline detector in the first place — the bar is a multiple of every harmonic rhythm that fits inside it. Harmony and voice leading

Asked for the rate, it answers a multiple

A reading that follows the chord rate rather than the bar can be asked what the rate is, and the shape of its errors is the whole of why it looked like a barline detector. Given two chords a bar it returns the right period a third of the time and something slower two thirds; given a chord every two bars it is right nine times in ten. It errs slow and essentially never fast, because a change every four slots also falls on every eighth slot and a slower grid inherits a faster rate's evidence — which is the same asymmetry that makes a pitch detector report an octave too low.

The tempo moves it further than the touch does. One measured slendro, scored among random scales of its size under a free bar, 4 s, at five tempi and under each touch. Left to ring it runs from 29 at 0.15 seconds a note to 83 at 2.4 — a span of 54 percentile points, where the two touches differ by at most 17. So the scale is smoother than most of its size when the music is fast and rougher than most when it is slow, and how the bar is damped is the smaller decision. The two touches converge at the slow end because a bar that has died before its successor is sounding against nothing whatever the player does. Scales and modes

The tempo moves a scale further than the touch

A gamelan is played two ways on the same bars: a saron's are damped as the next is struck and a gendèr's ring over their resonators. That decision moves a slendro's standing among random scales of its size by up to seventeen percentile points, which is real. Over the tempo levels a piece actually moves through it moves by fifty-four — from the twenty-ninth percentile at a fast elaboration to the eighty-third at a slow one. The same five pitches on the same bars are a smoother-than-average scale and a rougher-than-average one, and which depends on how fast they are played.

A sharper cue is worth nothing to a reading that will follow it anywhere. How often the reading names the right scale degree, against how much it costs to change key, at a bass worth 1.5 on a real bass line. the bass is the root: 20 per cent at a key cost of 0.5 and 68 at 4; the bass is some chord tone: 47 per cent at a key cost of 0.5 and 59 at 4; the bar in octaves: 46 per cent at a key cost of 0.5 and 91 at 4; roots, and a line of roots: 49 per cent at a key cost of 0.5 and 91 at 4. Where a key change is cheap the rule that names the chord outright reads no better than the rule that names three — 46 against 47 per cent — because a reading that will move key for one bar's evidence follows a sharp cue wherever it points. The sharper cue's whole advantage appears only once the reading is reluctant enough to stay put, and by a key cost of 2.2 it is 31 points ahead. Scales and modes

A sharper cue is worth nothing to a reading that moves

The rule that names the chord outright reads 92 per cent of scale degrees right where the published bass cue reads 68 — at the key cost these readings have always been run at. Sweep that cost and the advantage is not a property of the cue. Where a change of key is cheap the sharp rule reads 46 per cent and the vaguest rule 47, because a reading that will move key on one bar's evidence follows a sharp cue wherever it points. The cue's whole value is borrowed from the model's reluctance to be moved.

A displaced map is displaced by the same amount everywhere. How far a listener's heard direction is displaced, in units of the smallest angular change they could detect at that azimuth, for four constant offsets added to every interaural delay. Each curve is flat. an offset of 5 microseconds is worth 0.33 just-noticeable steps at every azimuth; an offset of 10 microseconds is worth 0.67 just-noticeable steps at every azimuth, and past 88° hands the listener a delay their own head cannot produce; an offset of 20 microseconds is worth 1.33 just-noticeable steps at every azimuth, and past 86° hands the listener a delay their own head cannot produce; an offset of 40 microseconds is worth 2.67 just-noticeable steps at every azimuth, and past 82° hands the listener a delay their own head cannot produce. The reason is exact: differentiating Woodworth's curve gives a slope proportional to (1 + cos θ), so the angular displacement a fixed offset produces carries a factor of 1/(1 + cos θ) — and so does the smallest detectable angle, so the ratio has no azimuth in it. That is the opposite of a wrong head radius, whose displacement is zero on the median plane and grows toward the side. Perception and the listener

The error that moves straight ahead

The essay before this one found that a listener whose internal head is the wrong size makes no error at all on the median plane, and has to look hard to the side to catch it. Every head drawn here has its ears at equal radii, which makes the delay curve odd and every error a factor — and a factor cannot move a zero. Real heads are not symmetric. A constant offset of twenty microseconds displaces a listener's straight ahead by two and a quarter degrees, and it displaces every other direction by the same number of just-noticeable steps, exactly.

The two parameters are not one, and the reason is a ceiling. The plane of the two parameters these readings have been swept one at a time: how much weight the bass cue carries, against what a change of key costs. Every cell is how often the reading names the right scale degree, and the lines are the contours of equal share. along the 50 per cent contour the product of the two coordinates runs from 0.10 to 0.50; along the 60 per cent contour the product of the two coordinates runs from 0.43 to 1.84; along the 70 per cent contour the product of the two coordinates runs from 0.66 to 2.56; along the 80 per cent contour the product of the two coordinates runs from 1.85 to 4.91; along the 90 per cent contour the product of the two coordinates runs from 2.71 to 15.20. If the two multiplied cleanly those products would be constant and the contours would be hyperbolae. They are not: every contour turns upward and then vertical, because past a bass weight of about 3 more of the cue buys nothing at all and only reluctance is left to buy anything with. The key cost has an interior best, at 3 on this grid, where the reading names 93 per cent of degrees — so a reading that will not change key at all is worse than one that will, which no sweep of a single parameter had found. Scales and modes

The two parameters turn out to have a ceiling between them

The essay before this one asked whether the bass cue's weight and the cost of changing key are one quantity with two names, and said the test was a contour: if they multiply, the curves of equal degree share are hyperbolae. They are not. Along the ninety per cent contour the product of the two runs from 2.7 to 15.2, because past a bass weight of about one and a half the reading saturates and more cue buys nothing. And the sweep finds something no single-parameter sweep here could: the key cost has a best value, and a reading that will never change key is worse than one that will.

A louder final chord stands higher and still stands for a fraction of a second. How far a final chord stands above the listener's running impression at the instant it is released, against how long it lasts. The chord is a struck six-note tonic; "a step louder" is the hammer velocity doubled, which raises its loudness 7.12 phons above the tutti's. At the tutti's level, after 1.2 s of silence, the impression is 8.09 dB below the chord and the chord stands highest, 5.14 dB, at 54 ms; a step louder, straight out of the tutti, the impression is 7.12 dB below the chord and the chord stands highest, 4.51 dB, at 36 ms; a step louder, after 1.2 s of silence, the impression is 15.21 dB below the chord and the chord stands highest, 9.46 dB, at 38 ms. Every curve returns to zero by half a second: the impression climbs to whatever level the chord is played at, so a louder mark is not a stand that lasts but a deeper fall to climb out of, and a silence and a mark add as depths. Form and structure

A louder final chord is a deeper silence and a brighter sound

A final chord marked a step louder than the passage was supposed to stand above a listener's running impression for as long as it sounded, since the impression can climb no higher than the chord. It climbs exactly that high, and the stand closes in half a second as it always did. What a louder mark actually buys is depth — about seven phons, the same depth a second of silence buys — and a spectrum whose balance point sits most of a whole tone higher, which, unlike the stand, lasts for the whole chord.

Checkpoints sharpen the middle of a form and leave its top vague. A piece of 480 seconds divided 7 times, with how many proportions between 1 : 1 and 3 : 1 a listener tells apart at each level: timed, counted in 2-second units, and counted with a second count of 16-second phrases that can mend a lapse in the first. 480 s: 2.1 timed, 2.2 counted, 3.9 with 0 per cent of lapses shared and 2.5 with 50 per cent of lapses shared; 240 s: 2.1 timed, 2.5 counted, 7.1 with 0 per cent of lapses shared and 3.1 with 50 per cent of lapses shared; 120 s: 2.1 timed, 3.1 counted, 13.2 with 0 per cent of lapses shared and 4.0 with 50 per cent of lapses shared; 60 s: 2.1 timed, 4.1 counted, 19.4 with 0 per cent of lapses shared and 5.4 with 50 per cent of lapses shared; 30 s: 2.1 timed, 5.4 counted, 18.1 with 0 per cent of lapses shared and 7.2 with 50 per cent of lapses shared; 15 s: 5.2 timed, 12.2 counted, 12.2 with 0 per cent of lapses shared and 12.2 with 50 per cent of lapses shared; 7.5 s: 5.2 timed, 10.3 counted, 10.3 with 0 per cent of lapses shared and 10.3 with 50 per cent of lapses shared. With the two counts failing independently, the level of 60 seconds goes from 4.1 to 19.4, and the whole piece only from 2.2 to 3.9. Form and structure

Checkpoints sharpen the middle of a form, not its top

A listener who counts bars loses the count somewhere in a long section and is thrown back on timing the whole of it. A listener who also counts phrases can mend the lapse at the last phrase. If the two counts fail independently, the level a minute long goes from four distinguishable proportions to nineteen; the whole eight-minute piece goes only from two to four, because thirty phrases are long enough to lose a count as well. And if a fifth of lapses take both counts at once, three quarters of the gain is gone.

A harder strike buys beats only if the drop does not deepen with it. The beats a mistuned octave on A3 delivers on a string that decays in two stages, against how hard both notes are struck, with the aftersound's drop below the strike 20 dB at 80 dB and changing by a stated number of decibels for each decibel of strike. -0.5 dB per dB: 70 dB 2.5, 75 dB 5.7, 80 dB 8.9, 85 dB 11.9, 90 dB 12.2, 95 dB 11.6, 100 dB 10.9; 0 dB per dB: 70 dB 4.7, 75 dB 6.9, 80 dB 8.9, 85 dB 11.0, 90 dB 12.3, 95 dB 12.2, 100 dB 11.8; +0.5 dB per dB: 70 dB 7.2, 75 dB 8.1, 80 dB 8.9, 85 dB 9.8, 90 dB 10.7, 95 dB 11.2, 100 dB 11.9; +1 dB per dB: 70 dB 9.6, 75 dB 9.2, 80 dB 8.9, 85 dB 8.4, 90 dB 8.1, 95 dB 7.8, 100 dB 7.5. Where the drop stays put, a harder strike keeps buying beats to about 90 dB; where it deepens by a decibel per decibel, every strike past 70 dB costs beats. Intervals and chords

A firm touch buys beats until the aftersound sinks with it

A piano string decays twice, and a mistuned octave's countable beats were found to rise with how hard the note is struck — on the assumption that the aftersound always starts twenty decibels below the strike. It need not. The moment the count shuts is set by the level the aftersound starts at and nothing else, so a harder strike buys beats only if its aftersound does not sink with it. If the drop deepens by more than 0.85 decibels for each decibel of strike, the firm touch costs beats instead.

The arch belongs to hearing, and the spacing only moves it. The share of a close major triad's twenty-four components that stand above what the rest of the chord masks, at 70 dB, with the root from C1 to C7, for three spectra given the same amplitude law and different frequencies: the harmonic series, a founder's bell, and a stiff string with B = 0.01. harmonic series: 0.04 at C1, peaking at 0.79 on E3, 0.42 at C7; a founder's bell: 0.04 at C1, peaking at 0.75 on C4, 0.38 at C7; a stiff string: 0.04 at C1, peaking at 0.71 on E3, 0.46 at C7. Only one of the three is a harmonic series, and all three rise out of the bass, peak in the middle of the compass and fall in the treble. Perception and the listener

The arch belongs to hearing, not to the series

A chord delivers most of its partials in the middle of the compass and loses them in the bass and the treble, and every spectrum that showed that arch was built on whole multiples of a fundamental. Give the same amplitudes to a bell's eight modes and to a stiff string's stretched partials and the arch is still there, peaking within a major third of where the harmonic series peaks. What the spacing changes is the detail: a bell crowds its tierce and quint into a quarter of a critical band in the bass and loses them, and a stiff string's stretch buys the bass back.

A listener who hears only the landmarks recognises the metre sooner. The share of metres recognised within the 16 steps of a 3.5-second present after coming in at a random step, against how many metres the listener knows, for five listeners: every onset with nothing marked, every onset with the long beats accented, only the downbeats and long beats with nothing marked, every onset with the downbeat accented, and only the downbeats and long beats with the downbeat marked. Every onset, nothing marked: 5 known, 49% within the present, 3% never; 10 known, 19% within the present, 8% never; 30 known, 1% within the present, 17% never; 100 known, 0% within the present, 34% never. Every onset, long beats accented: 5 known, 68% within the present, 3% never; 10 known, 39% within the present, 8% never; 30 known, 6% within the present, 17% never; 100 known, 0% within the present, 34% never. Landmarks only, nothing marked: 5 known, 69% within the present, 0% never; 10 known, 61% within the present, 3% never; 30 known, 26% within the present, 3% never; 100 known, 9% within the present, 9% never. Every onset, downbeat accented: 5 known, 88% within the present, 0% never; 10 known, 74% within the present, 0% never; 30 known, 54% within the present, 0% never; 100 known, 14% within the present, 0% never. Landmarks only, downbeat marked: 5 known, 91% within the present, 0% never; 10 known, 82% within the present, 0% never; 30 known, 56% within the present, 0% never; 100 known, 23% within the present, 0% never. Rhythm and metre

A late dancer needs the landmarks, not the rhythm

A listener who joins an additive-metre dance part-way recognises it far more reliably with the downbeat accented. Strip the stream down to its landmarks — the onsets that begin a bar or a long beat, with every other onset removed — and the listener does as well or better: knowing a hundred metres, 23 per cent are recognised within the present against 14 with every onset. Unmarked, the landmarks still beat every onset with the long beats accented. Neither half does it alone; what identifies a metre from a late entry is where its long beats sit relative to its bar.

A timed expectation would erase a slow cycle's cost, and a listener cannot time a slow cycle that well. Bits of position a listener with a 3.5-second memory is still missing over the first cycle of son clave, against how long the cycle takes, for a newcomer with no expectation, a listener timing the cycle with the Weber fraction a duration that long is judged with, and a listener timing it to ten per cent. a newcomer, no expectation: 2 s 0.94, 8 s 0.97, 24 s 1.41, 40 s 2.02, 60 s 2.47; timing as well as listeners do: 2 s 0.68 (w 0.150), 8 s 0.69 (w 0.150), 24 s 0.85 (w 0.150), 40 s 1.90 (w 0.375), 60 s 2.38 (w 0.375); timing the cycle to ten per cent: 2 s 0.47, 8 s 0.48, 24 s 0.55, 40 s 0.73, 60 s 1.02. At ten per cent even a sixty-second cycle is placed about as well as a newcomer places a two-second one. At the precision a listener actually has for durations of half a minute or more, the expectation is worth a tenth of a bit. Rhythm and metre

An expectation cannot rescue a cycle too slow to time

A listener who knows a piece arrives with an expectation of where in the cycle they are, and the size of that expectation was the number the last essay said nobody had. It can be given one: a listener who has been timing the cycle carries a spread of their Weber fraction times the cycle, which is the same number of steps at any tempo. Timed to ten per cent, a forty-second cycle would be placed better than a newcomer places a two-second one. But forty seconds is judged in the band where the Weber fraction is nearer forty per cent, and there the expectation is worth a tenth of a bit.

Counted over what arrives, the balanced bass is not the roughest register. The mean roughness of the I – vi – IV – V – I arrivals with each chord played as loud as the written register's, relative to the written register, counted over every partial and over the partials that stand above what the rest of the chord masks. Every partial: 70 −2 octaves, 7.48 −1 octave, 1.00 as written, 0.20 +1 octave. Delivered partials only: 6e-9 −2 octaves, 2.83 −1 octave, 1.00 as written, 0.19 +1 octave. Over every partial the lowest register is 343 times rougher than the highest; over what arrives it is the smoothest of the four, and the roughest is −1 octave, 2.8 times the written register. Perception and the listener

A bass chord low enough to balance has already hidden its tenor

Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.

All themes