Theme

The ear is not a microphone

Loudness is not amplitude, pitch is not frequency, and a note whose fundamental is missing still has one. Perception is a model, not a measurement.
Two tones in the ratio 3 to 2. Two sine tones whose frequencies are in the ratio 3 to 2, and their sum. The combined pattern repeats every time both waves return to the start together, which for a simple ratio is soon — here after 2 cycles of the lower tone and 3 of the upper. Intervals and chords

Two notes and a ratio, which is the whole of consonance

Sound two tones together and the pair either settles or does not. What decides it is the ratio of their frequencies, and the rule is that simpler ratios settle — which is two and a half thousand years old and still not quite an explanation.

220 Hz against 223 Hz. Two tones 3 hertz apart, added. The rapid oscillation is their average; the slow swelling is their difference, heard as 3 beats a second and used by every tuner who has ever worked by ear. Intervals and chords

Beats are arithmetic that anybody can hear

Two tones a few hertz apart swell and fade at exactly their difference. It is the most direct evidence available that the ear does sums on what reaches it, and it is how every instrument in the world gets tuned.

A note with its first partial removed. The spectrum of a 220 Hz tone with the lowest partial deleted, and the wave that remains. The wave still repeats 220 times a second, because the repeat rate of a sum of harmonics is fixed by the spacing between them and the spacing has not changed. The pitch heard is the one that is no longer in the sound. Intervals and chords

The note that is not there

A telephone reproduces nothing below about three hundred hertz, and a bass line down at eighty comes through it perfectly. The pitch that is heard is not a frequency present in the sound, and one nineteenth-century experiment settles what it is instead.

The critical band, measured in semitones. The width of the ear's frequency-analysis band at each pitch, converted from hertz into semitones. Two intervals drawn as horizontal lines cross the curves: below the crossing the interval fits inside one band and its notes are not resolved from each other, and above it they are. Intervals and chords

A third is rougher in the bass

Consonance is usually presented as a property an interval has. It is not. The same major third is muddy two octaves below middle C and clean two octaves above it, the ratio never changed, and the frequency where it stops being muddy can be solved for.

The first eight partials of a string. A string vibrating in one, two, three and more equal parts, with the frequency ratio and the nearest named note beside each. The seventh partial is a third of a semitone flat of anything on a keyboard, which is a fact about strings rather than about tuning. Timbre and acoustics

A string does everything at once

A plucked string does not vibrate at one frequency. It vibrates at all the whole-number multiples of one frequency simultaneously, and nearly everything in music theory is downstream of that fact.

Four spectra of the same note. The amplitude of each partial for 4 timbres at the same pitch — pure, string, clarinet, bell. These are the exact lists the sound buttons here synthesise from, so the picture and the sound are the same data. Timbre and acoustics

The ear hears the list, not the shape

Two sounds with the same partials and different phases have completely different waveforms and sound identical. What the ear extracts is a list of frequencies and strengths, and everything else is discarded.

Three envelopes. How loudness changes over the life of a note, for plucked, bowed and struck. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum. Timbre and acoustics

The shape of a note, which is most of what an instrument is

Cut the first fifty milliseconds off a recorded piano and listeners stop calling it a piano. The attack carries more identity than the steady tone it leads into, and it is the part every spectrum plot leaves out.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval. Intervals and chords

The ear sorts into boxes, and the boxes are the theory

Slide one note slowly upward against another and the interval between them changes continuously. What a listener reports does not. It stays a minor third, stays a minor third, and then in the space of about twenty cents becomes a major third — and nothing in the sound corresponds to the moment of the change.

Which partials two notes have in common. The first 12 partials of a note and of the note an interval above it, on a logarithmic frequency axis, with every partial of the upper tone that lands on one of the lower marked. octave: 12 of 12, fifth: 6 of 12, fourth: 4 of 12, major third: 3 of 12, minor third: 2 of 12, minor second: 0 of 12. Sharing partials is what fusion is: two voices whose spectra mostly coincide stop being heard as two voices, which is a measurable property of the interval rather than a matter of taste. Harmony and voice leading

Two voices that stop being two

The ban on parallel fifths is the most famous rule in Western music and it is usually taught as taste. It is not taste. At an octave the upper voice contributes no frequency the lower one did not already have, at a fifth it contributes half of them, and the number can be counted — which turns a prohibition into a measurement.

3 against 2, from a rhythm to an interval. The same 3:2 pattern at 6 rates, an event rate from 1.5 to 300 per second, on a logarithmic axis. Below about 20 events a second the two streams are counted; above about 40 they are heard as two pitches a 3:2 apart, which for 3:2 is 702 cents. Nothing in the pattern changes across that boundary. The slowest rate drawn is a tempo of 90 events a minute; the fastest is a pitch of 300 Hz against 450 Hz. Rhythm and metre

A rhythm fast enough to be a chord

Three against two is a polyrhythm. Speed the same pattern up until the events arrive faster than about twenty a second and it is a perfect fifth. The ratio never changed, the figure never changed, and the only thing that moved is the rate — which makes rhythm and pitch one continuum with a perceptual boundary across the middle of it.

One pattern, four metres. The same 12-step onset pattern read under 4 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are 4/4 -4, 3/4 12, 6/8 12, 12/8 as 3+3+3+3 12, so 3/4 and 6/8 and 12/8 as 3+3+3+3 tie and the model does not choose. Nothing about the sound differs between these readings; the bar line is supplied by the listener. Rhythm and metre

The beat is inferred, and sometimes wrongly

A bar line is not in the sound. One sequence of onsets supports a reading in four, in three and in six, and which one a listener settles on comes from a small set of preference rules rather than from anything in the signal — which is why the same recording can be heard two ways and why nobody can be talked out of either.

Three attacks, the first 50 ms. How loudness changes over the life of a note, for plucked, bowed and struck, drawn over the first 50 milliseconds. By the right-hand edge the plucked note is at 91%, the bowed note is at 36%, the struck note is at 97% — attack times of 4 ms, 140 ms, 2 ms, a spread of 70 to one, and the part a listener uses to tell them apart. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum. Timbre and acoustics

The first fifty milliseconds

A spectrum is supposed to be what makes a trumpet a trumpet. Cut the first fifty milliseconds off a recorded note and listeners stop being able to name the instrument — while the spectrum they are hearing is unchanged. Identity is in the part of the sound that ends before the note has properly started.

Every point on one curve sounds equally loud. The equal-loudness contours of ISO 226:2003, evaluated from the standard's own parameters. The lowest curve is the threshold of hearing. Because the curves are not parallel — they crowd together in the bass and spread apart in the middle — the same change in decibels is a different change in loudness at every frequency, and a spectrum that was balanced at one level is not balanced at another. Perception and the listener

The quietest thing audible, and why the volume knob is a tone control

A decibel is a fact about air. A phon is a fact about a listener, and the two do not line up — the map between them bends with frequency, and it bends differently at every level. One consequence is that turning a piece of music down does not turn all of it down equally, and the amount by which it does not is a number.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward. Perception and the listener

One sound hides another, and it hides upward

A tone can be made completely inaudible by a second tone that is nowhere near it in frequency, and the region it disappears into is lopsided. Masking spreads up the spectrum and barely down it, and the reach grows with the masker's level — so which line in an arrangement vanishes is a prediction, not a matter of taste.

The threshold before, during and after a burst. The level a brief probe needs in order to be heard, plotted against when it happens relative to a 70 dB burst that occupies the shaded band. To the right is forward masking, which decays over about 200 milliseconds. To the left is backward masking: the threshold is raised for a probe that has already finished before the masker begins. Perception and the listener

A sound hides what came before it

Masking does not stop when the masker does. A loud sound raises the threshold for about a fifth of a second after it ends, which is unremarkable, and for several milliseconds before it begins, which is not. The auditory present is a window rather than an instant, and inside the window the order of events is not the order they arrived.

Partial 3, mistuned by 3%. A 10-partial tone on 220 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 3.0 per cent, which is 19.8 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune. Timbre and acoustics

What makes two partials one note

A note is a stack of ten or twenty simultaneous tones and is heard as one thing. The obvious explanation is that they are whole-number multiples of a fundamental — and the obvious explanation is not sufficient. Mistune one partial by three per cent and it leaves the note; give a perfectly harmonic partial a thirty-millisecond head start and it leaves too. Shared behaviour beats arithmetic.

1000 and 1200 hertz, and what the ear adds. Two tones presented to a listener, and the frequencies a nonlinear ear generates from them. Nothing in the air is at any of the marked positions: they are products of the pair, at f₂ − f₁ = 200 Hz, 2f₁ − f₂ = 800 Hz, 3f₁ − 2f₂ = 600 Hz, 2f₂ − f₁ = 1400 Hz. The cubic difference tone sits just below the lower primary and is audible at modest levels; the quadratic one is far below both and needs a loud pair. Intervals and chords

The ear makes its own sound, and it is not the missing fundamental

Play two loud tones and a third pitch appears that is in neither of them. The ear is not a passive analyser: it is nonlinear, it generates frequencies of its own, and it emits sound back out of the ear canal. None of which explains the missing fundamental — the products land in the wrong place, and finding out where they land is the experiment that made the residue theory necessary.

Cycles of roughness completed in 125 milliseconds. For each interval and each register, how many cycles of its own beating fit inside a note 125 milliseconds long. Cells at or above 4 cycles are the ones a listener has time to hear as rough; the pattern runs the opposite way from roughness itself, which is largest in the bass. A short dissonance low down is the case where the two disagree. Intervals and chords

A dissonance has to last

Roughness is a fluctuation, and a fluctuation needs cycles. A minor second at the bottom of a cello fluctuates thirty-three times a second, so a semiquaver holds four of them and a demisemiquaver holds two — which turns the counterpoint rule that a dissonance may pass if it is short into a number, and puts that number at about seventy milliseconds through most of the range, once the pairs beating too slowly to be roughness at all are taken out of the average.

Four voices, placed by the arithmetic. The rules in force are: no parallel octaves; no parallel fifths; no voice crossing; no gap over an octave above the tenor; the leading note is not doubled; the leading note resolves, outer voices; no augmented melodic interval. The four parts are drawn lowest to highest — bass, tenor, alto, soprano. I – vi – ii – V – I in C major, realised in four voices by the cheapest set of voicings obeying 7 rules, at 16 semitones of motion in all. Every gap between adjacent voices narrower than the fission boundary of 5.2 semitones is marked, and below that boundary two parts cannot be heard as two however hard a listener tries. Perception and the listener

A voice is a stream, and the ear decides which

Seven earlier essays have assigned voices to notes. Whether a listener follows the assignment is a separate question with laboratory numbers attached, and the numbers are unkind to it: two parts closer than about five semitones cannot be heard as two at any speed, a third of the gaps in the cheapest four-part writing are inside that limit, and in a third of chord changes the ear's own rule for continuing a line does not recover the parts as written.

Which harmonics of a 200 Hz note arrive one to a filter. One row per harmonic of a 200.0 Hz tone. The bar is the ear's analysis band at that harmonic on the equivalent rectangular bandwidth model, and the two small marks either side are the neighbouring harmonics. A harmonic is counted as resolved when the spacing to its neighbours, 200.0 Hz, exceeds that bandwidth, and 8 of 12 are. The third to fifth harmonics are picked out because published measurements put the pitch's dominance region there; nothing in this drawing derives that. Intervals and chords

Which harmonics carry the pitch

A missing fundamental is inferred from a pattern, and a pattern has to be legible before it can be matched. Counting how many harmonics of a note land in separate auditory filters prices the inference — and the answer at the bottom of a bass guitar's range is none of them.

A 200-a-second click train, correlated with itself. The autocorrelation of a click train at 200 a second, smoothed by the ring of an auditory filter centred at 4000 Hz — an equivalent rectangular bandwidth of 456 Hz, so a ring of 2.2 ms. The regular train peaks at 5.0 ms, one period. With each click displaced by a standard deviation of 20 per cent of the period — 1.00 ms — the peak's contrast against the surrounding lags falls from 0.41 to 0.12. The average rate and the long-term spectrum are unchanged by the jitter; only the timing is. Perception and the listener

A pitch with nothing to match

Filter a click train into a band where no partial is separable from its neighbours and it still has a pitch at its repetition rate. Displace each click by a fraction of a millisecond, leaving the average rate and the long-term spectrum exactly where they were, and the pitch goes. The mechanism is reading the timing — which bounds the account endorsed here from the start.

Roughness and loudness of one interval against level. A minor third on C4 evaluated at every level from 20 to 100 dB SPL per note, with both quantities drawn relative to their own value at 60 dB. Roughness is the Plomp–Levelt sum over the partials that are above ISO 226's threshold at that level; loudness is the same partials in sones. Between 50 and 95 dB the interval grows 3.2 × 10⁴ times rougher and 24.9 times louder, so roughness grows like loudness raised to the power 3.2. Intervals and chords

The same chord is harsher when it is louder

Every roughness number so far was computed at a level nobody stated. Roughness is the product of two partial amplitudes, so it is quadratic in pressure, while loudness is compressive — which makes a minor third at middle C thirty-two thousand times rougher at fortissimo than at pianissimo and only twenty-five times louder. A chord has no single consonance to report.

The swing ratio, against the categories it passes through. The same swing curve read against the boundaries between duration categories rather than against notated values. A category's centre is a simple ratio — 1:1, 2:1, 3:1 — and the boundary between two of them is the midpoint, which is arithmetic. The curve crosses 2 of them: out of 2:1 and into 1:1 at 240 beats a minute, out of 3:1 and into 2:1 at 171 beats a minute. So the same notated figure is, by the categorical criterion, a different rhythm at each end of an ordinary tempo range, and the notation says triplet feel throughout. Perception and the listener

How late is a different note

A deviation of thirty milliseconds is expression and a deviation of two hundred is a wrong note, so there is an edge. The edges in time are arithmetic — the midpoints between the simple ratios — and the swing ratio crosses two of them as the tempo rises, at 171 and at 240 beats a minute, while the notation says triplet feel throughout.

One decay, two verdicts, and the line is the listener's. The early-to-late energy ratio against reverberation time, for 50 millisecond and 80 millisecond windows. Nothing about the room differs between the curves; only where the line is drawn across its decay. Zero comes at 1.00 seconds for the 50 ms window and 1.59 seconds for the 80 ms window — which are, to two figures, the published rules of thumb for a room for speech and a room for music. The design targets were not put in; they came out. Timbre and acoustics

The first eighty milliseconds are a different room

Draw a line across a room's decay and the energy on either side is two opposite verdicts about one building — clarity before it, reverberation after. The line is a property of the ear, not of the room, and putting it at 50 milliseconds and at 80 makes the two published design targets fall out — a room for speech at one second, a room for music at 1.6.

Ode to Joy as a path. Ode to Joy plotted as 30 notes against the 8 scale degrees it uses, one column per note. Its largest melodic interval is 2 semitones and it spans 7; the mean absolute step is 1.24 semitones. Beethoven, Ninth Symphony, finale, 1824 — the theme as first stated, eight bars. Scales and modes

A melody is a walk, not a set

Nine essays here are about which seven of the twelve a scale takes, and every one of them describes a set. A tune is not a set; it is a path across one, and the path is nearly all small steps. That is not a matter of taste. Above about eight notes a second the ear stops being able to hold a large interval and a small one in the same line, and at sixteen the choice disappears altogether — so a fast passage is scalar because a fast passage that leaps is two pieces of music.

The vowel in "hod", sung at 110 Hz. The partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung. Timbre and acoustics

The sound a listener knows best

A voice is recognisable across every vowel it says, across two octaves of pitch, down a bad telephone line and in a whisper where there is no pitch at all. Nothing that survives all of that can be a frequency. What survives is a ratio: the resonances of a vocal tract are set by its length, so a shorter tract multiplies every formant by the same factor, and identity is a scale on the spectral envelope rather than a position within it. Between an adult man and a child the whole pattern moves by a fifth, and the vowel does not change at all.

The fluctuation stays; the rate goes. A unison of n voices with a spread of 15 cents, averaged over 5 draws. The depth of the amplitude fluctuation does not fall as voices are added — a choir is no steadier than a duet — but the fraction of that fluctuation in any single modulation component falls from 77 per cent at two voices to 24 at 32. Two voices make one beat and it can be counted; 16 make 120 and none of them is a rate. That is why a choir cannot be tuned by nulling anything. Timbre and acoustics

What a choir does that a soloist cannot

Two singers on one note produce one beat and it can be counted. Sixteen produce a hundred and twenty at once, and the amplitude still fluctuates by as much as it did — a choir is no steadier than a duet. What has gone is not the fluctuation but its rate: the modulation energy that sat in a single line at two voices is spread across a band at sixteen, with no line in it. That is the choral sound, and it is also why the just-intonation drift this site measured describes only ensembles that hold their pitch still.

The chain of fifths in quarter-comma meantone. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and one of them — G♯ to D♯, the 12th link, where the chain is forced to close — is the wolf, at 35.7 cents. Intervals and chords

The same distance, under two names

Four hundred cents is a major third or a diminished fourth, and on a keyboard nothing in the sound distinguishes them. An earlier essay was about the boundary between two categories; this is about two categories at one acoustic value, and the surprise is where the ambiguity comes from. In quarter-comma meantone a major third is 386 cents and a diminished fourth is 427 — two names, two pitches, forty-one cents apart. Equal temperament collapsed them, and what a listener now supplies from context used to be in the sound.

A dynamic mark is an instruction about the spectrum. Six dynamic markings, given a hammer velocity each in a stated sequence of factors of two, with what the string then does. The level rises 35.1 decibels from pp to ff, which is the part everybody means. The contact time falls from 2.26 to 0.95 milliseconds, so the first null of the hammer's own pulse moves from partial 2.5 to partial 6.0 and the spectral centroid rises by 56 per cent. The partials between those two nulls are not quieter at pp; they are not there. Timbre and acoustics

The mark that is not a level

There are six of them, they carry no units, and a performer has to turn one into a number before it means anything. What they instruct is not loudness. On a struck string a harder blow shortens the hammer's contact from 2.26 milliseconds to 0.95, which moves the first null of its own pulse from the third partial to the sixth: the partials between those are not quieter at pianissimo, they are gone. A fortissimo is a different sound, and the page has one word for both things it changes.

The pitch moves and the repetition rate does not. Three partials around harmonic 10 of 200 hertz, shifted together by up to 200 hertz, with three curves. The flat line is the envelope repetition rate, which the shift cannot move at all. The rising line is the shift divided by the harmonic number — 200 hertz becoming 220.0 — which is what a harmonic template predicts and is what listeners report. The third curve is the next-best template, which overtakes the first partway along: the pitch is ambiguous, and it drops back rather than rising indefinitely. Perception and the listener

The pitch that moves the wrong distance

Take three partials two hundred hertz apart and move every one of them up by forty. The spacing has not changed, so anything reading the pitch off how often the waveform repeats must give the same answer as before. The pitch moves to 204 — the shift divided by the harmonic number — which is what a harmonic template predicts and what listeners report. Push the shift to a hundred and a second reading overtakes the first, so there are two pitches and neither is the spacing. This is the measurement that closes the question, and it closes it by ruling out one mechanism rather than by choosing between the two that are left.

The series has three tops. Where the harmonic series stops, asked three ways, at four fundamentals. Consecutive partials stop being separately resolvable around partial 8 — that one depends on the fundamental, since a critical band is a fixed width in hertz and the spacing is not. They stop being a semitone apart at partial 17 at every fundamental, because the ratio (n+1)/n does not know what n is measured in. And they stop being distinguishable in pitch at all between partials 34 and 140. Every claim about how far up the series something happens is a claim about which of these three was meant. Intervals and chords

The series has three tops

How far up the harmonic series can an ear go? The question has three answers and they are an order of magnitude apart. Consecutive partials stop being separately resolvable somewhere around the eighth, and where exactly depends on the fundamental. They stop being a semitone apart at the seventeenth, at every fundamental, because the ratio does not know what it is measured in. And they stop being distinguishable in pitch at all between the thirty-fourth and the hundred and fortieth. Every claim about where the series runs out is a claim about which of the three was meant.

How long a note has to be before its pitch is worth arguing about. The smallest audible frequency difference at 440 Hz, against how long the note lasts. The flat line is the steady-tone difference limen of 4.0 cents that every tuning argument on this site rests on. The falling line is the bound a finite duration imposes on its own frequency, 1/2T in cents, which no listener can beat. They cross at 486 milliseconds: below that the note is the limit and above it the listener is. A tenth of a second gives 19.6 cents and a quarter gives 7.9, against the commas drawn across the figure. Perception and the listener

How long a note has to be

Every difference limen quoted so far is for a tone that lasts as long as the listener needs, and no note in music does. A tone of duration T occupies a band about 1/2T wide whatever the ear does with it, so at 440 hertz the quoted five-cent limen is the right number only for notes longer than 486 milliseconds. A tenth of a second gives 19.6 cents, which does not clear the syntonic comma. Most of the tuning arguments in this collection are about a quantity that only exists in long notes, and the essays that made them said so about the listener and not about the note.

500 Hz in one ear, 504 in the other. Two tones 4 hertz apart, one to each ear. They never meet in the air, so neither eardrum sees any modulation at all and there is no acoustic beat to hear. What changes is the phase between the ears, which advances a whole cycle every 250 milliseconds — and the direction that phase implies sweeps with it, drawn here as azimuth against time. The sweep is clipped at the edges, because the implied delay leaves the range a head can produce. A head 17.5 cm across gives at most 656 microseconds, so the phase stops naming a direction above 762 Hz. Perception and the listener

The beat that is not in the air

Every sound this site synthesises reaches both ears identically, and that is the assumption none of its figures ever varied. Put 500 hertz in one ear and 504 in the other and nothing sums anywhere: each eardrum sees a steady sinusoid with no modulation on it at all. A listener still hears a four-per-second beat, which means the arithmetic is being done behind the ears rather than in the room. And it stops working above about a kilohertz — not where phase locking gives out at five, but where a head 17.5 centimetres across stops being able to name a direction, which is 762 hertz.

3 tones, one power, and the interval between them. 3 tones of fixed total power, spread symmetrically about 440 hertz, drawn against the interval between neighbours. Piled on one pitch they are one sound of that power; separated by more than a critical band — 4.5 semitones here — they are 3 sounds whose loudnesses add, and the same power reaches 2.08 times the loudness at 5 semitones. The two lines are two models of the same rule and they disagree about how abrupt the change is, not about where it goes. Perception and the listener

A chord is not as loud as its notes

The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.

One envelope, and the three places a listener might be said to hear it. The amplitude envelope of a note with a 90 millisecond exponential attack, with the three criteria the literature offers drawn across it. The heard moment is 8.0 ms at the detection criterion, 28 ms at the perceptual-onset criterion and 94 ms at the perceptual-attack criterion. The physical onset is at zero on this axis and no criterion puts the heard moment there. The buttons play this attack against a two-millisecond one, started at the same instant. Rhythm and metre

A note is heard after it starts

Every rhythm essay until now has treated a note's onset as the moment it happens. It is not: the instant a listener aligns a note with a beat is later than its physical start by an amount the note's own attack decides, and for a sung or bowed note that amount is about thirty milliseconds — the size of the whole quantity six essays on microtiming set out to measure.

How much earlier an accent is heard, by mechanism. An accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller. Perception and the listener

Playing louder is playing earlier

An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.

Crescendo, and what the impression does. A crescendo of 20 dB over 8 seconds, drawn as three loudnesses in sones. The pale line is what is physically sounding, the middle line is the short-term loudness of the moment and the heavy line is the long-term loudness, which is the passage's loudness as a listener would report it. The gap between the last two is the whole of the effect: at its widest the moment is 1.02 times the running impression, and the impression takes 2 seconds to come down against 99 milliseconds to go up. Form and structure

Loud is relative, and it comes down slowly

The account of loudness had a model of a moment and the account of closure asked it for a model of a form. The published one exists and its content is a pair of numbers that are not the same: a listener's running impression of how loud the music is rises to meet a step in a fifth of a second and takes seven seconds to come back down. A twenty-decibel crescendo spread over eight seconds therefore buys almost no contrast at all, and the same twenty decibels taken as a step buys a factor of two.

One pair, 12 beat rates. Two notes at 220 hertz, 15 cents apart, drawn as the beat rate between each pair of partials. The fundamentals beat at 1.91 hertz and the k-th partials at k times that, so the series of rates is a straight line and it crosses 15 hertz at the 8th partial. Below the line the fluctuation is counted; above it the same physical fluctuation is heard as roughness. 1 per cent of this spectrum's energy is on the roughness side. The bars are drawn at each partial's own amplitude, so a spectrum with little energy high up crosses the line where nothing is listening. Pitch and tuning

Every partial beats at its own rate

Five earlier essays have drawn one beat rate per figure, and every one of them is the rate between two fundamentals. Two real notes beat between all of their partials at once, the k-th pair beats k times as fast, and somewhere up the spectrum the rate passes the point at which a beat stops being a beat — so a chorused note is a beat at the bottom of itself and a roughness at the top, simultaneously, with a crossover partial that is arithmetic.

How far equal temperament puts each interval's coincidence out. For each interval, the pair of partials it brings together and how many cents equal temperament mistunes that coincidence by. A fifth's third-against-second is out by 2 cents and a major third's fifth-against-fourth by 14 — so the same temperament that is inaudible on a fifth produces, at 220 hertz, a beat of 8.7 per second between two sections singing a third, with every singer in both of them perfectly in tune. Timbre and acoustics

A section against another section

The choir has been treated as a unison, and no choir sings only unisons. Two sections an interval apart beat between partials rather than between fundamentals — the third brings the fifth partial of one against the fourth of the other — and equal temperament puts that coincidence fourteen cents out. So two sections singing a tempered third beat at nearly nine per second with every singer in both of them perfectly in tune, and the same temperament is inaudible on a fifth.

Where a room stops sending the two ears the same sound. The correlation between the two ears' signals against frequency, for a seat 15 metres from the source in a 15,000 cubic metre room with a 2 second reverberation time. The pale curve is the diffuse field alone — sin(kd)/(kd) for an ear separation of 17.5 centimetres, which first crosses zero at 980 hertz. The heavy curve adds the direct sound, which is coherent and lifts the whole thing by an amount the direct-to-reverberant ratio sets. At 125 hertz the coherence is 0.98 and at 1000 it is 0.16. Perception and the listener

Where the two ears stop agreeing

A room sends both ears versions of the same sound, alike at low frequencies and increasingly unlike at high ones. Where they stop resembling each other is 980 hertz, and it is set by the 17.5 centimetres between the ears rather than by anything about the room — which is within a quarter of a frequency found earlier for a completely different reason. One minus that correlation is spaciousness, and it is computable from a room's own reverberation.

The beat rate between two sections, second by second. The 5th partial of the lower section against the 4th of the upper, over 256 pairs of voices each sweeping 100 cents 6 times a second with its own phase. The band is the tenth to the ninetieth percentile of the instantaneous rate and the line is the median. With no vibrato the whole thing would be one flat line at 8.7 hertz, which is what was computed earlier. With it, the pair is inside the beating band 20 per cent of the time and above it for the rest — so what a listener gets is neither a beat nor a roughness but an alternation between them at the vibrato rate. Intervals and chords

Sixteen sweeps against sixteen

Every intonation figure about the voice treats a singer as a frequency. A singer is a frequency being swept a hundred cents wide six times a second, and two sections singing an interval are two hundred and fifty-six pairs of sweeps. The beat rate between the partials the interval brings together stops being a number and becomes a function of time — and the pair spends four fifths of its time above the rate at which beating is beating at all.

Why a concert hall is narrow. The lateral energy fraction at the middle seat as the same hall is widened, everything else held. It peaks at 12 metres across at 0.235 and falls to 0.000 at 44. A wide hall's side walls are further away, so their reflections arrive later, weaker and — this is the part Sabine's model cannot say — from nearer the front, where the sideways weighting discounts them. The shoebox halls the orchestral repertoire was written for are all between about eighteen and twenty-five metres wide, and this is the arithmetic they are the answer to. Perception and the listener

A room with directions in it

Every room until now has been a reservoir of energy that drains at a rate. That model has no directions in it at all, so it cannot say the one thing every published measure of spaciousness is about: how much of what arrives comes from the side. Mirror the source in six walls and every reflection acquires an angle and a time — and the answer to why a concert hall is narrow falls out at eighteen metres.

How much of each spectrum a listener can assemble into one note. Each partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there. Perception and the listener

The spectrum that will not fuse

A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.

Adding parts adds power, and very little loudness. Each part is played at the same level, and the chord is realised every way its parts allow and averaged over them, so the quantity is a property of the texture rather than of one arrangement. Going from 3 parts to 8 adds 4.3 decibels of power and 0.1 decibels of loudness, because the extra parts land in bands that are already occupied — the count of occupied critical bands FALLS from 6.0 to 3.9 as the parts crowd into the same register. Form and structure

The dynamics are in the score already

Count the parts in each bar, realise them in their ranges, put every partial in its critical band, sum the loudnesses and run the result through the two smoothers built earlier. What comes out is a dynamic curve for a piece with no performance in it anywhere — and it says that doubling the number of parts inside a fixed register adds three decibels of power and about one of loudness, because the extra parts land in bands that were already occupied. Let the register widen with the parts and the same arithmetic gives eight phon, which is what a tutti actually is.

Roughness and loudness do not rise together. One four-note chord, played at levels from 35 to 95 decibels, with both quantities drawn as multiples of what they are at the quietest. Roughness is quadratic in pressure, so 60 decibels multiply it by 1.0e+6. Loudness is compressive — about ten phons to a doubling of sones — so the same range multiplies it by 96. The gap between the two lines is the quantity: roughness per sone rises by a factor of 1.0e+4 between a pianissimo and a fortissimo of the same chord. Harmony and voice leading

The ranking survives the dynamic and the chord does not

Roughness is quadratic in pressure and loudness is compressive, so sixty decibels multiply a chord's roughness by a million and its loudness by ninety-six. Roughness per sone therefore rises ten thousandfold between a pianissimo and a fortissimo of the same four notes — and yet the ranking of which doubling is smoothest, over four hundred and eighty voicings, does not move by a single place.

Every arrival at one seat, against the delay at which it would be an echo. The echogram: each reflection at its delay after the direct sound and its level relative to it, in a 22 by 45 by 15 metre hall with 18 per cent absorption. The line is the published echo threshold for speech — 40 milliseconds at equal level and about 3.0 more for each decibel of attenuation — so anything to the RIGHT of it is late enough and loud enough to be heard separately. The once-reflected rear wall arrives at 210 milliseconds, 25 decibels down, against a threshold of 114 — well past it. Nothing here stands clear enough of its neighbours to be heard as an echo, and 181 arrivals are fused with the direct sound instead. Perception and the listener

An echo is prevented by the crowd around it

The echogram says when every reflection arrives and how loud it is; the published echo threshold says when a reflection that late and that quiet is heard separately. Put one against the other and the rear wall of every hall anybody builds is past the threshold — a 45-metre hall puts it 210 milliseconds late and 25 decibels down against a threshold of 114. It is not heard as an echo, and what saves it is not the geometry. It is everything else arriving at the same time.

Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together. Perception and the listener

The partials that do not die together

The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.

A vibrato flattens the dissonance curve. Each interval twice: hollow is its roughness computed at the two notes' nominal frequencies, filled is the average of its roughness over a vibrato cycle of 50 cents at 6 hertz. They are not the same number, because roughness is a curved function of the frequency difference and the average of a curve is not the curve of the average. The largest effect is at octave, where the moving average is 19.3 times the still value — an interval sitting in a deep narrow minimum is smeared out of it. The smallest is at major seventh, where it is 0.95: an interval near a maximum is smeared out of that too. Vibrato pushes every interval toward the middle, and what it takes away from the consonances is much more than what it takes away from the dissonances. Intervals and chords

A roughness with a rate of its own

Every roughness figure so far computes one number for a steady spectrum. Evaluate the same sum at every instant of a vibrato and there are three numbers instead — a mean, a depth and a rate — and the mean is not the roughness of the mean frequency. On an octave it is nineteen times it, because an octave sits in a deep narrow minimum and a vibrato smears it out of one.

Which interval gives a tuner the deepest null. A tuner listening to an interval p:q is listening to the lower note's p-th partial against the upper note's q-th, and how deep the beat's trough goes is decided by those two amplitudes rather than by the interval. For a string spectrum, whose partials fall as one over n, the minor third pairs partial 6 against partial 5 at a ratio of 1.18 for a dip of 21.8 decibels; the major third pairs partial 5 against partial 4 at a ratio of 1.25 for a dip of 19.1 decibels; the fourth pairs partial 4 against partial 3 at a ratio of 1.32 for a dip of 17.2 decibels; the fifth pairs partial 3 against partial 2 at a ratio of 1.52 for a dip of 13.8 decibels; the major sixth pairs partial 5 against partial 3 at a ratio of 1.65 for a dip of 12.2 decibels; the minor sixth pairs partial 8 against partial 5 at a ratio of 1.67 for a dip of 12.0 decibels; the octave pairs partial 2 against partial 1 at a ratio of 2.00 for a dip of 9.5 decibels. The best is the minor third at 21.8 and the worst is the octave at 9.5, which is the reverse of the order a tuner is usually taught to trust: the deepest null in the list is on the interval whose coincidence sits highest in the spectrum, where adjacent partials are nearly equal in strength. Intervals and chords

A beat has a depth, and six essays held it at one

Every beat figure so far adds two tones of equal amplitude, which is the single ratio at which the trough of a beat is a true null — and a null is what a tuner is actually listening for. Vary the ratio and the picture changes: at two to one the dip is nine and a half decibels, at ten to one it is under two, and the interval that gives the shallowest null of all is the octave.

The same roughness, before and after the window it has to be heard through. The instantaneous roughness of an interval under a vibrato, and the same quantity after a running average of 59 milliseconds — the time a dissonance has to last to be heard as one, which is 4 cycles of this interval's own 68-hertz fluctuation rather than a number chosen for the figure. The mean is identical to every digit, 0.1487 against 0.1487, because a running average cannot change an average — so the earlier Jensen factor of 1.0 survives the window untouched and its prediction that the window would shrink it is wrong. What the window destroys is the depth: 0.30 of the mean becomes 0.23, which is 77 per cent. The roughness a vibrato adds is heard; the fact that it is moving is mostly not. Intervals and chords

The mean survives the window

A roughness that moves has a mean, a depth and a rate — all three of which a listener could only have through a temporal window. Applying the window already to hand settles which of the three survives, and the answer refutes the guess: a running average cannot change an average, so the octave's factor of nineteen stands and the movement is what goes.

Where the sound is, and how wide. An earlier essay sorted a hall's arrivals into echoes and everything else. Everything else is not nothing: a reflection too early to be heard as a separate event still moves the apparent source, widens it and colours it, and all three come out of the same list of times, levels and angles. At this seat the direct sound arrives from 26.6 degrees off the front and the image is heard 8.5 degrees left of it, pulled by the near side wall. The apparent width is 33 degrees, from a lateral energy fraction of 0.23. And the strongest early reflection arrives 0.6 milliseconds behind off 1× the floor, which is a comb filter with notches every 1608 hertz and 25 decibels deep. The trading ratio and the discount on late arrivals are stipulated rather than measured, so the degrees are ordinal: what the figure claims is the direction and the shape, not the number. Perception and the listener

A position and a width

Sorting a hall's arrivals into echoes and everything else settled that fusion is not a yes or a no. Everything else is not nothing: a reflection too early to be heard separately still moves the apparent source, widens it and colours it. All three come out of the same list of times, levels and angles, and none of them needed a new input.

What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged. Perception and the listener

The exchange rate nobody has

There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.

A loud chord is a smaller chord. The share of a voicing's partials that stand above what the rest of it masks, and the share of its computed roughness that is between partials a listener actually has, from 30 decibels to 100. Both fall: 83 per cent of the partials survive at 30 decibels and 38 at 100, and the roughness share goes from 88 per cent to 67. The direction is the upward spread of masking, which grows faster than linearly with level: a loud partial masks a band above itself much wider than a quiet one does, so the chord's own top disappears into its own bottom. Two earlier essays are drawn at one level, and this is what they were holding. Harmony and voice leading

A loud chord is a smaller chord

Two earlier essays hold the level fixed, and the level decides how much of a chord a listener is given. At thirty decibels twenty of a triad's twenty-four partials stand above what the rest of it masks; at a hundred, nine do. Every roughness figure until now counts partials that are in the score, and a partial the chord masks is not a partial the listener has.

The period is still there, and it is wider. The autocorrelation of a 12-partial complex on 220 hertz, drawn twice: steady, and averaged over one cycle of a 71-cent vibrato. A vibrato moves every partial by the same number of cents, so the complex is exactly harmonic at every instant and nothing is mistuned — what moves is the period the extractor is looking for. The peak survives. It loses 6 per cent of its height above the surrounding lags and gains 11 per cent in width, because the vibrato swings the period by 0.37 milliseconds against a peak 0.90 wide. Its maximum also moves, to 2.4 cents sharp of the still tone's, which is a prediction with a sign in it. Instruments and their design

The pitch that does not wobble

Three earlier essays have treated a vibrato as a modulation of roughness. The reason singers use one is what it does to the note, and there is an extractor here that turns a set of partials into a pitch and has never been asked what it does with partials that will not hold still. The period survives, at a cost that rises with the extent — and the practice stops within a hair of where the cost becomes total.

Intonation is a unison problem and nothing else. The roughness between two instruments on one note, against how far apart they are in cents, drawn for a unison and for the intervals beside it. A perfect unison is 0.0007 — the partials coincide and there is nothing to beat. Five cents apart it is 0.0465, 65 times as rough, and ten cents apart it is rougher than a major third played exactly. The mechanism is that partial n of a note mistuned by c cents is mistuned by c cents as well, which is n times as many hertz — so the top of the spectrum enters the critical band long before the fundamental does. The other curves are flat, because a third's roughness is set by which partials nearly coincide and a few cents does not change which. Timbre and acoustics

Two players on one note

Six essays have put one instrument on each note of a chord, and the commonest thing an orchestrator actually does is put two on the same note. Two independent sources add in power, so the composite is neither of them — except that it nearly always is one of them, because the level at which ownership changes hands is rarely at zero. And a unison ten cents out is rougher than a major third dead in tune.

One contrast survives every tempo anybody plays and the other does not. How much of each quantity's contrast between chords a listener still has at the end of each chord, against how long a chord lasts. The roughness curve is flat at one down to 45 milliseconds a chord and then falls off a cliff, because its window is 37 milliseconds and a boxcar either fits inside a chord or does not. The loudness curve is already losing at a second a chord and keeps 83 per cent at the slowest pace here, 39 at the fastest. Nothing in music is faster than the roughness window and a great deal of music is faster than the loudness one, so a passage delivers its dissonance and averages its dynamics. Form and structure

The dissonance arrives and the dynamic does not

A scoring decides two things at once and both of them have to be integrated by a listener before they exist. The loudness smoother's release is two seconds and the roughness window is thirty-seven milliseconds, and that ratio of fifty decides which of the two survives at the pace music is actually played. Nothing anybody performs is fast enough to blur a dissonance, and a great deal of it is fast enough to average a dynamic.

The hall, as two numbers a listener has. Every reflection at this seat, placed by the interaural delay it produces rather than by the direction it came from. Time runs down; the dot's size is its energy. The direct sound is at 0 microseconds and the reverberation spreads over the whole available range, with a root-mean-square width of 322 against a geometric maximum of 656. That is the position and width computed earlier, in the units a listener has instead of the vectors used until now. 10 of the 56 reflections arrive from behind and carry 9 per cent of the energy — and they are drawn where they are because the interaural delay of a reflection from 120 degrees is identical to one from 60. Perception and the listener

The hall through a head

Every direction computed so far is a vector from a seat to an image source, and a listener has no vectors. Run the echogram through the head computed six essays ago and two things happen: the position and width become microseconds, and a third of the room disappears — because both cues fold at ninety degrees and a reflection from behind is identical to one in front.

Nothing at all until fifteen decibels, and then it depends on the tempo. The fraction of a line's partials that stay above threshold, over how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement — forward masking leaves a threshold at most ten decibels below the masker, and a note's own partials mask each other from the same components at full level. The boundary is between twelve and eighteen decibels, and beyond it the loss grows with the tempo: at 280 to the crotchet and 36 decibels of contrast, 26 per cent of the line's partials are gone. Fifteen decibels is about the gap between a forte and a piano. Perception and the listener

An equal note cannot be masked

Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.

The cue that settles it. Every spectrum to hand, arbitrated by the earlier competition and then again with the onset cue added at equal weight. 3 of the 5 change their verdict, and all 3 change the same way — from splitting into two streams to staying as one: a piano string, a bell, a bar. Nothing changes the other way, because the onset cue on a struck source votes for fusion on every partial and can only ever push toward one stream. The bell is the case worth naming: its partials are wildly inharmonic and it is heard as one sound, which is a fact the harmonicity cue alone cannot produce. Perception and the listener

The cue that settles it

Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.

A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance. Perception and the listener

A page has two decibels

The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.

A turn of 0.99 degrees tells front from back. A source 45 degrees off centre and its mirror image 135 degrees off, which produce the same interaural delay and are therefore the same signal to a listener who does not move. As the head turns the two predictions separate: the front source's delay falls and the rear source's rises, because the fold at ninety degrees puts them on opposite branches of the same curve. They differ by the 15-microsecond threshold after 0.99 degrees of turn — which is exactly half the 1.97 degrees a source would have to move for the same listener to notice it moving, and it is half for a reason: a turn displaces the two hypotheses from each other by twice what it displaces either of them from where it started. Perception and the listener

The turn is half the angle

A stationary head cannot tell a sound in front from the same sound behind, and an earlier essay said so at length. The turn that breaks the confusion is 0.84 degrees — exactly half the angle a source would have to move for the same listener to notice it moving, and half for a reason. In a hall the same turn does something else: the source swings at 8.9 microseconds a degree and the room swings at 2.7, so a listener who moves is separating the soloist from the reverberation as well as the front from the back.

The same geometry at four sizes of head. Woodworth's interaural delay against direction, for 4 head radii from 5.8 to 9.8 centimetres. The whole range runs from 431 microseconds for a newborn to 735 for a large adult, and it scales exactly with the radius because the delay is (r/c)(θ + sin θ) and r is a multiplier. The detection threshold does not scale with the listener, so the number of distinguishable delays across the whole range falls from 98 to 57: a smaller head has the same directions in front of it and a shorter ruler to measure them with. Perception and the listener

A smaller head in the same hall

Ten earlier essays draw one head. Every parameter belonging to the room has been varied by some figure and the one belonging to the listener never has, and it is the only one whose change the detection threshold does not follow: a six-year-old in the same seat receives the same fifty-four reflections at the same instants and reads them onto an axis with seventy distinguishable positions instead of eighty-seven. The speed of sound, swept over every temperature a hall is ever at, changes nothing at all — and the reason it cannot is the reason head size can.

A clarinet's partials, each on its own resonance. The 8 partials of the clarinet's chalumeau D that ride an impedance peak, each building toward its steady amplitude as 1 − exp(−t/τ) with τ = Q/πf from that peak's own Q. The time constants run from 11.9 milliseconds to 64.5, so the partials do not arrive at different times — they all begin the instant the reed does — and what differs is how fast each approaches its final level. The horizontal bars are how far apart the first and last are at three criteria: 5.5 ms at 10 per cent, 36.4 ms at 50 per cent, 121.0 ms at 90 per cent. A twenty-millisecond asynchrony is the threshold for hearing a partial out of a note, and this note crosses it at 32 per cent of steady amplitude — so whether a blown note's onset cue is unanimous or divided is decided entirely by how far along a partial has to be before it counts as having started. Perception and the listener

A blown note does not start late, it starts slowly

Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.

Eleven partials is one partial too many. What fraction of a spectrum the harmonicity census finds fused, against how many partials it is asked to census. At ten a perfect harmonic series fuses 10 of 10 and the fundamental it finds is the right one. At eleven it fuses 5 of 11 and the fundamental jumps to exactly 2.00 — the octave above. The cause is the cap the census carries for a reason established earlier: without it a bell fuses perfectly at a fundamental nobody could hear, so the search refuses any fundamental more than about ten harmonics below the top partial. At eleven partials the first thing that cap excludes is the series' own fundamental, and the census then takes the octave and calls every odd partial inharmonic. So the number of partials and the cap are the same number, and nothing had ever said so, because every earlier figure censuses ten. Perception and the listener

Eleven partials is one too many

Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.

An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none. Perception and the listener

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

Articulation is worth 2.6 phons, and nobody counts it. A passage of notes at 80 decibels, 2 to the beat, at 7 tempi and 3 articulations, scored against the same level held continuously. The variable is the fraction of each inter-onset interval that is sounding — 0.95 is a legato, 0.4 a staccato — and the vertical axis is what that costs the passage's running loudness in phons. Nothing here is anybody playing harder or softer. At 40 to the beat the span from legato to staccato is 1.47 phons; at 200 it is 2.61, because a staccato note there lasts 60 milliseconds and no longer reaches its own loudness either. Perception and the listener

A staccato is a dynamic mark

Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.

The top voice arrives whole and the bottom one arrives as a sine. 4 parts sounding together, each of 8 partials, with every partial tested against the summed masked threshold of every component in the texture. A filled mark is a partial the listener receives and an open one is a partial the part would have had alone and does not have here. The bass at C3 keeps 1 of 8, the tenor at C4 keeps 2 of 8, the alto at E4 keeps 5 of 8, the soprano at G5 keeps 8 of 8, every one of them at 70 decibels. Every part is at the same level and the difference is entirely where each one sits: masking spreads upward, so the part at the top of the texture has nothing above it to be masked by and the part at the bottom has everything. Perception and the listener

The listener is given the top voice, and the bass as a sine

Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.

The roughest chord on the page is at C1 and the roughest one heard is at A♭2. One close triad at 70 decibels through 17 registers, with its roughness computed twice: over every partial in the score, and over only the partials that stand above what the chord itself masks. The written curve rises all the way down and its maximum is the lowest register drawn, C1, which is the low-interval rule as it has always been computed here. The delivered curve turns over at A♭2 and falls to nothing below A♭1: a close triad down there is not rough, because it is not arriving as a chord — 1 of its 24 partials survives at C1 and there is almost nothing left to beat against anything. Perception and the listener

A low chord stops being rough by stopping being a chord

Every count of audible partials until now is of a chord at middle C, and the three registers it did compare span C3 to C5 — a third of the range a chord is written in. Move the same triad down and the count collapses: 79 per cent of its partials arrive at E3 and 4 per cent at C1. So the roughest chord on the page is the lowest one and the roughest chord a listener receives is at G2, and where that maximum sits moves nearly two octaves with the dynamic.

Who owns clarinet, oboe, voice at every balance. The composite of three players at 392 hertz belongs to whichever of them it is nearest in log-spectral distance, and here that is drawn over the whole plane of balances a conductor could set — the second and third players from 24 decibels below the first to 24 above. voice owns 79 per cent of the square. The three regions meet where all three distances are equal, which is the only balance at which the composite belongs to nobody: it is at -0.3 and -19.1 decibels, inside the square and therefore a balance an ensemble could actually be asked for. A trio has a colour of its own at one point, not over a region. Timbre and acoustics

A section has a loudest member, not a colour

Two players on one note have a balance at which the composite belongs to neither, and that is what blending means. Three should have three such balances and no reason for them to agree — a trio with a rock-paper-scissors ownership would have no strongest member at all. Twenty trios, sixty pairwise comparisons, and not one disagreement: the possibility is real, arbitrary spectra do it once in twenty, and instruments never do.

Which pairs blend is a question about the note. The level at which a doubled pair's composite changes owner, drawn for all 15 pairs of 6 radiators over 2.6 octaves from 131 to 784 hertz. A pair blends when that level is inside the shaded band, which is the twenty-four decibels either way two players can manage; a curve outside it, or absent, is a pair one instrument owns at every balance. 6 of 15 pairs blend at the bottom of the range and 12 at the top. Every filter in this collection is fixed in frequency and the fundamental is not, so a radiator's shape is a function of pitch and so is everything computed from two of them — the blend ranking at the bottom and at the top disagree on 70 of 105 comparisons, which is more than half, so the order has turned over rather than merely shuffled. Timbre and acoustics

The blend table has a row for every note

Eight earlier essays sound their instruments at one note, and one of them says why that cannot be innocent: every filter here is fixed in frequency and the fundamental is not. Swept over four octaves, the number of pairs that blend doubles from six to twelve, the ranking turns over rather than shuffles — seventy of a hundred and five comparisons swap — and a clarinet with an oboe goes from the best pair in the collection to the eleventh.

What the chord before takes out of the chord after. Five chords at 1.2 seconds each, with the roughness each one has on its own — its simultaneous masking and the threshold of hearing already applied — and the roughness it actually has once the chord in front of it has raised the threshold. Four of the five are untouched. The fifth, i, two parts, follows the only step in this passage that falls more than fifteen decibels, and it arrives into a hole: it is entirely below threshold for its first 13 milliseconds and takes 240 to get all of itself back. Masking can only remove partials, so it can only lower a roughness — and the chords it can reach are the ones a written dynamic has just made quiet, which are already the smooth ones. Across the passage the dissonance contrast goes from 3021 to 3113: the mask widens it by 3.0 per cent rather than eating it. Form and structure

A soft chord has to fade in

Forward masking sits between the two integration times already in play — two hundred milliseconds against a thirty-seven millisecond roughness window and a two-second loudness release — and it was owed as the term that might eat the dissonance contrast. It does not. It widens it, by three per cent at a chorale's pace and fifty-nine at four chords a second, because it can only ever remove partials and it can only reach the chord a dynamic has already made quiet. What it does instead is stranger: one chord in the passage is entirely inaudible for its first twelve milliseconds and takes a quarter of a second to arrive whole.

Of 8 beats at A3, 3 can be attended to. Every member of the beat family a 2:1 mistuned by 6.0 cents makes on a real string at A3, placed by how separable its coincidence is from the partials beside it — in auditory filter widths, across — and by how fast it beats, up. The shaded region is the set a listener can receive: wider than one filter, and between 0.4 and 15 fluctuations a second. 4 of 8 members fall inside it, and once rates within a factor of 2 are counted as one modulation channel there are 3. The members that fail do so for two different reasons: the low ones sit under the roughness ceiling but their coincidences are buried in a filter that holds three partials, and the high ones are resolved and far too fast. Intervals and chords

Three beats at most, and only in the middle of the keyboard

A mistuned octave on a real piano makes eight beats at once, and a listener attending to one of them is doing something that has a threshold. Two thresholds, in fact — a rate and a place — and once both are applied the eight become four at A3, one at A1 and one at A5. Every interval a tuner sets goes to zero at both ends of the compass and peaks at eight countable beats in the octave the bearing is laid in.

The top of a struck note's series falls while the note lasts. The highest partial still above the threshold of hearing, against time, for a note struck at 80 decibels on a fundamental of 130.8 hertz with a 1/n spectrum and a loss rising as the partial number to the power 1. It starts at partial 152 and it is falling from the first millisecond. The horizontal lines are the three tops computed from frequency alone, all of which assume a note that never ends: the note's own top drops past the difference-limen top after 0.03 seconds, past the semitone top after 0.32, and past the resolvable top after 0.64. After that the series is shorter than the ear could have resolved, and what stops it is the clock. Timbre and acoustics

The top that falls while the note lasts

Four earlier essays have asked where the harmonic series stops, and all four answered with a number computed for a tone that never ends. A struck C3 has 152 audible partials at the strike and eight after two-thirds of a second, so the ear's own resolution limit governs the first ten per cent of the note and the decay governs the rest. Playing ten decibels louder buys the ear a further tenth of a second, and doubling its reign would take fifty-eight.

Partials 3, 4, 12 of "hod", over one vibrato cycle. The level of three partials of a 220 hertz note on the vowel in "hod", each about its own mean, over one cycle of a vibrato of ±71 cents at 6.0 hertz. The pale curve is the frequency deviation itself, for phase reference. Partial 3 at 660 hertz swings 4.22 decibels and peaks with the frequency; Partial 4 at 880 hertz swings 0.41 decibels and peaks twice a cycle; Partial 12 at 2640 hertz swings 8.62 decibels and peaks against it. The formants of this vowel are at 730, 1090, 2440 hertz and do not move; a partial below one rises as the frequency rises and one above it falls, so the modulations of a single note run in opposite directions at the same instant. Instruments and their design

The partial that gets louder as it goes sharp

Eleven earlier essays sweep a set of partials and hold their amplitudes still, and a real tract does not move with the fundamental. Put the formant account and the sweeping account together and every partial acquires an amplitude modulation at the vibrato rate: 0.07 decibels on the fundamental and 8.6 on the twelfth partial of the same note. They are in phase below a formant and anti-phase above one — not ninety degrees apart — and the whole note swings 0.82 decibels, because they cancel.

Where a section's fluctuation stops being a beat, on 220 hertz. Two rates up the spectrum of a 220 hertz note sung by a section whose voices are spread by 15 cents. The rising line is the beat rate between a typical pair of them, which grows with the partial because a mistuning in cents is a difference in hertz that scales with frequency; it reaches the 15 hertz at which a beat stops being a beat by partial 5.5, at 1217 hertz. The flat line is the amplitude modulation the vibrato imposes through the formants, which is 6.0 hertz at every partial because the vibrato modulates every partial by the same number of cents at the same rate. The two are equal at 487 hertz. Above 1217 hertz the beating has become roughness and the only fluctuation left is the vibrato's — and that frequency is the same one an octave up, where it is partial 2.8 instead. Instruments and their design

The rate that does not rise with the partial

Twelve earlier essays give every vibrato the same six hertz, and the measured spread is 5.5 to 7.5. Putting the two fluctuations a choir contains on one axis shows why the rate matters: the beating between mistuned voices rises with the partial and leaves the range a listener follows as fluctuation at 1,217 hertz, while the vibrato's own modulation is six hertz at every partial. Above that frequency a section fluctuates by vibrato alone — and if every singer had the same rate, it would barely fluctuate at all.

14 players summed, against the one at their average onset. Each thin line is one player's rising envelope, started at its own moment, with a spread of 30 milliseconds about the beat and a 90-millisecond attack. The heavy line is the section: nominally identical sources add incoherently, so their powers add and the sum is the root of the mean of their squares, drawn here as a fraction of the section's own peak. The dashed line is the single player who started at the section's average onset. The section reaches the 6 dB below peak criterion at 18.1 milliseconds and that player at 27.6, a difference of 9.5. The section is early because the players who started first are already sounding while the average one is still building, and nothing a late player does can make the sum quieter. Rhythm and metre

Twelve violins are more punctual than one

Every essay until now treats a part as one player, and an orchestral part is a dozen. Sectioning does two things at once and only one of them was expected: it pulls the part's heard moment forward, by four milliseconds against a map spanning twenty-six, and it makes the part's arrival more accurate by very nearly the root of the number of players. So the map of required leads applies to an orchestra better than it applies to a quartet, and the case where it fails is three trumpets rather than fourteen violins.

A head model a few millimetres out reads every azimuth but the front. The azimuth a listener reports against the azimuth a source is at, for internal head radii from 8.22 to 9.28 centimetres against a true radius of 8.75. The delay a source produces is (r/c)(θ + sin θ) and the listener inverts it with the radius they believe they have, so their answer solves θ̂ + sin θ̂ = (r/r̂)(θ + sin θ). Every curve passes exactly through the origin: on the median plane there is no delay and therefore no error, whatever the head model is. The error grows with azimuth and is largest at the side. An internal head 5.3 millimetres too small runs out of azimuth at 82 degrees: beyond that the world is delivering a delay larger than any its owner's model can produce, and every source out there collapses onto the side. Perception and the listener

Where a wrong head gives itself away

Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.

A 20-decibel crescendo is 27 phons on a bass note and 20 on a high one. The same change of level, from 60 to 80 decibels, converted to loudness at each register through ISO 226's equal-loudness contours rather than at one kilohertz. The heavy curve gives each note a string spectrum, so its partials are converted in their own bands and summed; the pale one is the fundamental alone. On the spectrum-aware curve the crescendo is worth 27.3 phons at C1 and 20.3 at C7. On the fundamental alone it is 77 at C1, which is not a finding but an artefact: a 60-decibel tone at 33 hertz sits 1.8 decibels above the threshold of hearing and is very nearly nothing. The honest correction is the smaller one, and it is still a difference of 7.1 phons across the compass for a mark written in the same ink. Perception and the listener

A subito piano is four seconds longer in the bass

Every loudness figure with time in it converts level to loudness at one kilohertz, and the equal-loudness contours say that no other frequency works that way. Joining the two sorts the published numbers into those that were about the treble and those that were not. Three move a great deal — a twenty-decibel crescendo is worth 27 phons on a bass note and 20 on a high one, and the seven seconds a subito piano takes becomes eleven and a third. Three do not move at all, and the reason they do not is the same reason in every case.

Scored the way these figures score it, a clarinet is the worst of the six. One close triad at 70 decibels through 17 registers, drawn once for each of the 6 spectra to hand. The score is the share of every partial written, which is the quantity the register figure published. pure 100 per cent at best, string 79 per cent at best, clarinet 54 per cent at best, reed 71 per cent at best, bell 52 per cent at best, organ 72 per cent at best. A clarinet's four even partials are twenty-eight decibels below its odd ones and are inaudible beside their own neighbours before any chord is built, so counting them in the denominator makes the spectrum that survives its own masking best look like the one that survives it worst. Perception and the listener

A clarinet keeps what a string loses

Every masker, probe, chord, line and texture until now is eight partials falling as 1/n, and it was not even an option a placement could pass. Sweeping the six spectra to hand says the clarinet is the worst of them — 54 per cent of itself at best against a string's 79 — and that answer is an artefact of the score. Counted against what each note keeps on its own, the clarinet keeps 100 per cent where the string keeps 79, because its components stand a twelfth apart rather than an octave. The missing parameter was the spectrum; the second missing parameter was the denominator.

Both notes of a fifth, and the coincidences going out. The highest partial still above the threshold of hearing for each note of a fifth struck at 80 decibels on 130.8 hertz, against time, with a 1/n spectrum and a loss rising as the partial number to the power 1. Both curves come down from off the top of the frame — the lower note starts with 152 partials and the upper with 102, because twenty kilohertz is a ceiling in frequency and not in partial number. The rings are the interval's partial coincidences at the moment they stop existing — successive multiples of one ratio, which go out from the top down: 12:8 at 0.47 s, then 9:6 at 0.64 s, then 6:4 at 1.04 s, then 3:2 at 2.14 s. The lowest, 3:2, is the last, and after it the two notes have no partial in common that either of them can still supply. Timbre and acoustics

A fifth on a piano is not a fifth a second later

Nine essays draw one note in silence, and a listener is given a texture. Put two struck notes an interval apart and consonance comes apart into two quantities that had agreed while nothing moved: the partial coincidence that names the interval outlives the strike in exactly the order common-practice theory ranks its intervals, and the roughness that scores it reorders itself inside a third of a second, with the fifth overtaken by four intervals every treatise calls harsher.

The attack is the balance dial, turned by the clock. The level of a violin against a clarinet on one note at 392 hertz, moment by moment through the attack, with both players starting together. Two envelopes rising at different rates are a balance, so this axis is the same dial a conductor turns — and its whole travel is 6.02 decibels, which is twenty times the log of the ratio of the two attack times, 45 against 90 milliseconds, and nothing else. The pair does not begin as one player alone: both envelopes leave zero at the same slope ratio, so the dial starts at a finite offset rather than at silence. The dashed line is the balance at which the composite changes owner, -3.48 decibels — inside the travel, so the note belongs to a clarinet for its first 29 milliseconds and to a violin for the rest of its life. Timbre and acoustics

The blend arrives before the note does

Nine essays on spectrum draw a steady state, and the strongest cue that two instruments are two instruments is that they do not start together. Two envelopes rising at different rates turn out to be a balance — the same dial an earlier essay swept — so the attack is that dial moved by the clock, and its whole travel is fixed at twenty times the log of the two attack times. It is six decibels for a clarinet with a violin against a crossing twelve to twenty-two decibels out, so one pair in ten changes hands during its own attack, and which one depends on a convention rather than on the instruments.

Every beat in a family is the same depth, and none is near the threshold. The 8 members of the beat family a octave mistuned by 6.0 cents makes at A3 on a real string, each placed at the rate it beats and at the modulation index a listener's filter delivers there. The rising curve is the published detection threshold for amplitude modulation, which is flat at 0.03 below about fifty fluctuations a second and rises above it. The flat dashed line at 0.667 is the depth the pair has in isolation, and it is the same for every member: on a spectrum falling as one over n, the k-th member pairs partials 2k and 1k, whose ratio is 2.00 whatever k is. What the filled points show is the smaller effect that does depend on the member — the partials on either side of the coincidence leak into the same filter, add level without adding fluctuation, and dilute the index from 0.662 to 0.405. Inside the countable rate window the narrowest margin over the threshold is a factor of 16.7, on member 4. The depth criterion removes nothing. Intervals and chords

Every member of a beat family is the same depth

A mistuned octave's beats have been counted on their rate and their place, with the depth recorded as the thing left out and a prediction that the shallow upper members would take the count from four to two. The depth turns out not to fall at all: on any power-law spectrum every member of a family has exactly the modulation index its interval's own ratio gives, at every register, on every wire. The count does fall to two, and the thing that takes it there is the criterion that essay was already using.

An entrance stops being a loudness event and never stops being a colour one. The same oboe entering on the same note at the same level, against how many players were already sounding. Its contribution to the loudness falls from 15.5 phons to 0.32 — a factor of 48 — and crosses the one-phon difference limen at 5 players already playing. Its contribution to the roughness rises by a factor of 12.3 over the same range, because roughness is a sum over pairs and the entrant makes one new pair with everybody. Both curves are drawn as a share of their own largest value, since a phon and a squared pascal have no exchange rate. The claim is the two directions, not the crossing point of two units. Form and structure

An entrance is a change of colour

Eight essays on orchestration move the assignment and hold the ensemble still, and a score does the opposite: it brings players in and takes them out. Loudness is a sum over parts and roughness is a sum over pairs, so the player who joins adds one term to the first and one to the second for everybody already there. What the entrance is worth in phons falls by a factor of forty-eight across the range an ensemble spans and crosses the difference limen at five players; what it is worth in roughness rises by twelve, and by a further factor of ten for every ten decibels the passage is played at.

Three clocks receive one entrance, and they do not agree about when. An oboe joining 5 players already sounding, at time zero, with each of the listener's three readings drawn as its own share of the change it eventually makes. The roughness window is 49 milliseconds wide and has half the change at 25; the short-term loudness smoother has half at 15; the long-term one, whose release is the two seconds an earlier essay is about, has half at 90. The two-second release is on the wrong side of the smoother to hide an entrance. Its attack is 99 milliseconds, so an entrance is received promptly and it is a departure that is not. Form and structure

The release is on the wrong side

Whether the loudness model's two-second release makes an entrance inaudible has the answer no, for a reason the question did not anticipate. The smoother is asymmetric — ninety-nine milliseconds going up and two seconds coming down — so a rise is tracked twenty times faster than a fall, and an entrance is received promptly by every one of a listener's three readings. The colour of it arrives first, at twenty-five milliseconds against ninety, and the reading that moves with the ensemble is the one nobody would have picked.

Which chord of a passage has room for the part that is entering. An oboe entering on one note, tried at each chord of a five-chord passage, scored by how far its own partials sit above the threshold the ensemble already sounding puts over them. The best moment gives it 9.0 decibels of margin and the worst 0.3, a spread of 8.7 — and the best moment is not the quietest chord, which is vi, close below. Room for an entrance is spectral rather than dynamic. A chord with a hole in its written spacing need not have one in its spectrum, because the partials of its bass fill the middle whatever the notes above it do. Form and structure

The chord that has room for an entrance

Three essays have made the ensemble something a score can change and none of them has asked when. The ensemble already sounding puts a masked threshold over whatever register an entering part takes, and that threshold is set by the voicing rather than by the dynamic — so the five chords of one passage differ by 8.7 decibels in how much of an entering oboe survives them, and the quietest chord of the five is the worst place in the passage to bring somebody in. Swept over the entrant's own pitch, the choice of moment is worth as much as the choice of register.

The channels a fluctuation is analysed into, and how wide they are. A modulation filterbank of quality factor 1, drawn every half octave, with two of the beats a mistuned octave at middle C produces marked at 0.89 and 1.26 a second. A channel of quality one has its half-power points at 0.618 and 1.618 of its centre, so two fluctuations closer than a factor of 1.618 never end up in different channels. The two marked rates differ by a factor of 1.42, which is inside that. They are one fluctuation and not two — which is the question an earlier essay asked and left open, answered by a bank rather than by a convention. Pitch and tuning

One fluctuation or two

Whether two members of a beat family are one thing or two was decided by asking whether their rates differ by a factor of two, and the factor was written down as a stand-in for a modulation filterbank nobody had run. Run, the bank gives 1.618 — the golden section, and not by accident, since a channel of quality one has its half-power points there. The stand-in was conservative rather than optimistic, and the recomputed counts do not change at a single register, because the criterion was never what was binding.

A loudly struck note hides the beats it is being struck to reveal. The share of the fluctuation in its own auditory filter that belongs to each of the first five members of a mistuned octave's beat family, against how loud the note is. The filter's lower skirt shallows by about 38 per cent of its 51-decibel value every ten decibels, so more of the neighbouring partials get into the filter and the pedestal each member sits on grows. The mean share falls from 0.55 at 40 decibels to 0.07 at 100. Past about 90 decibels the curves are flat because the model's skirt is clamped rather than because anything stops changing — that clamp is the model's floor and not a measurement. Pitch and tuning

How hard the note was struck

The auditory filter is not a fixed shape: its lower skirt shallows by about 38 per cent of its reference value every ten decibels, so a loud note is analysed through a wider filter than a quiet one. Every share computed so far was quoted at a moderate level, and a tuner does not strike moderately. Recomputed, the mean share of a mistuned octave's filter falls from 0.55 at forty decibels to 0.07 at seventy, and the count of separable beats goes from two to none — which is a prediction too strong to be right, and the way it fails is the useful part.

A struck octave becomes countable a second after the strike, or never. The number of separable beats a mistuned octave on A3 delivers, second by second after both notes are struck at 80 decibels, with every partial dying at its own rate (a 12-second fundamental, losses rising as frequency to the power 0.7). Read with the filter broadened by the level of the whole note, which is how the level-dependent count was first computed, the count is zero at every instant: the partials fall below audibility before the filter has narrowed enough to separate them. Read with the filter broadened by the level inside itself, which is what the published parameterisation was fitted against, the count is 0 at the strike, reaches 2 at 1.0 s and falls to nothing at 3.3 s. Pitch and tuning

Counted in the decay, or not at all

A mistuned octave struck hard delivers no countable beat at the strike, and the reconciliation offered for that was that a tuner listens to the decay. Computed through a real decay it fails on its own terms: the partials fall silent before the filter has narrowed enough to separate them. It succeeds only when the filter is broadened by the level inside it, which is what the published parameterisation was fitted against — and then the window opens at a twelfth of the note's life and shuts at a quarter.

A tempered interval moves its difference tone several times further than itself. For every interval inside the octave tuned to twelve equal steps, how far the difference tone f₂ − f₁ lands from where the just interval would put it, in cents, with the interval's own departure from just drawn as the thin bar beside it. minor second −198.0 (the interval −11.7); major second −35.5 (the interval −3.9); minor third −96.0 (the interval −15.6); major third +67.4 (the interval +13.7); fourth +7.8 (the interval +2.0); fifth −5.9 (the interval −2.0); minor sixth −36.7 (the interval −13.7); major sixth +38.8 (the interval +15.6); minor seventh −39.8 (the interval −17.6); major seventh +25.0 (the interval +11.7). The major third's product is +67.4 cents out and the minor third's −96.0, and the largest error is the minor second's, at −198: a product moves p/(p − q) times as far as the interval p:q that made it. Intervals and chords

The third sound magnifies cents, not hertz

Tartini's third sound is said to be a few cents off on a tempered interval. It is sixty-seven cents off on a major third and ninety-six on a minor third, because a difference tone moves p/(p − q) times as many cents as the interval p:q that made it. In hertz it moves exactly as far as the note that moved, and no further — so what the magnifier is worth is the ear's finer resolution at the low frequency where the product lands, which is a factor of two for a long note and nothing at all for a short one.

Every product of a just interval is a harmonic of the note it implies. Each interval drawn as two harmonics of a fundamental it does not contain — the lower note is harmonic q and the upper harmonic p — with its three combination tones placed on the same numbering: the difference tone at p − q, the cubic product below the pair at 2q − p, and the one above at 2p − q. minor second 16:15: 1, 14, 17; major second 9:8: 1, 7, 10; minor third 6:5: 1, 4, 7; major third 5:4: 1, 3, 6; fourth 4:3: 1, 2, 5; fifth 3:2: 1, 1, 4; minor sixth 8:5: 3, 2, 11; major sixth 5:3: 2, 1, 7; minor seventh 9:5: 4, 1, 13; major seventh 15:8: 7, 1, 22. The shaded column is the fundamental itself. The difference tone sits on it for every interval up to the fifth, the cubic product for the fifth and every interval above except the minor sixth, whose products are its fundamental's octave and twelfth. Intervals and chords

The tone on the root changes hands at the fifth

Every combination tone of a just interval is a harmonic of a fundamental neither note contains, and which harmonic is fixed by the ratio. The difference tone lands on that fundamental for every interval up to the fifth; the cubic product lands on it for the fifth and every interval above except the minor sixth. So the loud product names the root of a narrow interval and the quiet one names the root of a wide one — and a just major seventh's difference tone is a note seven harmonics up that no keyboard has.

A major triad's combination tones, against its own notes. The three notes of a major triad on C4 in root position, voiced C4–E4–G4, as tall lines, and every combination tone its pairs make, as short ones: difference tones lowest, cubic products taller. In just intonation 2 cubic products land exactly on a note of the chord, and none comes within forty hertz of one. In equal temperament no cubic product lands on a note of the chord, and the nearest miss is 5.63 hertz. Intervals and chords

A major triad's combination tones are its own notes

Play a just major triad of pure tones and two of the ear's cubic products land exactly on its root and its fifth. The reason is a condition rather than a coincidence — a chord's cubic products fall on its own notes when its middle note is the mean of the outer two in hertz — and it holds for the major triad in root position and in the six-four, and for no minor triad in any position or tuning. Equal temperament misses the landing by one number, 5.6 hertz on middle C, which is a beat that belongs to no pair of notes in the chord.

A scale in parallel thirds has a line underneath it that nobody plays. A major scale on C4 harmonised in parallel diatonic thirds, with the difference tone f₂ − f₁ of each pair drawn as a third line. In five-limit just intonation that line is C2, A1, C2, F2, G2, F2, G2, C3. In equal temperament it moves to C♯2, A1, B1, F♯2, A♭2, E2, F♯2, C♯3, departing from the just line by +67, +33, −82, +69, +65, −80, −84, +67 cents. Intervals and chords

The bass line under a passage in thirds

A major scale harmonised in parallel thirds gives the ear a difference tone under every pair, and in five-limit just intonation those tones are a diatonic bass line — C, A, C, F, G, F, G, C — made of the scale's own notes. Tempered, the same line moves only by whole tones, a neutral third and a fourth stretched to 650 cents, and wobbles by up to 84 cents from note to note. In sixths the bass is drawn by the other product, because the cubic product of a pair is the difference tone of the same pair inverted.

A doubled pizzicato gives its note away while it is still the louder. The power of a violin plucked, against a flue pipe holding the same note at 392 hertz, through the first 600 milliseconds of the pluck, with the pluck starting 12 decibels up and its fundamental decaying over 1 second. With each partial losing level in proportion to its number, the composite stops resembling the pluck at 70 ms, when the pluck is still 5.2 decibels the louder. With every partial fading together it would keep the note until 543 ms. The dashed line is the balance at which the steady-state doubling changes owner, minus 20.6 decibels: the release crosses the owner long before its balance gets there, because what hands the note over is the pluck's upper partials going, not its level. Timbre and acoustics

A doubled pizzicato gives its note away early

The attack turns the balance between two players on one note by a few decibels and stops. A pluck does not stop — every partial of it decays, so a pizzicato doubled by a held instrument walks the balance for the whole note, and the expectation was a handover as slow as the decay. It is fast. A one-second pizzicato over a flute loses its note in 70 milliseconds, while it is still five decibels the louder, because what hands the note over is its upper partials going first. A uniform fade would have kept it eight times as long.

The schedule that hears every entrance best holds the high parts back. Six parts waiting to enter a five-chord passage over four sounding players, each entering once and staying: the schedule under which the least audible entrance is as audible as it can be made. brass on E3 enters at I, open with -1.4 decibels of mean margin over the mask; oboe on E4 enters at vi, close below with -2.6 decibels of mean margin over the mask; clarinet on G4 enters at I, open with -1.4 decibels of mean margin over the mask; voice on C5 enters at IV, close above with 2.1 decibels of mean margin over the mask; violin on G5 enters at V, bracketing with 6.9 decibels of mean margin over the mask; flue pipe on C6 enters at I, hollow with 4.0 decibels of mean margin over the mask. The least audible entrance is at -2.6 decibels and the margins sum to 7.6; of all 15625 schedules 0 have a better least audible entrance and 576 a larger sum. Form and structure

Room is used up by whoever enters first

The chord with the most room for a part entering alone is a fact about that chord. It stops being a fact the moment two parts want it, because each part that comes in raises the mask over everybody after it. Given six parts waiting to enter a five-chord passage, choosing each part's moment the way one part's moment is chosen puts three of them into the same chord and lands in the bottom fifth of all 15,625 schedules. Placing them one at a time does no better. The schedule under which the least audible entrance is heard best is unique, and it brings the low and middle parts in while the texture is thin and holds the three highest back for the last three chords — because a high part keeps its room over a full texture and a middle part does not.

Level does not dilute the register's roughness, it multiplies it. The mean roughness of the I – vi – IV – V – I arrivals at 4 registers, each relative to the register as written, read three ways. Level-free, the bass is 8.6 times rougher than the treble. With every note at 70 dB it is 8.6 times, the same factor, because one level rescales every pair alike. With each chord played at the level that makes it as loud as the written register's chords — 82.8 dB −2 octaves, 75.5 dB −1 octave, 70.0 dB as written, 66.8 dB +1 octave — the bass is 343 times rougher than the treble, because roughness grows with the square of the pressure and the bass needs more of it to be heard at the same loudness. Harmony and voice leading

A rough arrival is rough because of its spacing

The pair the expectation essays report for every chord — how surprising it was, how rough its voicing is — has no level in it. Putting level back in answers the question it left open, and not the way it was framed. At one written dynamic the arrivals keep their order from 40 to 90 dB at three registers of four, and the bass stays 8.6 times rougher than the treble. Made equally loud, the bass has to be played 12.8 dB harder, and it is 343 times rougher: level does not explain the register's roughness away, it multiplies it.

The thirds' difference-tone line, note by note, against the dynamic. Each note of the line the difference tone f₂ − f₁ draws under a scale in just thirds, as its level above the higher of the threshold of hearing and the primaries' masking, against the level of the primaries. The product sits 50 dB below the primaries at 60 dB and grows twice as fast as they do. C2 above the limit from 73.5 dB; A1 above the limit from 76 dB; C2 above the limit from 73.5 dB; F2 above the limit from 70 dB; G2 above the limit from 68.5 dB; F2 above the limit from 70 dB; G2 above the limit from 68.5 dB; C3 above the limit from 66 dB. Intervals and chords

A combination-tone bass needs a forte

A scale in just thirds draws a diatonic bass line through its difference tones, and in sixths the cubic product draws one. Given the two published level laws, with their constants swept, the thirds' bass is not heard at all below primaries of about 66 dB and is heard whole only from 71 to 81. The cubic products are a different kind of object: the primaries mask them decibel for decibel as they rise, so no dynamic changes whether they are heard. Most of the thirds' inner line never is, and the sixths' bass needs a forte and a gentle law.

A string that decays twice opens its count at once and shuts it early. The count of separable beats a mistuned octave on A3 delivers after both notes are struck at 80 decibels, with each filter read at the level inside it, under one exponential decay of 12 seconds and under two stages — a prompt sound of 1.5 seconds carrying all but the last 20 decibels, and an aftersound of 12 seconds. One exponential: open from 1.00 s to 3.30 s, 11.8 beats. Two stages: open from 0.15 s to 1.77 s, 8.8 beats — and the single exponential struck 20 decibels softer closes at 1.77 s. Pitch and tuning

A string that decays twice is counted early

A mistuned octave's beats were found countable only between a twelfth and a quarter of a note's life, on a note decaying once. A piano string decays twice, a fast prompt sound over a slow aftersound, and the prediction was that this would open the count sooner and close it later. It opens sooner — at a seventh of a second rather than a second — and closes exactly where a single decay struck twenty decibels softer closes, so at 80 dB it holds 8.8 beats instead of 11.8. The count now rises with the strike to 90 dB, and a tuner who strikes hard is right.

How long a doubled pizzicato keeps its note, seat by seat, in two rooms. How long a doubled violin pizzicato on 392 hertz keeps its note against the metres from the players, the pluck starting 12 dB up and decaying over 1 s with a loss exponent of 1. a concert hall, a flue pipe: 1 → 86 ms, 1.5 → 123 ms, 2 → 226 ms, 3 → 359 ms, 5 → 445 ms, 7 → 481 ms, 10 → 506 ms, 15 → 522 ms, 20 → 528 ms, 30 → 532 ms; a concert hall, an oboe: 1 → 52 ms, 1.5 → 55 ms, 2 → 59 ms, 3 → 77 ms, 5 → 149 ms, 7 → 195 ms, 10 → 224 ms, 15 → 242 ms, 20 → 248 ms, 30 → 254 ms; a concert hall, a clarinet: 1 → 44 ms, 1.5 → 45 ms, 2 → 45 ms, 3 → 47 ms, 5 → 53 ms, 7 → 65 ms, 10 → 86 ms, 15 → 105 ms, 20 → 112 ms, 30 → 118 ms; a large stone church, a flue pipe: 1 → 440 ms, 1.5 → 578 ms, 2 → 651 ms, 3 → 728 ms, 5 → 784 ms, 7 → 803 ms, 10 → 814 ms, 15 → 820 ms, 20 → 822 ms, 30 → 824 ms; a large stone church, an oboe: 1 → 56 ms, 1.5 → 65 ms, 2 → 89 ms, 3 → 207 ms, 5 → 281 ms, 7 → 303 ms, 10 → 315 ms, 15 → 322 ms, 20 → 325 ms, 30 → 326 ms; a large stone church, a clarinet: 1 → 44 ms, 1.5 → 43 ms, 2 → 43 ms, 3 → 44 ms, 5 → 53 ms, 7 → 67 ms, 10 → 80 ms, 15 → 89 ms, 20 → 92 ms, 30 → 94 ms. The mid-band critical distance is 5.3 m in a concert hall and 2.3 m in a large stone church. In none of the 60 cases does the note return to the pluck once it has left. Timbre and acoustics

A room keeps a pizzicato from giving its note away

Doubled by a flute, a one-second pizzicato loses its note in 70 milliseconds dry, because its upper partials go first. The question left open was whether a room, whose reverberation keeps those partials alive, gives the note back afterwards. It does not give it back. It stops the note going: ten metres into a concert hall the pluck keeps it for 506 milliseconds, in a stone church for 814, and the room's own uneven decay takes back between a quarter and two fifths of that. In a room the loss law that decided everything dry matters a tenth as much, because the room's decay has become the clock.

A bow shows the loss law a blow conceals. The same string on 130.8 hertz under three loss laws, drawn twice each: struck, and held by a continuous drive. The pale marks are the spectrum a blow produces, and they are identical in all three rows — a strike is the source spectrum and has no loss in it yet, which is why both endpoints of a struck note were found to carry nothing about the law. The solid marks are where each partial settles when a drive balances its own loss, at drive over loss, so the steady spectrum rolls off as the source's roll-off plus the exponent. At an exponent of 0.5 the held spectrum's centroid sits at 167 hertz, 5.7 semitones under the strike's 232; At an exponent of 1 the held spectrum's centroid sits at 144 hertz, 8.2 semitones under the strike's 232; At an exponent of 2 the held spectrum's centroid sits at 133 hertz, 9.6 semitones under the strike's 232. The quantity that is invisible at both ends of a struck note is the slope of a bowed one, for as long as the bow moves. Timbre and acoustics

A bow holds the number a blow hides

Two struck notes with different loss laws are identical at the strike and identical at the end, which is why separating them at all meant looking in the middle. Drive the same two strings continuously and the loss law stops being a rate and becomes a slope: each partial settles at its drive over its own loss, so the exponent adds to the source's roll-off and sits in the spectrum for as long as the bow moves. It is 18.7 decibels of separation available from the first instant, against 41.3 that a blow delivers after four tenths of a second and then takes away.

Every product of every pair of partials is a harmonic of one absent note. A 4 : 5 interval on 261.6 and 327.0 hertz, just, each note carrying 6 partials at one over n, with the fundamental at 80 decibels. Every pair of partials makes its own difference tone, and there are 19 distinct frequencies among them. All of them are exact multiples of 65.41 hertz — the note neither instrument is playing — because partial j of a p·f₀ note and partial k of a q·f₀ note differ by (kq − jp)·f₀ whatever j and k are. The heavy line is what each product has to clear — the threshold of hearing at its own frequency, or the masked threshold the two primaries cast there, whichever is higher. 11 of the 19 get through. Intervals and chords

Three harmonics of the bass arrive before the bass

Every product priced until now was between two pure tones, and nothing that plays thirds is pure. Give each note a spectrum and the ear receives every pair of partials — and for a just interval p:q every one of their products is an exact multiple of the same absent fundamental. That crowd lands where the threshold of hearing is tens of decibels cheaper, so it names the bass at 71 decibels where the component at the bass's own frequency needs 74, and at 75 against 85 an octave lower. Tempered, the crowd still forms and names a note seventy cents flat.

The ghost bass drops when the passage gets louder. The note the whole crowd of products names, as a multiple of the fundamental the interval implies, against how loudly the interval is played. a major third: 2.9999999999999996 times the fundamental below 70 decibels and 1 times above it, a drop of 19 semitones; a minor third: 4 times the fundamental below 62 decibels and 2 times above it, a drop of 12 semitones; a fourth: 2 times the fundamental below 72 decibels and 1 times above it, a drop of 12 semitones. Softly, only the cubic products clear their thresholds, and they are an exact series on (2p − q) times the fundamental with no gaps in it. Loudly, the difference tones fill in the low harmonics, no template on the higher note can explain them, and the fit falls. Nothing about the interval has changed; the listener is simply being given a different subset of the same harmonic series. Intervals and chords

The ghost bass drops a twelfth at a forte

Both crowds arrive at once and every member of both is a multiple of the same absent fundamental, so a listener is never given a choice between them — only a different subset of one harmonic series at every dynamic. Softly, the subset is an exact gapless series on three times the fundamental. Loudly, the difference tones fill in the low harmonics and no template on the higher note survives them. Between 62 and 72 decibels, depending on the interval, the note the crowd names falls by an octave or a twelfth, and the two qualities of third cross at different levels.

A displaced map is displaced by the same amount everywhere. How far a listener's heard direction is displaced, in units of the smallest angular change they could detect at that azimuth, for four constant offsets added to every interaural delay. Each curve is flat. an offset of 5 microseconds is worth 0.33 just-noticeable steps at every azimuth; an offset of 10 microseconds is worth 0.67 just-noticeable steps at every azimuth, and past 88° hands the listener a delay their own head cannot produce; an offset of 20 microseconds is worth 1.33 just-noticeable steps at every azimuth, and past 86° hands the listener a delay their own head cannot produce; an offset of 40 microseconds is worth 2.67 just-noticeable steps at every azimuth, and past 82° hands the listener a delay their own head cannot produce. The reason is exact: differentiating Woodworth's curve gives a slope proportional to (1 + cos θ), so the angular displacement a fixed offset produces carries a factor of 1/(1 + cos θ) — and so does the smallest detectable angle, so the ratio has no azimuth in it. That is the opposite of a wrong head radius, whose displacement is zero on the median plane and grows toward the side. Perception and the listener

The error that moves straight ahead

The essay before this one found that a listener whose internal head is the wrong size makes no error at all on the median plane, and has to look hard to the side to catch it. Every head drawn here has its ears at equal radii, which makes the delay curve odd and every error a factor — and a factor cannot move a zero. Real heads are not symmetric. A constant offset of twenty microseconds displaces a listener's straight ahead by two and a quarter degrees, and it displaces every other direction by the same number of just-noticeable steps, exactly.

Counted in beats, the abandoned chord comes back sooner and by steps. The piece in standard tuning with its G major, open chord given more and more of the piece, searched twice: once minimising the mean departure from just in cents, once minimising the mean beat rate. The vertical axis is how far that chord sits from just at each optimum. Counted in cents it stays at 13.29 cents until it holds exactly 44 beats — the same as the rest of the piece — and then drops to zero in one step. Counted in beats it drops at 30.4 beats to 4.96, then at 47.4 beats to 3.91, then at 73.1 beats to 0.00 cents. At the piece's own 8 beats the two counts agree exactly: the corner does not move, and what the beats change is where the steps are. Pitch and tuning

Counting beats moves the price of a chord, not the tuning

The search that found a guitar piece's best tuning counted every cent of error alike, and the obvious objection is that a cent of a third beats faster than a cent of an octave. Counted in beats instead, every piece gets exactly the same tuning back. What changes is how much of the piece the abandoned chord has to hold before it is rescued — thirty beats instead of forty-four, arriving in three steps instead of one.

A louder final chord stands higher and still stands for a fraction of a second. How far a final chord stands above the listener's running impression at the instant it is released, against how long it lasts. The chord is a struck six-note tonic; "a step louder" is the hammer velocity doubled, which raises its loudness 7.12 phons above the tutti's. At the tutti's level, after 1.2 s of silence, the impression is 8.09 dB below the chord and the chord stands highest, 5.14 dB, at 54 ms; a step louder, straight out of the tutti, the impression is 7.12 dB below the chord and the chord stands highest, 4.51 dB, at 36 ms; a step louder, after 1.2 s of silence, the impression is 15.21 dB below the chord and the chord stands highest, 9.46 dB, at 38 ms. Every curve returns to zero by half a second: the impression climbs to whatever level the chord is played at, so a louder mark is not a stand that lasts but a deeper fall to climb out of, and a silence and a mark add as depths. Form and structure

A louder final chord is a deeper silence and a brighter sound

A final chord marked a step louder than the passage was supposed to stand above a listener's running impression for as long as it sounded, since the impression can climb no higher than the chord. It climbs exactly that high, and the stand closes in half a second as it always did. What a louder mark actually buys is depth — about seven phons, the same depth a second of silence buys — and a spectrum whose balance point sits most of a whole tone higher, which, unlike the stand, lasts for the whole chord.

With 4 of six required at the end, the best schedule hands parts over. Six parts entering, leaving and re-entering a five-chord passage over four sounding players, in the walk through all 64 sets of sounding parts that makes the least audible entrance as audible as possible, with no memory of the chord before and at least 4 of the six sounding at the last chord. I, open: clarinet on G4 enters at 0.6 dB; vi, close below: oboe on E4 enters at 0.3 dB, and clarinet leaves; IV, close above: brass on E3 enters at -0.7 dB, voice on C5 enters at 2.2 dB, and oboe leaves; V, bracketing: violin on G5 enters at 7.2 dB; I, hollow: flue pipe on C6 enters at 4.0 dB. The least audible entrance is -0.72 dB against -2.65 for the best schedule in which nobody leaves; the walk has 2 exits and 6 entrances. Form and structure

An exit is worth nothing until the tutti is given up

Six parts entering a five-chord passage have a best schedule when each enters once and stays, and letting parts leave and come back was supposed to improve it. Searched over every set of sounding parts at every chord, it improves it by exactly nothing, with or without the chord before still masking — as long as all six must be playing at the end. Let one part be missing from the final chord and the weakest entrance gains 1.4 decibels; let two be missing and it gains 1.9, by a relay in which the parts with least room come in, are heard for one chord, and give way.

Put back beside its notes, the crowd names the bass at every dynamic. A just major third on complex tones, drawn on one axis of harmonic numbers of the fundamental its ratio implies: the partials of the two played notes, and the products of those partials that clear threshold, at 55 and 80 dB. At 55 dB the products alone name 3 times the fundamental with 0 empty slots; the products and the notes together name 1 times it with 4 empty slots, and the notes alone name it with 9. At 80 dB the products alone name 1 times the fundamental with 0 empty slots; the products and the notes together name 1 times it with 0 empty slots, and the notes alone name it with 9. The soft reading on a higher note exists only when the loud notes are set aside. Taken together, the products do not decide which fundamental is named; they decide how many holes its template has. Intervals and chords

The played notes already name the ghost bass

The products of a just third's partials, fitted on their own, name a note a twelfth above the bass when the interval is soft and drop to the bass when it is loud. Put the two played notes back beside them and the drop disappears: the notes and their products name the bass at every dynamic, because the notes' own partials are harmonics of it already. What the dynamic changes is not which note is implied but how complete its harmonic series is — nine holes from the notes alone, four when soft, none when loud.

A staccato is the direct sound's, and the room takes it within a fifth of the critical distance. A note on 130.8 Hz held 0.4 s and damped, in a room of 2 s reverberation, heard at distances from 0.02 to 5 times the critical distance: how long after the release the note takes to fall 10 dB and 20 dB. To fall 10 dB: 24 ms at the source, 333 ms far away; 0.02: 24 ms, 0.05: 25 ms, 0.1: 25 ms, 0.15: 26 ms, 0.2: 28 ms, 0.3: 35 ms, 0.5: 103 ms, 0.75: 187 ms, 1: 234 ms, 1.5: 281 ms, 2: 302 ms, 3: 318 ms, 5: 328 ms; doubled by 0.38 of the critical distance. To fall 20 dB: 49 ms at the source, 667 ms far away; 0.02: 49 ms, 0.05: 51 ms, 0.1: 60 ms, 0.15: 117 ms, 0.2: 198 ms, 0.3: 308 ms, 0.5: 436 ms, 0.75: 520 ms, 1: 568 ms, 1.5: 614 ms, 2: 635 ms, 3: 652 ms, 5: 661 ms; doubled by 0.14 of the critical distance. Where the direct sound and the room are equal, the damper's work is already hidden: the room's copy is only 20 dB below the direct sound at a tenth of the critical distance, and a 20 dB fall reaches it there. Timbre and acoustics

Only the player hears a staccato end

A damper stops a string in a seventh of a second, and in a hall the room goes on for two. A listener hears both, mixed in proportion to how close they sit, and the question was at what distance the short part stops mattering. The answer is closer than any seat. A damped note's twenty-decibel fall has doubled in length by a seventh of a hall's critical distance — 77 centimetres in a two-second concert hall — and by a quarter of it in a jazz club. The end of a staccato is something the pianist hears and the front row does not.

The arch belongs to hearing, and the spacing only moves it. The share of a close major triad's twenty-four components that stand above what the rest of the chord masks, at 70 dB, with the root from C1 to C7, for three spectra given the same amplitude law and different frequencies: the harmonic series, a founder's bell, and a stiff string with B = 0.01. harmonic series: 0.04 at C1, peaking at 0.79 on E3, 0.42 at C7; a founder's bell: 0.04 at C1, peaking at 0.75 on C4, 0.38 at C7; a stiff string: 0.04 at C1, peaking at 0.71 on E3, 0.46 at C7. Only one of the three is a harmonic series, and all three rise out of the bass, peak in the middle of the compass and fall in the treble. Perception and the listener

The arch belongs to hearing, not to the series

A chord delivers most of its partials in the middle of the compass and loses them in the bass and the treble, and every spectrum that showed that arch was built on whole multiples of a fundamental. Give the same amplitudes to a bell's eight modes and to a stiff string's stretched partials and the arch is still there, peaking within a major third of where the harmonic series peaks. What the spacing changes is the detail: a bell crowds its tierce and quint into a quarter of a critical band in the bass and loses them, and a stiff string's stretch buys the bass back.

Counted over what arrives, the balanced bass is not the roughest register. The mean roughness of the I – vi – IV – V – I arrivals with each chord played as loud as the written register's, relative to the written register, counted over every partial and over the partials that stand above what the rest of the chord masks. Every partial: 70 −2 octaves, 7.48 −1 octave, 1.00 as written, 0.20 +1 octave. Delivered partials only: 6e-9 −2 octaves, 2.83 −1 octave, 1.00 as written, 0.19 +1 octave. Over every partial the lowest register is 343 times rougher than the highest; over what arrives it is the smoothest of the four, and the roughest is −1 octave, 2.8 times the written register. Perception and the listener

A bass chord low enough to balance has already hidden its tenor

Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.

All themes