Field

Perception and the listener

Loudness is not amplitude and a beat is not in the signal. What the ear adds, hides, merges and supplies — measured, with the numbers every other field has been assuming.
Every point on one curve sounds equally loud. The equal-loudness contours of ISO 226:2003, evaluated from the standard's own parameters. The lowest curve is the threshold of hearing. Because the curves are not parallel — they crowd together in the bass and spread apart in the middle — the same change in decibels is a different change in loudness at every frequency, and a spectrum that was balanced at one level is not balanced at another.

The quietest thing audible, and why the volume knob is a tone control

A decibel is a fact about air. A phon is a fact about a listener, and the two do not line up — the map between them bends with frequency, and it bends differently at every level. One consequence is that turning a piece of music down does not turn all of it down equally, and the amount by which it does not is a number.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward.

One sound hides another, and it hides upward

A tone can be made completely inaudible by a second tone that is nowhere near it in frequency, and the region it disappears into is lopsided. Masking spreads up the spectrum and barely down it, and the reach grows with the masker's level — so which line in an arrangement vanishes is a prediction, not a matter of taste.

The threshold before, during and after a burst. The level a brief probe needs in order to be heard, plotted against when it happens relative to a 70 dB burst that occupies the shaded band. To the right is forward masking, which decays over about 200 milliseconds. To the left is backward masking: the threshold is raised for a probe that has already finished before the masker begins.

A sound hides what came before it

Masking does not stop when the masker does. A loud sound raises the threshold for about a fifth of a second after it ends, which is unremarkable, and for several milliseconds before it begins, which is not. The auditory present is a window rather than an instant, and inside the window the order of events is not the order they arrived.

A 7-semitone sequence at 120 ms a tone. Tones drawn as pitch against time, one bar per tone. The events are the same in both readings of this pattern; what changes is whether a listener assigns them to one line that leaps back and forth or to two lines that each stay put. Nothing in the drawing decides which, and nothing in the sound does either.

The ear builds objects, and sometimes offers a choice

What arrives at an ear is one pressure signal. What a listener gets is a set of separate things — a violin, a voice, a car outside. The assignment is a construction, and the clearest evidence is that it can be flipped by changing nothing but the speed: one sequence of tones is a single line when slow and two lines when fast, with a wide region in between where the listener may choose.

A source 45° off centre, and the path difference it makes. A head from above with a source to one side. The near ear is reached first; the far ear's path runs round the head, and the difference between the two is 13.1 centimetres, which at 343 metres a second is 381 microseconds. That number, and the level difference the head's shadow produces, are the whole of what the ear has to work with.

Two ears, and the whole of the difference is 655 microseconds

Direction is computed from two numbers — when a sound reaches each ear and how loud it is at each — and which of the two is usable is decided by the wavelength against the width of a head. The changeover frequency is not a design choice. It falls out of 343 metres a second and 17.5 centimetres, and it is why the mechanism of hearing where something is changes halfway up the piano.

How late a reflection has to be before it is an echo. What a single reflection does to the sound it follows, against its delay, on a logarithmic axis. Under a millisecond the two combine into one image that is pulled towards the earlier source. From there out to a few tens of milliseconds the reflection is not heard as a separate event at all and does not move the image — it only changes the timbre. Past the echo threshold it becomes a second sound, and the threshold is five times later for speech than for a click.

The first wavefront wins

A room sends a hundred copies of every note to a listener from a hundred directions, and the listener hears one note in one place. The mechanism that does it is brutal and simple: for the first few tens of milliseconds after a sound arrives, everything that follows is denied a vote on where it came from — even when it is louder than the original.

The smallest audible difference, and what has to clear it. The difference limen for frequency, converted from Wier, Jesteadt and Green's 1977 fit into cents, against the intervals and commas the rest of these essays argue about. Anything drawn below the curve is a quantity nobody can hear as a change of pitch; anything well above it is a quantity a listener can be asked about. The limen is for pure tones, successive, with trained listeners — the most favourable case there is, and therefore the right one to test a claim against.

How small a difference is audible

Every essay here about tuning has assumed a listener who can hear the difference between two systems. The assumption has a number: about five cents in the middle of the range. It clears the two commas fourfold and it does not clear the schisma at all, which sorts the whole subject of tuning into the part that is about music and the part that is about arithmetic.

Three octaves, and none of them is 2:1. How far above an exact doubling the upper note of an octave is set, against frequency. The listener's octave is measured with pure tones, which have no partials to beat against each other, so nothing about a stiff string can account for it. The piano's stretch is a different quantity with a different cause, and the two are drawn together only so that the difference is visible.

The octave that is not two to one

The octave is the one interval nobody argues about: two to one, exact, in every tradition that has one. Asked to set an octave by ear, listeners set it wide — and they do it with pure tones, which have no partials to beat against each other. Whatever is stretching the octave, it is not the stiffness of a piano string.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.

The chord is still major, and that is why temperament works

A major third can be seventeen cents wrong and still be a major third. That tolerance is not a failure of hearing — it is the reason the whole subject of tuning is a discussion rather than a catastrophe. Every temperament ever proposed moves intervals around inside their categories, and the one thing none of them may do is push one across a boundary.

The probe-tone profile, major key. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.

Counting produced the hierarchy

Ask listeners how well each of the twelve notes fits after a passage in C major and the answers are not a smooth gradient. They fall into four groups with no overlap at all: the tonic, then the rest of the tonic triad, then the rest of the scale, then everything else — categories the subject had names for centuries before anybody ran the experiment.

Two models of consonance, and where they disagree. Every chord scored twice: horizontally by summed Plomp–Levelt roughness, vertically by the largest integer needed to write it as members of one harmonic series. Both are supposed to be measuring consonance and both are computed here from the chord itself. They correlate, but not tightly enough to be the same claim, and the chords furthest from the diagonal are the ones any experiment has to be run on.

Consonance is half learned, and this is the half

This site's founding claim is that consonance is small whole numbers. Two computable models say so and they disagree about which chords — which is already awkward. The cross-cultural evidence is worse: listeners with little exposure to Western music discriminate roughness exactly as anyone does, match octaves exactly as anyone does, and rate consonant and dissonant chords as equally pleasant.

Direct and reverberant sound in a shoebox concert hall. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 5.5 metres in a room of 18700 cubic metres with a 2-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument.

How far away the room takes over

Direct sound falls six decibels every time the distance doubles and the reverberant field does not fall at all, so the two cross once. In a concert hall the crossing is at about five and a half metres, which is nearer than nearly every seat — so almost everybody in almost every hall is hearing the building more than the players, and the number that says so is built from two quantities already computed — a room's reverberation time and an instrument's directivity.

The ranking is settled either side of one narrow band. Remembered repetition — each bar's best match to an earlier bar, discounted by exp(−Δt/τ) with Δt in seconds — for 6 schemes at 108 beats a minute, against the decay constant τ on a logarithmic axis. The order of the schemes changes only between 8 and 13 seconds; outside that band it is fixed, so an estimate of τ wrong by any amount that stays outside it leaves the ranking alone.

A return has to be remembered

A stripe four bars off the diagonal and a stripe twenty-four bars off it are the same ink and are not the same experience. Convert the lag axis to seconds, discount every comparison by how long ago it was, and the ranking of these six schemes by how repetitive they are changes — and the decay constant and the tempo turn out to enter the arithmetic as one number rather than two.

Four voices, placed by the arithmetic. The rules in force are: no parallel octaves; no parallel fifths; no voice crossing; no gap over an octave above the tenor; the leading note is not doubled; the leading note resolves, outer voices; no augmented melodic interval. The four parts are drawn lowest to highest — bass, tenor, alto, soprano. I – vi – ii – V – I in C major, realised in four voices by the cheapest set of voicings obeying 7 rules, at 16 semitones of motion in all. Every gap between adjacent voices narrower than the fission boundary of 5.2 semitones is marked, and below that boundary two parts cannot be heard as two however hard a listener tries.

A voice is a stream, and the ear decides which

Seven earlier essays have assigned voices to notes. Whether a listener follows the assignment is a separate question with laboratory numbers attached, and the numbers are unkind to it: two parts closer than about five semitones cannot be heard as two at any speed, a third of the gaps in the cheapest four-part writing are inside that limit, and in a third of chord changes the ear's own rule for continuing a line does not recover the parts as written.

A 200-a-second click train, correlated with itself. The autocorrelation of a click train at 200 a second, smoothed by the ring of an auditory filter centred at 4000 Hz — an equivalent rectangular bandwidth of 456 Hz, so a ring of 2.2 ms. The regular train peaks at 5.0 ms, one period. With each click displaced by a standard deviation of 20 per cent of the period — 1.00 ms — the peak's contrast against the surrounding lags falls from 0.41 to 0.12. The average rate and the long-term spectrum are unchanged by the jitter; only the timing is.

A pitch with nothing to match

Filter a click train into a band where no partial is separable from its neighbours and it still has a pitch at its repetition rate. Displace each click by a fraction of a millisecond, leaving the average rate and the long-term spectrum exactly where they were, and the pitch goes. The mechanism is reading the timing — which bounds the account endorsed here from the start.

One pattern, four metres. The same 16-step onset pattern read under 4 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are downbeat on step 1 0, downbeat on step 2 -12, downbeat on step 3 -12, downbeat on step 4 0, so downbeat on step 1 and downbeat on step 4 tie and the model does not choose. Nothing about the sound differs between these readings; the bar line is supplied by the listener.

The beat that is never sounded

A listener who has heard four bars of a groove and then hears two bars with the downbeats taken out does not move the downbeat. This site's rule set does, every time, on every pattern tried — and the direction it moves in says exactly what kind of model would be needed instead.

The swing ratio, against the categories it passes through. The same swing curve read against the boundaries between duration categories rather than against notated values. A category's centre is a simple ratio — 1:1, 2:1, 3:1 — and the boundary between two of them is the midpoint, which is arithmetic. The curve crosses 2 of them: out of 2:1 and into 1:1 at 240 beats a minute, out of 3:1 and into 2:1 at 171 beats a minute. So the same notated figure is, by the categorical criterion, a different rhythm at each end of an ordinary tempo range, and the notation says triplet feel throughout.

How late is a different note

A deviation of thirty milliseconds is expression and a deviation of two hundred is a wrong note, so there is an edge. The edges in time are arithmetic — the midpoints between the simple ratios — and the swing ratio crosses two of them as the tempo rises, at 171 and at 240 beats a minute, while the notation says triplet feel throughout.

The price of a tonic. Every note of Dorian is given the same duration except its tonic, which is lengthened; the horizontal axis is the share of the total that goes to it. The key-finder answers with the parent key until 25.0 per cent of the time is spent on the modal tonic, and with D minor above it. At the left-hand edge every note has equal weight, which is the pitch-class set itself — and with every weight identical the correlation is not merely low but undefined, because a flat histogram has no variance to correlate with anything.

What a tonic costs in seconds

The standard key-finding algorithm cannot be run on a pitch-class set at all — a flat histogram has no variance and the correlation is undefined. Give it durations and it answers with the parent key for all seven modes identically, and it takes between 15.8 and 30.0 per cent of the total time spent on one note before it names that note instead.

Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.

Three answers to how finely a pitch can be heard

Two notes one after the other are told apart at about four cents at A440. Whether a melodic interval is in tune is a judgement an order of magnitude coarser. And two notes held together are heard to beat at a third of a cent, because the question is answered by counting rather than by hearing pitch at all. Every equal division ever built sits between the coarsest and the finest.

How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is.

How much evidence a modulation needs

Run a key-finder bar by bar over a progression that moves to the dominant at bar six. With four bars of history the answer becomes the new key at bar seven and holds. With three bars it never gets there at all, and reports E minor and B minor on the way. The window decides the lag as much as the music does.

What a contour costs to remember. A melody of n notes over 8 degrees carries 3 bits a note. Its contour carries fewer, and fewer than the number of distinct contours suggests, because the contours are not equally likely: at 6 notes there are 243 of them but the entropy is 6.59 bits, an effective alphabet of 96. Each further note adds 1.28 bits of contour against three of melody, so the shape keeps a stable 37 per cent of what is there however long the tune.

The part of the tune that is kept

Contour survives transposition, retuning, a change of instrument and a doubling of every interval, and the usual explanation is that it is what a listener retains. That can be counted rather than assumed. A six-note melody over eight degrees carries eighteen bits; its contour carries 6.59 — not the 7.92 the number of distinct shapes suggests, because the shapes are wildly unequal — and the effective alphabet is ninety-six out of two hundred and forty-three. Each further note adds 1.28 bits of shape against three of melody, and at about nine notes a contour is specific enough to pick one tune out of a thousand.

The pitch moves and the repetition rate does not. Three partials around harmonic 10 of 200 hertz, shifted together by up to 200 hertz, with three curves. The flat line is the envelope repetition rate, which the shift cannot move at all. The rising line is the shift divided by the harmonic number — 200 hertz becoming 220.0 — which is what a harmonic template predicts and is what listeners report. The third curve is the next-best template, which overtakes the first partway along: the pitch is ambiguous, and it drops back rather than rising indefinitely.

The pitch that moves the wrong distance

Take three partials two hundred hertz apart and move every one of them up by forty. The spacing has not changed, so anything reading the pitch off how often the waveform repeats must give the same answer as before. The pitch moves to 204 — the shift divided by the harmonic number — which is what a harmonic template predicts and what listeners report. Push the shift to a hundred and a second reading overtakes the first, so there are two pitches and neither is the spacing. This is the measurement that closes the question, and it closes it by ruling out one mechanism rather than by choosing between the two that are left.

How long a note has to be before its pitch is worth arguing about. The smallest audible frequency difference at 440 Hz, against how long the note lasts. The flat line is the steady-tone difference limen of 4.0 cents that every tuning argument on this site rests on. The falling line is the bound a finite duration imposes on its own frequency, 1/2T in cents, which no listener can beat. They cross at 486 milliseconds: below that the note is the limit and above it the listener is. A tenth of a second gives 19.6 cents and a quarter gives 7.9, against the commas drawn across the figure.

How long a note has to be

Every difference limen quoted so far is for a tone that lasts as long as the listener needs, and no note in music does. A tone of duration T occupies a band about 1/2T wide whatever the ear does with it, so at 440 hertz the quoted five-cent limen is the right number only for notes longer than 486 milliseconds. A tenth of a second gives 19.6 cents, which does not clear the syntonic comma. Most of the tuning arguments in this collection are about a quantity that only exists in long notes, and the essays that made them said so about the listener and not about the note.

500 Hz in one ear, 504 in the other. Two tones 4 hertz apart, one to each ear. They never meet in the air, so neither eardrum sees any modulation at all and there is no acoustic beat to hear. What changes is the phase between the ears, which advances a whole cycle every 250 milliseconds — and the direction that phase implies sweeps with it, drawn here as azimuth against time. The sweep is clipped at the edges, because the implied delay leaves the range a head can produce. A head 17.5 cm across gives at most 656 microseconds, so the phase stops naming a direction above 762 Hz.

The beat that is not in the air

Every sound this site synthesises reaches both ears identically, and that is the assumption none of its figures ever varied. Put 500 hertz in one ear and 504 in the other and nothing sums anywhere: each eardrum sees a steady sinusoid with no modulation on it at all. A listener still hears a four-per-second beat, which means the arithmetic is being done behind the ears rather than in the room. And it stops working above about a kilohertz — not where phase locking gives out at five, but where a head 17.5 centimetres across stops being able to name a direction, which is 762 hertz.

3 tones, one power, and the interval between them. 3 tones of fixed total power, spread symmetrically about 440 hertz, drawn against the interval between neighbours. Piled on one pitch they are one sound of that power; separated by more than a critical band — 4.5 semitones here — they are 3 sounds whose loudnesses add, and the same power reaches 2.08 times the loudness at 5 semitones. The two lines are two models of the same rule and they disagree about how abrupt the change is, not about where it goes.

A chord is not as loud as its notes

The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.

Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 2.8 cents at ten to one and 3.6 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it.

The boundary that barely moves

Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.

How much earlier an accent is heard, by mechanism. An accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller.

Playing louder is playing earlier

An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.

Where a room stops sending the two ears the same sound. The correlation between the two ears' signals against frequency, for a seat 15 metres from the source in a 15,000 cubic metre room with a 2 second reverberation time. The pale curve is the diffuse field alone — sin(kd)/(kd) for an ear separation of 17.5 centimetres, which first crosses zero at 980 hertz. The heavy curve adds the direct sound, which is coherent and lifts the whole thing by an amount the direct-to-reverberant ratio sets. At 125 hertz the coherence is 0.98 and at 1000 it is 0.16.

Where the two ears stop agreeing

A room sends both ears versions of the same sound, alike at low frequencies and increasingly unlike at high ones. Where they stop resembling each other is 980 hertz, and it is set by the 17.5 centimetres between the ears rather than by anything about the room — which is within a quarter of a frequency found earlier for a completely different reason. One minus that correlation is spaciousness, and it is computable from a room's own reverberation.

Why a concert hall is narrow. The lateral energy fraction at the middle seat as the same hall is widened, everything else held. It peaks at 12 metres across at 0.235 and falls to 0.000 at 44. A wide hall's side walls are further away, so their reflections arrive later, weaker and — this is the part Sabine's model cannot say — from nearer the front, where the sideways weighting discounts them. The shoebox halls the orchestral repertoire was written for are all between about eighteen and twenty-five metres wide, and this is the arithmetic they are the answer to.

A room with directions in it

Every room until now has been a reservoir of energy that drains at a rate. That model has no directions in it at all, so it cannot say the one thing every published measure of spaciousness is about: how much of what arrives comes from the side. Mirror the source in six walls and every reflection acquires an angle and a time — and the answer to why a concert hall is narrow falls out at eighteen metres.

How much of each spectrum a listener can assemble into one note. Each partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there.

The spectrum that will not fuse

A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.

Every arrival at one seat, against the delay at which it would be an echo. The echogram: each reflection at its delay after the direct sound and its level relative to it, in a 22 by 45 by 15 metre hall with 18 per cent absorption. The line is the published echo threshold for speech — 40 milliseconds at equal level and about 3.0 more for each decibel of attenuation — so anything to the RIGHT of it is late enough and loud enough to be heard separately. The once-reflected rear wall arrives at 210 milliseconds, 25 decibels down, against a threshold of 114 — well past it. Nothing here stands clear enough of its neighbours to be heard as an echo, and 181 arrivals are fused with the direct sound instead.

An echo is prevented by the crowd around it

The echogram says when every reflection arrives and how loud it is; the published echo threshold says when a reflection that late and that quiet is heard separately. Put one against the other and the rear wall of every hall anybody builds is past the threshold — a 45-metre hall puts it 210 milliseconds late and 25 decibels down against a threshold of 114. It is not heard as an echo, and what saves it is not the geometry. It is everything else arriving at the same time.

Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together.

The partials that do not die together

The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.

Where the sound is, and how wide. An earlier essay sorted a hall's arrivals into echoes and everything else. Everything else is not nothing: a reflection too early to be heard as a separate event still moves the apparent source, widens it and colours it, and all three come out of the same list of times, levels and angles. At this seat the direct sound arrives from 26.6 degrees off the front and the image is heard 8.5 degrees left of it, pulled by the near side wall. The apparent width is 33 degrees, from a lateral energy fraction of 0.23. And the strongest early reflection arrives 0.6 milliseconds behind off 1× the floor, which is a comb filter with notches every 1608 hertz and 25 decibels deep. The trading ratio and the discount on late arrivals are stipulated rather than measured, so the degrees are ordinal: what the figure claims is the direction and the shape, not the number.

A position and a width

Sorting a hall's arrivals into echoes and everything else settled that fusion is not a yes or a no. Everything else is not nothing: a reflection too early to be heard separately still moves the apparent source, widens it and colours it. All three come out of the same list of times, levels and angles, and none of them needed a new input.

What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged.

The exchange rate nobody has

There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.

The hall, as two numbers a listener has. Every reflection at this seat, placed by the interaural delay it produces rather than by the direction it came from. Time runs down; the dot's size is its energy. The direct sound is at 0 microseconds and the reverberation spreads over the whole available range, with a root-mean-square width of 322 against a geometric maximum of 656. That is the position and width computed earlier, in the units a listener has instead of the vectors used until now. 10 of the 56 reflections arrive from behind and carry 9 per cent of the energy — and they are drawn where they are because the interaural delay of a reflection from 120 degrees is identical to one from 60.

The hall through a head

Every direction computed so far is a vector from a seat to an image source, and a listener has no vectors. Run the echogram through the head computed six essays ago and two things happen: the position and width become microseconds, and a third of the room disappears — because both cues fold at ninety degrees and a reflection from behind is identical to one in front.

Nothing at all until fifteen decibels, and then it depends on the tempo. The fraction of a line's partials that stay above threshold, over how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement — forward masking leaves a threshold at most ten decibels below the masker, and a note's own partials mask each other from the same components at full level. The boundary is between twelve and eighteen decibels, and beyond it the loss grows with the tempo: at 280 to the crotchet and 36 decibels of contrast, 26 per cent of the line's partials are gone. Fifteen decibels is about the gap between a forte and a piano.

An equal note cannot be masked

Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.

The cue that settles it. Every spectrum to hand, arbitrated by the earlier competition and then again with the onset cue added at equal weight. 3 of the 5 change their verdict, and all 3 change the same way — from splitting into two streams to staying as one: a piano string, a bell, a bar. Nothing changes the other way, because the onset cue on a struck source votes for fusion on every partial and can only ever push toward one stream. The bell is the case worth naming: its partials are wildly inharmonic and it is heard as one sound, which is a fact the harmonicity cue alone cannot produce.

The cue that settles it

Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.

A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance.

A page has two decibels

The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.

A turn of 0.99 degrees tells front from back. A source 45 degrees off centre and its mirror image 135 degrees off, which produce the same interaural delay and are therefore the same signal to a listener who does not move. As the head turns the two predictions separate: the front source's delay falls and the rear source's rises, because the fold at ninety degrees puts them on opposite branches of the same curve. They differ by the 15-microsecond threshold after 0.99 degrees of turn — which is exactly half the 1.97 degrees a source would have to move for the same listener to notice it moving, and it is half for a reason: a turn displaces the two hypotheses from each other by twice what it displaces either of them from where it started.

The turn is half the angle

A stationary head cannot tell a sound in front from the same sound behind, and an earlier essay said so at length. The turn that breaks the confusion is 0.84 degrees — exactly half the angle a source would have to move for the same listener to notice it moving, and half for a reason. In a hall the same turn does something else: the source swings at 8.9 microseconds a degree and the room swings at 2.7, so a listener who moves is separating the soloist from the reverberation as well as the front from the back.

The same geometry at four sizes of head. Woodworth's interaural delay against direction, for 4 head radii from 5.8 to 9.8 centimetres. The whole range runs from 431 microseconds for a newborn to 735 for a large adult, and it scales exactly with the radius because the delay is (r/c)(θ + sin θ) and r is a multiplier. The detection threshold does not scale with the listener, so the number of distinguishable delays across the whole range falls from 98 to 57: a smaller head has the same directions in front of it and a shorter ruler to measure them with.

A smaller head in the same hall

Ten earlier essays draw one head. Every parameter belonging to the room has been varied by some figure and the one belonging to the listener never has, and it is the only one whose change the detection threshold does not follow: a six-year-old in the same seat receives the same fifty-four reflections at the same instants and reads them onto an axis with seventy distinguishable positions instead of eighty-seven. The speed of sound, swept over every temperature a hall is ever at, changes nothing at all — and the reason it cannot is the reason head size can.

A clarinet's partials, each on its own resonance. The 8 partials of the clarinet's chalumeau D that ride an impedance peak, each building toward its steady amplitude as 1 − exp(−t/τ) with τ = Q/πf from that peak's own Q. The time constants run from 11.9 milliseconds to 64.5, so the partials do not arrive at different times — they all begin the instant the reed does — and what differs is how fast each approaches its final level. The horizontal bars are how far apart the first and last are at three criteria: 5.5 ms at 10 per cent, 36.4 ms at 50 per cent, 121.0 ms at 90 per cent. A twenty-millisecond asynchrony is the threshold for hearing a partial out of a note, and this note crosses it at 32 per cent of steady amplitude — so whether a blown note's onset cue is unanimous or divided is decided entirely by how far along a partial has to be before it counts as having started.

A blown note does not start late, it starts slowly

Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.

Eleven partials is one partial too many. What fraction of a spectrum the harmonicity census finds fused, against how many partials it is asked to census. At ten a perfect harmonic series fuses 10 of 10 and the fundamental it finds is the right one. At eleven it fuses 5 of 11 and the fundamental jumps to exactly 2.00 — the octave above. The cause is the cap the census carries for a reason established earlier: without it a bell fuses perfectly at a fundamental nobody could hear, so the search refuses any fundamental more than about ten harmonics below the top partial. At eleven partials the first thing that cap excludes is the series' own fundamental, and the census then takes the octave and calls every odd partial inharmonic. So the number of partials and the cap are the same number, and nothing had ever said so, because every earlier figure censuses ten.

Eleven partials is one too many

Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.

The census with the criterion moved under it. Every instrument's speaking time in milliseconds, against the fraction of the steady amplitude counted as speaking. The criterion is in the wind instruments alone: a resonance takes −ln(1−p)·Q/(πf) to reach a fraction p, so those lines rise across the whole picture, while a bow's capture and an exciter's contact contain no criterion at all and are flat. Every earlier figure sits at 0.9, where the wind instruments are the slowest things in the collection by a factor of 20.5. At 0.05 they are the fastest: a violin's G3 string is the slowest at 15.5 milliseconds and a trumpet takes 1.4. The three clusters cross at a criterion between 0.18 and 0.39, which is inside the range the perceptual measurements work in — their three named criteria are 15 decibels below peak, 6 decibels below peak, and ninety per cent — and the settling figures have only ever used the third of them.

Read at two different heights

Ten placements of these figures, one value: the settling criterion is nine tenths in every one of them, and nothing is measured behind it. It is a multiplicative constant only inside the mechanism that has it — a bow's capture and an exciter's contact contain no criterion at all — so moving it rescales one of three clusters against two that stand still. The most-quoted number here, a factor of sixty-nine between the instrument's account and the listener's, is 5.9 at the criterion the listener's own measurements use, and the ordering an earlier essay was written about does not exist below a fifth.

An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none.

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

Articulation is worth 2.6 phons, and nobody counts it. A passage of notes at 80 decibels, 2 to the beat, at 7 tempi and 3 articulations, scored against the same level held continuously. The variable is the fraction of each inter-onset interval that is sounding — 0.95 is a legato, 0.4 a staccato — and the vertical axis is what that costs the passage's running loudness in phons. Nothing here is anybody playing harder or softer. At 40 to the beat the span from legato to staccato is 1.47 phons; at 200 it is 2.61, because a staccato note there lasts 60 milliseconds and no longer reaches its own loudness either.

A staccato is a dynamic mark

Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.

The top voice arrives whole and the bottom one arrives as a sine. 4 parts sounding together, each of 8 partials, with every partial tested against the summed masked threshold of every component in the texture. A filled mark is a partial the listener receives and an open one is a partial the part would have had alone and does not have here. The bass at C3 keeps 1 of 8, the tenor at C4 keeps 2 of 8, the alto at E4 keeps 5 of 8, the soprano at G5 keeps 8 of 8, every one of them at 70 decibels. Every part is at the same level and the difference is entirely where each one sits: masking spreads upward, so the part at the top of the texture has nothing above it to be masked by and the part at the bottom has everything.

The listener is given the top voice, and the bass as a sine

Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.

The roughest chord on the page is at C1 and the roughest one heard is at A♭2. One close triad at 70 decibels through 17 registers, with its roughness computed twice: over every partial in the score, and over only the partials that stand above what the chord itself masks. The written curve rises all the way down and its maximum is the lowest register drawn, C1, which is the low-interval rule as it has always been computed here. The delivered curve turns over at A♭2 and falls to nothing below A♭1: a close triad down there is not rough, because it is not arriving as a chord — 1 of its 24 partials survives at C1 and there is almost nothing left to beat against anything.

A low chord stops being rough by stopping being a chord

Every count of audible partials until now is of a chord at middle C, and the three registers it did compare span C3 to C5 — a third of the range a chord is written in. Move the same triad down and the count collapses: 79 per cent of its partials arrive at E3 and 4 per cent at C1. So the roughest chord on the page is the lowest one and the roughest chord a listener receives is at G2, and where that maximum sits moves nearly two octaves with the dynamic.

One attack time, three shapes, 48 ms of disagreement. Three amplitude envelopes with the same 90-millisecond attack time, which is the only quantity the published tables report. a resonator from a step rises as one minus a decaying exponential; an excitation ramping rises in a straight line; a ramp through a resonator is a raised cosine. Each is normalised so that its own 10-to-90 per cent rise takes exactly 90 milliseconds, so all three are the same measurement. The horizontal rules are the three criteria a heard moment is read off. At 6 dB below peak the three shapes put the heard moment at 28.5, 56.4, 76.3 milliseconds — a spread of 48, on one attack time, from a property nothing in the table records.

An attack time is not an attack

Eight earlier essays read a heard moment off an envelope, and every one of them used the same envelope shape without saying so: the source table records one curve for all nine of its families, and the map's own arithmetic does not carry the parameter at all. A published attack time fixes a ten-to-ninety time and nothing else. Under the two other shapes the same measurement admits, every millisecond computed so far doubles — and the constructed passage called inaudible earlier becomes three times a listener's threshold.

A head model a few millimetres out reads every azimuth but the front. The azimuth a listener reports against the azimuth a source is at, for internal head radii from 8.22 to 9.28 centimetres against a true radius of 8.75. The delay a source produces is (r/c)(θ + sin θ) and the listener inverts it with the radius they believe they have, so their answer solves θ̂ + sin θ̂ = (r/r̂)(θ + sin θ). Every curve passes exactly through the origin: on the median plane there is no delay and therefore no error, whatever the head model is. The error grows with azimuth and is largest at the side. An internal head 5.3 millimetres too small runs out of azimuth at 82 degrees: beyond that the world is delivering a delay larger than any its owner's model can produce, and every source out there collapses onto the side.

Where a wrong head gives itself away

Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.

A 20-decibel crescendo is 27 phons on a bass note and 20 on a high one. The same change of level, from 60 to 80 decibels, converted to loudness at each register through ISO 226's equal-loudness contours rather than at one kilohertz. The heavy curve gives each note a string spectrum, so its partials are converted in their own bands and summed; the pale one is the fundamental alone. On the spectrum-aware curve the crescendo is worth 27.3 phons at C1 and 20.3 at C7. On the fundamental alone it is 77 at C1, which is not a finding but an artefact: a 60-decibel tone at 33 hertz sits 1.8 decibels above the threshold of hearing and is very nearly nothing. The honest correction is the smaller one, and it is still a difference of 7.1 phons across the compass for a mark written in the same ink.

A subito piano is four seconds longer in the bass

Every loudness figure with time in it converts level to loudness at one kilohertz, and the equal-loudness contours say that no other frequency works that way. Joining the two sorts the published numbers into those that were about the treble and those that were not. Three move a great deal — a twenty-decibel crescendo is worth 27 phons on a bass note and 20 on a high one, and the seven seconds a subito piano takes becomes eleven and a third. Three do not move at all, and the reason they do not is the same reason in every case.

Scored the way these figures score it, a clarinet is the worst of the six. One close triad at 70 decibels through 17 registers, drawn once for each of the 6 spectra to hand. The score is the share of every partial written, which is the quantity the register figure published. pure 100 per cent at best, string 79 per cent at best, clarinet 54 per cent at best, reed 71 per cent at best, bell 52 per cent at best, organ 72 per cent at best. A clarinet's four even partials are twenty-eight decibels below its odd ones and are inaudible beside their own neighbours before any chord is built, so counting them in the denominator makes the spectrum that survives its own masking best look like the one that survives it worst.

A clarinet keeps what a string loses

Every masker, probe, chord, line and texture until now is eight partials falling as 1/n, and it was not even an option a placement could pass. Sweeping the six spectra to hand says the clarinet is the worst of them — 54 per cent of itself at best against a string's 79 — and that answer is an artefact of the score. Counted against what each note keeps on its own, the clarinet keeps 100 per cent where the string keeps 79, because its components stand a twelfth apart rather than an octave. The missing parameter was the spectrum; the second missing parameter was the denominator.

A displaced map is displaced by the same amount everywhere. How far a listener's heard direction is displaced, in units of the smallest angular change they could detect at that azimuth, for four constant offsets added to every interaural delay. Each curve is flat. an offset of 5 microseconds is worth 0.33 just-noticeable steps at every azimuth; an offset of 10 microseconds is worth 0.67 just-noticeable steps at every azimuth, and past 88° hands the listener a delay their own head cannot produce; an offset of 20 microseconds is worth 1.33 just-noticeable steps at every azimuth, and past 86° hands the listener a delay their own head cannot produce; an offset of 40 microseconds is worth 2.67 just-noticeable steps at every azimuth, and past 82° hands the listener a delay their own head cannot produce. The reason is exact: differentiating Woodworth's curve gives a slope proportional to (1 + cos θ), so the angular displacement a fixed offset produces carries a factor of 1/(1 + cos θ) — and so does the smallest detectable angle, so the ratio has no azimuth in it. That is the opposite of a wrong head radius, whose displacement is zero on the median plane and grows toward the side.

The error that moves straight ahead

The essay before this one found that a listener whose internal head is the wrong size makes no error at all on the median plane, and has to look hard to the side to catch it. Every head drawn here has its ears at equal radii, which makes the delay curve odd and every error a factor — and a factor cannot move a zero. Real heads are not symmetric. A constant offset of twenty microseconds displaces a listener's straight ahead by two and a quarter degrees, and it displaces every other direction by the same number of just-noticeable steps, exactly.

The arch belongs to hearing, and the spacing only moves it. The share of a close major triad's twenty-four components that stand above what the rest of the chord masks, at 70 dB, with the root from C1 to C7, for three spectra given the same amplitude law and different frequencies: the harmonic series, a founder's bell, and a stiff string with B = 0.01. harmonic series: 0.04 at C1, peaking at 0.79 on E3, 0.42 at C7; a founder's bell: 0.04 at C1, peaking at 0.75 on C4, 0.38 at C7; a stiff string: 0.04 at C1, peaking at 0.71 on E3, 0.46 at C7. Only one of the three is a harmonic series, and all three rise out of the bass, peak in the middle of the compass and fall in the treble.

The arch belongs to hearing, not to the series

A chord delivers most of its partials in the middle of the compass and loses them in the bass and the treble, and every spectrum that showed that arch was built on whole multiples of a fundamental. Give the same amplitudes to a bell's eight modes and to a stiff string's stretched partials and the arch is still there, peaking within a major third of where the harmonic series peaks. What the spacing changes is the detail: a bell crowds its tierce and quint into a quarter of a critical band in the bass and loses them, and a stiff string's stretch buys the bass back.

Counted over what arrives, the balanced bass is not the roughest register. The mean roughness of the I – vi – IV – V – I arrivals with each chord played as loud as the written register's, relative to the written register, counted over every partial and over the partials that stand above what the rest of the chord masks. Every partial: 70 −2 octaves, 7.48 −1 octave, 1.00 as written, 0.20 +1 octave. Delivered partials only: 6e-9 −2 octaves, 2.83 −1 octave, 1.00 as written, 0.19 +1 octave. Over every partial the lowest register is 343 times rougher than the highest; over what arrives it is the smoothest of the four, and the roughest is −1 octave, 2.8 times the written register.

A bass chord low enough to balance has already hidden its tenor

Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.

All fields · All essays