The collection

Every essay — page 15

Page 15 of 21, continuing through the fields in the same order.

Pitch and tuning Intervals and chords Scales and modes Harmony and voice leading Rhythm and metre Timbre and acoustics Perception and the listener Instruments and their design Form and structure Series Objects Sounds Search

Perception and the listener

Loudness is not amplitude and a beat is not in the signal. What the ear adds, hides, merges and supplies — measured, with the numbers every other field has been assuming.

A 200-a-second click train, correlated with itself. The autocorrelation of a click train at 200 a second, smoothed by the ring of an auditory filter centred at 4000 Hz — an equivalent rectangular bandwidth of 456 Hz, so a ring of 2.2 ms. The regular train peaks at 5.0 ms, one period. With each click displaced by a standard deviation of 20 per cent of the period — 1.00 ms — the peak's contrast against the surrounding lags falls from 0.41 to 0.12. The average rate and the long-term spectrum are unchanged by the jitter; only the timing is.

A pitch with nothing to match

Filter a click train into a band where no partial is separable from its neighbours and it still has a pitch at its repetition rate. Displace each click by a fraction of a millisecond, leaving the average rate and the long-term spectrum exactly where they were, and the pitch goes. The mechanism is reading the timing — which bounds the account endorsed here from the start.

8 figures
One pattern, four metres. The same 16-step onset pattern read under 4 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are downbeat on step 1 0, downbeat on step 2 -12, downbeat on step 3 -12, downbeat on step 4 0, so downbeat on step 1 and downbeat on step 4 tie and the model does not choose. Nothing about the sound differs between these readings; the bar line is supplied by the listener.

The beat that is never sounded

A listener who has heard four bars of a groove and then hears two bars with the downbeats taken out does not move the downbeat. This site's rule set does, every time, on every pattern tried — and the direction it moves in says exactly what kind of model would be needed instead.

8 figures
The swing ratio, against the categories it passes through. The same swing curve read against the boundaries between duration categories rather than against notated values. A category's centre is a simple ratio — 1:1, 2:1, 3:1 — and the boundary between two of them is the midpoint, which is arithmetic. The curve crosses 2 of them: out of 2:1 and into 1:1 at 240 beats a minute, out of 3:1 and into 2:1 at 171 beats a minute. So the same notated figure is, by the categorical criterion, a different rhythm at each end of an ordinary tempo range, and the notation says triplet feel throughout.

How late is a different note

A deviation of thirty milliseconds is expression and a deviation of two hundred is a wrong note, so there is an edge. The edges in time are arithmetic — the midpoints between the simple ratios — and the swing ratio crosses two of them as the tempo rises, at 171 and at 240 beats a minute, while the notation says triplet feel throughout.

6 figures
The price of a tonic. Every note of Dorian is given the same duration except its tonic, which is lengthened; the horizontal axis is the share of the total that goes to it. The key-finder answers with the parent key until 25.0 per cent of the time is spent on the modal tonic, and with D minor above it. At the left-hand edge every note has equal weight, which is the pitch-class set itself — and with every weight identical the correlation is not merely low but undefined, because a flat histogram has no variance to correlate with anything.

What a tonic costs in seconds

The standard key-finding algorithm cannot be run on a pitch-class set at all — a flat histogram has no variance and the correlation is undefined. Give it durations and it answers with the parent key for all seven modes identically, and it takes between 15.8 and 30.0 per cent of the total time spent on one note before it names that note instead.

7 figures
Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.

Three answers to how finely a pitch can be heard

Two notes one after the other are told apart at about four cents at A440. Whether a melodic interval is in tune is a judgement an order of magnitude coarser. And two notes held together are heard to beat at a third of a cent, because the question is answered by counting rather than by hearing pitch at all. Every equal division ever built sits between the coarsest and the finest.

5 figures
How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is.

How much evidence a modulation needs

Run a key-finder bar by bar over a progression that moves to the dominant at bar six. With four bars of history the answer becomes the new key at bar seven and holds. With three bars it never gets there at all, and reports E minor and B minor on the way. The window decides the lag as much as the music does.

7 figures
What a contour costs to remember. A melody of n notes over 8 degrees carries 3 bits a note. Its contour carries fewer, and fewer than the number of distinct contours suggests, because the contours are not equally likely: at 6 notes there are 243 of them but the entropy is 6.59 bits, an effective alphabet of 96. Each further note adds 1.28 bits of contour against three of melody, so the shape keeps a stable 37 per cent of what is there however long the tune.

The part of the tune that is kept

Contour survives transposition, retuning, a change of instrument and a doubling of every interval, and the usual explanation is that it is what a listener retains. That can be counted rather than assumed. A six-note melody over eight degrees carries eighteen bits; its contour carries 6.59 — not the 7.92 the number of distinct shapes suggests, because the shapes are wildly unequal — and the effective alphabet is ninety-six out of two hundred and forty-three. Each further note adds 1.28 bits of shape against three of melody, and at about nine notes a contour is specific enough to pick one tune out of a thousand.

7 figures
The pitch moves and the repetition rate does not. Three partials around harmonic 10 of 200 hertz, shifted together by up to 200 hertz, with three curves. The flat line is the envelope repetition rate, which the shift cannot move at all. The rising line is the shift divided by the harmonic number — 200 hertz becoming 220.0 — which is what a harmonic template predicts and is what listeners report. The third curve is the next-best template, which overtakes the first partway along: the pitch is ambiguous, and it drops back rather than rising indefinitely.

The pitch that moves the wrong distance

Take three partials two hundred hertz apart and move every one of them up by forty. The spacing has not changed, so anything reading the pitch off how often the waveform repeats must give the same answer as before. The pitch moves to 204 — the shift divided by the harmonic number — which is what a harmonic template predicts and what listeners report. Push the shift to a hundred and a second reading overtakes the first, so there are two pitches and neither is the spacing. This is the measurement that closes the question, and it closes it by ruling out one mechanism rather than by choosing between the two that are left.

6 figures
How long a note has to be before its pitch is worth arguing about. The smallest audible frequency difference at 440 Hz, against how long the note lasts. The flat line is the steady-tone difference limen of 4.0 cents that every tuning argument on this site rests on. The falling line is the bound a finite duration imposes on its own frequency, 1/2T in cents, which no listener can beat. They cross at 486 milliseconds: below that the note is the limit and above it the listener is. A tenth of a second gives 19.6 cents and a quarter gives 7.9, against the commas drawn across the figure.

How long a note has to be

Every difference limen quoted so far is for a tone that lasts as long as the listener needs, and no note in music does. A tone of duration T occupies a band about 1/2T wide whatever the ear does with it, so at 440 hertz the quoted five-cent limen is the right number only for notes longer than 486 milliseconds. A tenth of a second gives 19.6 cents, which does not clear the syntonic comma. Most of the tuning arguments in this collection are about a quantity that only exists in long notes, and the essays that made them said so about the listener and not about the note.

8 figures
500 Hz in one ear, 504 in the other. Two tones 4 hertz apart, one to each ear. They never meet in the air, so neither eardrum sees any modulation at all and there is no acoustic beat to hear. What changes is the phase between the ears, which advances a whole cycle every 250 milliseconds — and the direction that phase implies sweeps with it, drawn here as azimuth against time. The sweep is clipped at the edges, because the implied delay leaves the range a head can produce. A head 17.5 cm across gives at most 656 microseconds, so the phase stops naming a direction above 762 Hz.

The beat that is not in the air

Every sound this site synthesises reaches both ears identically, and that is the assumption none of its figures ever varied. Put 500 hertz in one ear and 504 in the other and nothing sums anywhere: each eardrum sees a steady sinusoid with no modulation on it at all. A listener still hears a four-per-second beat, which means the arithmetic is being done behind the ears rather than in the room. And it stops working above about a kilohertz — not where phase locking gives out at five, but where a head 17.5 centimetres across stops being able to name a direction, which is 762 hertz.

7 figures
3 tones, one power, and the interval between them. 3 tones of fixed total power, spread symmetrically about 440 hertz, drawn against the interval between neighbours. Piled on one pitch they are one sound of that power; separated by more than a critical band — 4.5 semitones here — they are 3 sounds whose loudnesses add, and the same power reaches 2.08 times the loudness at 5 semitones. The two lines are two models of the same rule and they disagree about how abrupt the change is, not about where it goes.

A chord is not as loud as its notes

The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.

8 figures
Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 2.8 cents at ten to one and 3.6 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it.

The boundary that barely moves

Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.

7 figures
How much earlier an accent is heard, by mechanism. An accented note on an instrument with a 90 millisecond attack, drawn against how many decibels louder it is, with the three ways it can arrive early separated. A criterion tied to the note's own peak on an unchanging envelope gives exactly nothing. The same criterion on the shorter rise a harder-driven instrument has gives 5.3 milliseconds at 12 decibels. A criterion at a fixed level gives 23.0. Both together give 24.0, and the rise at that dynamic is 73 milliseconds rather than 90. The rise-shortening exponent is stipulated at 0.15 rather than measured, and the two upper curves would separate further if it were smaller.

Playing louder is playing earlier

An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.

8 figures
Where a room stops sending the two ears the same sound. The correlation between the two ears' signals against frequency, for a seat 15 metres from the source in a 15,000 cubic metre room with a 2 second reverberation time. The pale curve is the diffuse field alone — sin(kd)/(kd) for an ear separation of 17.5 centimetres, which first crosses zero at 980 hertz. The heavy curve adds the direct sound, which is coherent and lifts the whole thing by an amount the direct-to-reverberant ratio sets. At 125 hertz the coherence is 0.98 and at 1000 it is 0.16.

Where the two ears stop agreeing

A room sends both ears versions of the same sound, alike at low frequencies and increasingly unlike at high ones. Where they stop resembling each other is 980 hertz, and it is set by the 17.5 centimetres between the ears rather than by anything about the room — which is within a quarter of a frequency found earlier for a completely different reason. One minus that correlation is spaciousness, and it is computable from a room's own reverberation.

6 figures
Why a concert hall is narrow. The lateral energy fraction at the middle seat as the same hall is widened, everything else held. It peaks at 12 metres across at 0.235 and falls to 0.000 at 44. A wide hall's side walls are further away, so their reflections arrive later, weaker and — this is the part Sabine's model cannot say — from nearer the front, where the sideways weighting discounts them. The shoebox halls the orchestral repertoire was written for are all between about eighteen and twenty-five metres wide, and this is the arithmetic they are the answer to.

A room with directions in it

Every room until now has been a reservoir of energy that drains at a rate. That model has no directions in it at all, so it cannot say the one thing every published measure of spaciousness is about: how much of what arrives comes from the side. Mirror the source in six walls and every reflection acquires an angle and a time — and the answer to why a concert hall is narrow falls out at eighteen metres.

7 figures
How much of each spectrum a listener can assemble into one note. Each partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there.

The spectrum that will not fuse

A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.

6 figures
Every arrival at one seat, against the delay at which it would be an echo. The echogram: each reflection at its delay after the direct sound and its level relative to it, in a 22 by 45 by 15 metre hall with 18 per cent absorption. The line is the published echo threshold for speech — 40 milliseconds at equal level and about 3.0 more for each decibel of attenuation — so anything to the RIGHT of it is late enough and loud enough to be heard separately. The once-reflected rear wall arrives at 210 milliseconds, 25 decibels down, against a threshold of 114 — well past it. Nothing here stands clear enough of its neighbours to be heard as an echo, and 181 arrivals are fused with the direct sound instead.

An echo is prevented by the crowd around it

The echogram says when every reflection arrives and how loud it is; the published echo threshold says when a reflection that late and that quiet is heard separately. Put one against the other and the rear wall of every hall anybody builds is past the threshold — a 45-metre hall puts it 210 milliseconds late and 25 decibels down against a threshold of 114. It is not heard as an echo, and what saves it is not the geometry. It is everything else arriving at the same time.

6 figures
Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together.

The partials that do not die together

The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.

6 figures
Where the sound is, and how wide. An earlier essay sorted a hall's arrivals into echoes and everything else. Everything else is not nothing: a reflection too early to be heard as a separate event still moves the apparent source, widens it and colours it, and all three come out of the same list of times, levels and angles. At this seat the direct sound arrives from 26.6 degrees off the front and the image is heard 8.5 degrees left of it, pulled by the near side wall. The apparent width is 33 degrees, from a lateral energy fraction of 0.23. And the strongest early reflection arrives 0.6 milliseconds behind off 1× the floor, which is a comb filter with notches every 1608 hertz and 25 decibels deep. The trading ratio and the discount on late arrivals are stipulated rather than measured, so the degrees are ordinal: what the figure claims is the direction and the shape, not the number.

A position and a width

Sorting a hall's arrivals into echoes and everything else settled that fusion is not a yes or a no. Everything else is not nothing: a reflection too early to be heard separately still moves the apparent source, widens it and colours it. All three come out of the same list of times, levels and angles, and none of them needed a new input.

6 figures
What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged.

The exchange rate nobody has

There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.

7 figures
The hall, as two numbers a listener has. Every reflection at this seat, placed by the interaural delay it produces rather than by the direction it came from. Time runs down; the dot's size is its energy. The direct sound is at 0 microseconds and the reverberation spreads over the whole available range, with a root-mean-square width of 322 against a geometric maximum of 656. That is the position and width computed earlier, in the units a listener has instead of the vectors used until now. 10 of the 56 reflections arrive from behind and carry 9 per cent of the energy — and they are drawn where they are because the interaural delay of a reflection from 120 degrees is identical to one from 60.

The hall through a head

Every direction computed so far is a vector from a seat to an image source, and a listener has no vectors. Run the echogram through the head computed six essays ago and two things happen: the position and width become microseconds, and a third of the room disappears — because both cues fold at ninety degrees and a reflection from behind is identical to one in front.

7 figures
Nothing at all until fifteen decibels, and then it depends on the tempo. The fraction of a line's partials that stay above threshold, over how fast the line moves and how much louder everything before each note is. Darker is more lost. The whole left-hand side is white: at equal levels a note cannot be masked by its predecessor at any tempo, and that is a proof rather than a measurement — forward masking leaves a threshold at most ten decibels below the masker, and a note's own partials mask each other from the same components at full level. The boundary is between twelve and eighteen decibels, and beyond it the loss grows with the tempo: at 280 to the crotchet and 36 decibels of contrast, 26 per cent of the line's partials are gone. Fifteen decibels is about the gap between a forte and a piano.

An equal note cannot be masked

Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.

7 figures
The cue that settles it. Every spectrum to hand, arbitrated by the earlier competition and then again with the onset cue added at equal weight. 3 of the 5 change their verdict, and all 3 change the same way — from splitting into two streams to staying as one: a piano string, a bell, a bar. Nothing changes the other way, because the onset cue on a struck source votes for fusion on every partial and can only ever push toward one stream. The bell is the case worth naming: its partials are wildly inharmonic and it is heard as one sound, which is a fact the harmonicity cue alone cannot produce.

The cue that settles it

Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.

7 figures
A page has two decibels and a player has sixty. Across, parts added to a final chord one at a time, each at the same level; up, the loudness that results, on a logarithmic scale. Going from one part to eight moves the total by 1.8 decibels and does not move it monotonically — four parts are louder than five and than eight. The faint line is what a naive power sum would give: 9.0 decibels. The band down the right is the same chord played by people, from forty to a hundred decibels, which spans 62. So a texture that thins from eight parts to one is not a diminuendo. It is a change of colour at constant loudness, and everything the closure figures call a dynamic belongs to the performance.

A page has two decibels

The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.

7 figures