The collection

Every essay — page 4

Page 4 of 21, continuing through the fields in the same order.

Pitch and tuning Intervals and chords Scales and modes Harmony and voice leading Rhythm and metre Timbre and acoustics Perception and the listener Instruments and their design Form and structure Series Objects Sounds Search

Intervals and chords

Two notes at once, why some of them beat, and the geometry of the ones that do not.

Where to stop, and what stopping there gives. Each prefix of the harmonic series sounded as a chord of pure tones, with its roughness per pair and the pitch classes it contains. The smoothest prefix is the first 4, which is a bare fifth and an octave; roughness rises monotonically from there, so no roughness argument selects six. The prefixes that contain a major triad and nothing else are 5 and 6. One partial further and it is a dominant seventh; four further and there are five pitch classes. Six is the last stopping point that gives the answer the derivation wants, and nothing else in the arithmetic picks it out.

The series is not a chord

The major triad is derived from the first six partials of the harmonic series in every textbook that derives it from anything, and the derivation works: the first six contain a major triad and nothing else. Take the first four instead and there is no third at all; take seven and there is a dominant seventh with a flat seventh thirty-one cents below any keyboard note; take sixteen and there are eight pitch classes. Six is the last stopping point that gives the answer, roughness prefers three, four or ten depending on the register, and no criterion in the arithmetic selects six over any of them.

7 figures
The chain of fifths in quarter-comma meantone. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and one of them — G♯ to D♯, the 12th link, where the chain is forced to close — is the wolf, at 35.7 cents.

The same distance, under two names

Four hundred cents is a major third or a diminished fourth, and on a keyboard nothing in the sound distinguishes them. An earlier essay was about the boundary between two categories; this is about two categories at one acoustic value, and the surprise is where the ambiguity comes from. In quarter-comma meantone a major third is 386 cents and a diminished fourth is 427 — two names, two pitches, forty-one cents apart. Equal temperament collapsed them, and what a listener now supplies from context used to be in the sound.

7 figures
The series has three tops. Where the harmonic series stops, asked three ways, at four fundamentals. Consecutive partials stop being separately resolvable around partial 8 — that one depends on the fundamental, since a critical band is a fixed width in hertz and the spacing is not. They stop being a semitone apart at partial 17 at every fundamental, because the ratio (n+1)/n does not know what n is measured in. And they stop being distinguishable in pitch at all between partials 34 and 140. Every claim about how far up the series something happens is a claim about which of these three was meant.

The series has three tops

How far up the harmonic series can an ear go? The question has three answers and they are an order of magnitude apart. Consecutive partials stop being separately resolvable somewhere around the eighth, and where exactly depends on the fundamental. They stop being a semitone apart at the seventeenth, at every fundamental, because the ratio does not know what it is measured in. And they stop being distinguishable in pitch at all between the thirty-fourth and the hundred and fortieth. Every claim about where the series runs out is a claim about which of the three was meant.

6 figures
How many notes an octave can hold, asked twice. The share of trials on which a category is named correctly, against how many equal categories the octave is cut into, for a listener whose internal estimate carries 11 cents of noise — the logistic scale this site's identification figures already use, which is the thirty-cent transition the studies report. At 95 per cent accuracy the ceiling is 6 categories, and seven scores 94.9 per cent — on the line. The other ceiling is resolution: 151 to 356 difference limens fit in an octave depending on register, which is a factor of forty larger. Every system marked below sits between the two, and the marks separate: the number of degrees a mode uses clears the criterion, and the size of the gamut it chooses them from does not.

How many boxes an octave holds

An identification model of pitch is usually handed twelve categories, and nothing ever asked how many an octave can hold. There are two answers and they are a factor of forty apart. Resolution allows between 151 and 356 — a listener can tell that many pitches apart in a direct comparison. Naming one of them without a comparison is a different faculty and it runs out at six or seven, which is where every mode in every tradition compared here sits. Turkish theory names fifty-three commas to the octave and a makam uses seven of them, and the gap between those two numbers is the whole of the argument.

8 figures
What a spread costs a chord. a major triad of 3 notes lasting 600 ms each, with the onsets spread by up to 320 ms. The upper line is the share of each note's length during which every note is sounding; the lower is the chord's roughness weighted by that share, since roughness is a property of two partials sounding at the same time. At a spread of 30 ms — the asynchrony at which a mistimed partial stops belonging to its note — the chord is still 89 per cent simultaneous. It stops being simultaneous at all at 300 ms, which is where the last note arrives after the first has finished.

The chord that is not played at once

Every chord until now starts its notes at the same instant, and no figure ever set the asynchrony to anything else. A spread chord is not a defective simultaneity: at forty milliseconds a triad of half-second notes is still eighty-seven per cent simultaneous, so it carries almost all its roughness, and it stops being a chord at all only when the last note arrives after the first has finished. What none of this can explain is the one thing every keyboard player knows — that a chord is rolled upward. The masking asymmetry that ought to explain it is 2.4 decibels at a close voicing, which is not enough.

8 figures
How finely a fifth can be heard, against how fast it goes by. The smallest audible mistuning of a melodic fifth above 440 hertz, against how long each of its two notes lasts. The lower solid line is one note's own limen; the upper one is the interval's, larger because two independent errors add in quadrature. A tenth of a second gives 23.5 cents against the note's own 19.6, the floor for long notes is 5.4, and the syntonic comma is not cleared until each note lasts 111 milliseconds. The dotted line at 1.31 cents is the same interval heard as a simultaneity, where partials 3 and 2 coincide at 1320 hertz and one beat every 2 seconds can be counted there.

An interval is two errors

Two notes in succession are two pitch estimates, and what a listener judges is their difference — so a melodic interval is heard less finely than either of the notes in it. At a semiquaver the limen is nineteen cents, which lands inside a published range long quoted from the literature and never derived. Sounded together instead of one after the other, the same interval is judged eighteen times more finely.

7 figures
Which note a chord would rather have twice. Every complete four-part voicing of each chord inside the SATB ranges, grouped by which member sounds twice and scored for roughness — 480 voicings for a triad. The order for a major triad is root < fifth < third, which is the rule every part-writing treatise states. For a minor triad it is fifth < root < third, which is not. The numbers printed under each bar are the mean error, in cents, with which the four sounding notes fit a single harmonic series, and that measure separates the three far more sharply than roughness does.

The note that sounds twice

A triad has three notes and a four-part texture has four voices, so one note is doubled — and the voicing model used here leaves the choice free because the rules have an opinion about it. Asked properly, the arithmetic agrees with the treatises for the first time in nine essays: root, then fifth, then third. For a minor triad it does not agree, and for a symmetric chord it correctly has nothing to say.

7 figures
How much correlation it would take to matter. The limen of a 7-semitone interval at a note length of 0.25 seconds, against the correlation between the two notes' errors. The independent model at the left gives 9.44 cents. Halving that needs a correlation of 0.75; a fifth off it needs 0.31. The curve is √(1 − ρ) and nothing else, so the correlation required for a stated improvement is arithmetic — which turns the question from “does a key help?” into “by how much, and here is the number it must reach”.

How much an anchor would have to be worth

Two pitch errors added in quadrature assume an independence nobody measured — a listener inside a key hears a note as a scale degree, and a shared reference is exactly a correlated error. Turning the dial is not evidence. What is evidence is that the dial is not free: a shared error cancels out of a difference completely, so a listener's single-note limen and their interval limen give the two components with nothing left over, and halving the interval limen needs a correlation of exactly 0.75.

6 figures
Every modern bell puts the boundary at the same partial. For each brass instrument, the partial at which its bell stops turning the wave round — the frequency at which half the energy escapes at the mouth, divided by the tube's own fundamental. The five modern instruments land between partial 8.4 and partial 10.2 despite tube lengths differing by a factor of four, because a bell is scaled with its instrument. The natural trumpet of the baroque, whose bell is small and whose tube is long, puts it at partial 17.2.

The fourth top is the maker's

The series has three tops and all three are the ear's: where consecutive partials stop being resolved, where their spacing falls below a semitone, and where it falls below what the ear can hear at all. The fourth is how far up a player can go, and it is the only one somebody chose. Every modern brass bell puts the boundary at about the ninth partial across a family whose tubes differ by a factor of four — and the baroque natural trumpet, whose bell is small, puts it at the seventeenth, which is exactly the top of the clarino register.

6 figures
The beat rate between two sections, second by second. The 5th partial of the lower section against the 4th of the upper, over 256 pairs of voices each sweeping 100 cents 6 times a second with its own phase. The band is the tenth to the ninetieth percentile of the instantaneous rate and the line is the median. With no vibrato the whole thing would be one flat line at 8.7 hertz, which is what was computed earlier. With it, the pair is inside the beating band 20 per cent of the time and above it for the rest — so what a listener gets is neither a beat nor a roughness but an alternation between them at the vibrato rate.

Sixteen sweeps against sixteen

Every intonation figure about the voice treats a singer as a frequency. A singer is a frequency being swept a hundred cents wide six times a second, and two sections singing an interval are two hundred and fifty-six pairs of sweeps. The beat rate between the partials the interval brings together stops being a number and becomes a function of time — and the pair spends four fifths of its time above the rate at which beating is beating at all.

6 figures
What the notes in between do to the anchor the interval is measured against. How finely a 7-semitone interval can be judged when its two notes are separated by other notes rather than by silence, under the two published accounts. Confirming material restates the key and refreshes the shared reference, so the correlation climbs from 0.5 toward a ceiling and the limen falls to 4.81 cents. Overwriting material competes for the same memory, so the correlation decays to 0.04 and the limen rises to 9.28. By 8 notes the two accounts differ by 4.5 cents, which is 47 per cent of the limen with no anchor at all — and no experiment here distinguishes them.

The notes in between

Every figure until now is about two notes with nothing between them, and a melody is notes with other notes between them. Two published accounts of what the intervening material does predict opposite signs — one says the key is restated and the shared reference is refreshed, the other says each note competes for the same memory and it decays. By eight notes they differ by four and a half cents, which is nearly half the limen the interval would have with no anchor at all.

6 figures
A vibrato flattens the dissonance curve. Each interval twice: hollow is its roughness computed at the two notes' nominal frequencies, filled is the average of its roughness over a vibrato cycle of 50 cents at 6 hertz. They are not the same number, because roughness is a curved function of the frequency difference and the average of a curve is not the curve of the average. The largest effect is at octave, where the moving average is 19.3 times the still value — an interval sitting in a deep narrow minimum is smeared out of it. The smallest is at major seventh, where it is 0.95: an interval near a maximum is smeared out of that too. Vibrato pushes every interval toward the middle, and what it takes away from the consonances is much more than what it takes away from the dissonances.

A roughness with a rate of its own

Every roughness figure so far computes one number for a steady spectrum. Evaluate the same sum at every instant of a vibrato and there are three numbers instead — a mean, a depth and a rate — and the mean is not the roughness of the mean frequency. On an octave it is nineteen times it, because an octave sits in a deep narrow minimum and a vibrato smears it out of one.

6 figures
Which interval gives a tuner the deepest null. A tuner listening to an interval p:q is listening to the lower note's p-th partial against the upper note's q-th, and how deep the beat's trough goes is decided by those two amplitudes rather than by the interval. For a string spectrum, whose partials fall as one over n, the minor third pairs partial 6 against partial 5 at a ratio of 1.18 for a dip of 21.8 decibels; the major third pairs partial 5 against partial 4 at a ratio of 1.25 for a dip of 19.1 decibels; the fourth pairs partial 4 against partial 3 at a ratio of 1.32 for a dip of 17.2 decibels; the fifth pairs partial 3 against partial 2 at a ratio of 1.52 for a dip of 13.8 decibels; the major sixth pairs partial 5 against partial 3 at a ratio of 1.65 for a dip of 12.2 decibels; the minor sixth pairs partial 8 against partial 5 at a ratio of 1.67 for a dip of 12.0 decibels; the octave pairs partial 2 against partial 1 at a ratio of 2.00 for a dip of 9.5 decibels. The best is the minor third at 21.8 and the worst is the octave at 9.5, which is the reverse of the order a tuner is usually taught to trust: the deepest null in the list is on the interval whose coincidence sits highest in the spectrum, where adjacent partials are nearly equal in strength.

A beat has a depth, and six essays held it at one

Every beat figure so far adds two tones of equal amplitude, which is the single ratio at which the trough of a beat is a true null — and a null is what a tuner is actually listening for. Vary the ratio and the picture changes: at two to one the dip is nine and a half decibels, at ten to one it is under two, and the interval that gives the shallowest null of all is the octave.

6 figures
The same interval, started on each of the twelve. An interval of 7 semitones started on each pitch class of a major key, against how strongly the key specifies its two notes — the mean of the probe-tone profile at each. The interval account says the listener encodes a distance, so the key cannot enter and the prediction is a horizontal line at 5.4 cents. The degree account says the listener refers each note to the key, so its precision on a note falls as the key's specification of that note weakens; scaled to agree at the most stable start, it rises from 5.4 cents on C to 8.5 on E♭. Every earlier figure measures a quantity the second account says is not being formed at all.

The quantity a rival account says is not there

Two earlier essays measure how much two notes' errors are correlated through a shared anchor, and price what that correlation would be worth. There is a rival account in which a listener refers each note to a key and never forms the distance at all — under which the correlation is not small, it is a description of something that is not happening. The two accounts agree on almost everything and disagree on one manipulation, and the manipulation costs an afternoon.

5 figures
How far out of tune an interval has to be before its beat is usable. An earlier essay produced a modulation index for every interval on a stated timbre; whether a fluctuation of that index at that rate can be detected is a published function of both. Running one against the other turns a table of decibels into a window in cents. The bar is the mistuning over which the beat is both deep enough to notice and at a rate a tuner can use — not so slow that a beat takes half a minute to complete, not so fast that it has stopped being a beat. the major third gives the widest window, 1.3 to 60.0 cents, and the minor sixth the narrowest, 0.8 to 38.9. The ordering is the opposite of the dip's: the octave has the shallowest dip on a string spectrum and the widest usable window, because its coincidence sits at a low partial and a given mistuning therefore produces a slower beat. Depth and rate pull opposite ways, and it is the rate that decides.

The beat a tuner can actually use

There is a modulation index for every interval on every timbre, and the published threshold for detecting a fluctuation is a function of exactly that and its rate. Running one against the other turns a table of decibels into a window in cents — and reverses the ordering, because the interval with the shallowest dip has the widest window.

7 figures
The same roughness, before and after the window it has to be heard through. The instantaneous roughness of an interval under a vibrato, and the same quantity after a running average of 59 milliseconds — the time a dissonance has to last to be heard as one, which is 4 cycles of this interval's own 68-hertz fluctuation rather than a number chosen for the figure. The mean is identical to every digit, 0.1487 against 0.1487, because a running average cannot change an average — so the earlier Jensen factor of 1.0 survives the window untouched and its prediction that the window would shrink it is wrong. What the window destroys is the depth: 0.30 of the mean becomes 0.23, which is 77 per cent. The roughness a vibrato adds is heard; the fact that it is moving is mostly not.

The mean survives the window

A roughness that moves has a mean, a depth and a rate — all three of which a listener could only have through a temporal window. Applying the window already to hand settles which of the three survives, and the answer refutes the guess: a running average cannot change an average, so the octave's factor of nineteen stands and the movement is what goes.

6 figures
A mistuned 2:1 makes 8 beats, not one. Every pair of partials that coincides in a 2:1 mistuned by 6 cents, with the rate each one beats at. On a perfectly flexible string the rates are 1.5, 3.1, 4.6, 6.1 and so on — an exact harmonic series of the slowest, because partial 2k is exactly twice partial k. On a stiff one they are 1.3, 0.9, 2.7, 11.1, and the 8th member is at 122 hertz. The family grows as the cube of the member number rather than linearly, so 4 of the 8 are slow enough to be beats at all and the rest are roughness.

A beat is never one beat

Every beat counted until now has come from one pair of partials. A real spectrum has many, so a mistuned octave produces eight beats at once — and on a perfectly flexible string those eight are an exact harmonic series of the slowest. On a real piano string they are a cubic, half of them are not beats at all, and no single width of octave silences more than one. That is where the tuner's five octaves come from.

7 figures
Of 8 beats at A3, 3 can be attended to. Every member of the beat family a 2:1 mistuned by 6.0 cents makes on a real string at A3, placed by how separable its coincidence is from the partials beside it — in auditory filter widths, across — and by how fast it beats, up. The shaded region is the set a listener can receive: wider than one filter, and between 0.4 and 15 fluctuations a second. 4 of 8 members fall inside it, and once rates within a factor of 2 are counted as one modulation channel there are 3. The members that fail do so for two different reasons: the low ones sit under the roughness ceiling but their coincidences are buried in a filter that holds three partials, and the high ones are resolved and far too fast.

Three beats at most, and only in the middle of the keyboard

A mistuned octave on a real piano makes eight beats at once, and a listener attending to one of them is doing something that has a threshold. Two thresholds, in fact — a rate and a place — and once both are applied the eight become four at A3, one at A1 and one at A5. Every interval a tuner sets goes to zero at both ends of the compass and peaks at eight countable beats in the octave the bearing is laid in.

7 figures
How much of A4's pitch error a key could possibly remove. The largest correlation two notes' pitch errors can have at A4, against how long each note lasts. A note's error has two parts and only one of them is the listener's: the steady-tone limen of 4.04 cents, which context might reduce, and the bound a note of length T puts on its own frequency, which context cannot touch. Taking them in quadrature, the shareable fraction is the curve. At a quarter-second note it is 0.21, so the one-half priced earlier is not available at all until each note lasts 486 milliseconds — which is exactly the crossover found by a different route, because a correlation of a half is the two parts being equal. The step is the convention used here, which takes the larger of the two rather than their sum and therefore says the shareable fraction below the crossover is zero.

The part of the error a key cannot touch

Four earlier essays turn one dial — the correlation between two notes' pitch errors — and apply it to the whole of a note's limen. Half of that limen is not the listener's: a note of finite length does not carry its frequency more finely than 1/2T, and no context can put information into a signal that is not there. So the correlation has a ceiling, it is 0.21 at a quarter-second note at A4 and 0.07 at A2, and the figure that prices a correlation of one half is drawn where one half is unavailable.

7 figures
Thirty cents out of tune is heard as 8 on a 125 ms note and 25 on a 1 s one. How far out of tune a note sounds against how far out of tune it is, at 4 note lengths, at 440 hertz. The key is treated as a prior over pitch: a mixture of Gaussians on the twelve scale degrees, weighted by Krumhansl and Kessler's probe-tone profile and given the width the degree account already uses. The likelihood is the note's own effective limen, which for a short note is the Fourier bound 1/2T. The estimate is the posterior mean, and the shrinkage toward a prior is one line of arithmetic. At 1 s a thirty-cent mistuning is heard as 25.0 cents and at 125 ms as 7.8. Every curve turns back up near the middle of the semitone, because past there the nearest degree is the other one and the pull reverses. The buttons sound at A4, which is the pitch the figure is computed at, at the shortest note length it draws.

A short note is heard more in tune than it is

Four earlier essays treat a key as something that reduces the noise in a pitch judgement. Treat it instead as a prior and the prediction changes kind: not a smaller error but a systematic bias, pulling a short note toward the nearest scale degree by an amount the Fourier bound sets. Thirty cents out of tune on an eighth-of-a-second note is heard as eight. And the part the debt got wrong is the part that matters — the bias does not vanish on a long note. It stops at 17 per cent at A4 and at 48 per cent at A2, because the likelihood's width has a floor that no duration removes.

7 figures
Every beat in a family is the same depth, and none is near the threshold. The 8 members of the beat family a octave mistuned by 6.0 cents makes at A3 on a real string, each placed at the rate it beats and at the modulation index a listener's filter delivers there. The rising curve is the published detection threshold for amplitude modulation, which is flat at 0.03 below about fifty fluctuations a second and rises above it. The flat dashed line at 0.667 is the depth the pair has in isolation, and it is the same for every member: on a spectrum falling as one over n, the k-th member pairs partials 2k and 1k, whose ratio is 2.00 whatever k is. What the filled points show is the smaller effect that does depend on the member — the partials on either side of the coincidence leak into the same filter, add level without adding fluctuation, and dilute the index from 0.662 to 0.405. Inside the countable rate window the narrowest margin over the threshold is a factor of 16.7, on member 4. The depth criterion removes nothing.

Every member of a beat family is the same depth

A mistuned octave's beats have been counted on their rate and their place, with the depth recorded as the thing left out and a prediction that the shallow upper members would take the count from four to two. The depth turns out not to fall at all: on any power-law spectrum every member of a family has exactly the modulation index its interval's own ratio gives, at every register, on every wire. The count does fall to two, and the thing that takes it there is the criterion that essay was already using.

7 figures
The same interval, mistuned by the same amount, at each of its two ends. A C to G in the major key, played 25 cents wrong, with the departure carried by the lower note, split between the two, and carried by the upper note. All three are the same interval size; what differs is which note is off the scale. The share of the departure that survives into what a listener hears is 32 per cent when the lower note carries it and 56 when the upper does. The middle bar is the mean of the other two to within a hundredth, so the averaging is linear and the asymmetry is the whole of the effect. Two things produce it: the prior is 9.0 cents wide at the C and 10.0 at the G, and the likelihood is 13.2 cents wide at the lower pitch and 8.8 at the higher.

An interval is two posteriors subtracted

Treating a key as a prior over one note predicts that an interval's pull is not the single-note pull doubled, because the two degrees are not equally weighted. Half of that is wrong: splitting a mistuning between the two notes gives exactly the mean of what each end gives alone, to a thousandth, at every one of the twenty-one intervals in the scale. What is not the mean is which end carries it — and the pull turns out to be largest not on the shortest notes but on notes of about an eighth of a second, where the likelihood is a quarter of a semitone wide.

7 figures
Which intervals in a key can be mistuned invisibly, and from which end. Every interval between two degrees of the major scale, ranked by how differently its two ends treat a 25-cent departure on notes of 0.25 seconds. The C to B is the most lopsided, at 47 points: a mistuning on its upper note reaches the listener nearly 2.5 times as strongly as the same mistuning on its lower one. The D to E is the most even, at -0. A negative bar is an interval whose LOWER note is the one that carries a mistuning into the listener, which happens whenever the lower degree is the less specified of the two. No account of interval perception predicts a table like this, because an interval is usually treated as one quantity rather than as a difference of two estimates.

Which end the mistuning is on

Twenty-one intervals in the major scale, each with two ends, and the same twenty-five cents reaches a listener at anywhere between 32 and 79 per cent of its size depending on which of the two notes carries it. The most lopsided is the tonic to the leading note, where a departure on the upper note arrives two and a half times as strongly as the same departure on the lower. Two mechanisms produce it and they can be separated by one flag: two thirds of the asymmetry is register and one third is the key.

7 figures
The setting a listener reports is not the interval they preferred. Five published preferences for an interval size, and the value a listener would have to play in a key context for that preference to be what they hear. The key pulls a heard interval back toward the scale, so a listener adjusting until it sounds right has to overshoot — and the reported setting, which is what they played, exaggerates the preference. The corrections run from 1.6 cents to 12.8 on notes of 0.25 seconds. The largest is nearly a syntonic comma, on a preference of thirteen cents, which is to say that the correction is bigger than the effect it is a correction to.

The setting is not the preference

Every published number for an interval listeners prefer — the pure third a quartet is said to find, the raised leading note, the harmonic seventh — is a value somebody adjusted until it sounded right. A listener adjusting inside a key is adjusting through the posterior computed just before, so the value they stopped at is not the one they preferred: it is the one whose heard size equals it. On quarter-second notes the correction for a pure major third is 12.8 cents, which is nearly a syntonic comma and is larger than the 13.7-cent preference it corrects.

6 figures