The collection

Every essay — page 14

Page 14 of 21, continuing through the fields in the same order.

Pitch and tuning Intervals and chords Scales and modes Harmony and voice leading Rhythm and metre Timbre and acoustics Perception and the listener Instruments and their design Form and structure Series Objects Sounds Search

Timbre and acoustics

What makes a clarinet a clarinet, and what a room does to it before it reaches an ear.

An inversion lasts as long as its outer sixth. The six three-note voicings of a major and a minor triad, each over a bass of C3 struck at 80 decibels, with how long the strongest partial coincidence of each of its three intervals survives the strike. The interval that goes first is marked, and its time is how long the chord keeps the evidence of all its intervals at once. major, root position: major third 5:4 1.26 s, minor third 6:5 1.03 s, fifth 3:2 2.14 s; the chord 1.03 s. major, sixth chord: minor third 6:5 1.04 s, fourth 4:3 1.62 s, minor sixth 8:5 0.74 s; the chord 0.74 s. major, six-four: fourth 4:3 1.60 s, major third 5:4 1.27 s, major sixth 5:3 1.26 s; the chord 1.26 s. minor, root position: minor third 6:5 1.04 s, major third 5:4 1.27 s, fifth 3:2 2.14 s; the chord 1.04 s. minor, sixth chord: major third 5:4 1.26 s, fourth 4:3 1.63 s, major sixth 5:3 1.26 s; the chord 1.26 s. minor, six-four: fourth 4:3 1.60 s, minor third 6:5 1.03 s, minor sixth 8:5 0.74 s; the chord 0.74 s. The major six-four lasts longest and the major sixth chord shortest; every root position is held to its minor third's life.

An inversion lasts as long as its outer sixth

A struck interval keeps the partial coincidence that names it for a time set by its ratio, and a chord is three intervals at once. Voiced over one bass and struck on a piano, a triad keeps the evidence of all three only as long as its weakest one lasts, and for an inversion that is the sixth on the outside: a major sixth lasts as long as a major third, a minor sixth dies first. So the major six-four and the minor sixth chord are the most durable voicings of their triads and the major sixth chord and the minor six-four the least — and unlike a dyad, a triad's inversions keep their order by roughness through almost the whole decay.

6 figures
A doubled pizzicato gives its note away while it is still the louder. The power of a violin plucked, against a flue pipe holding the same note at 392 hertz, through the first 600 milliseconds of the pluck, with the pluck starting 12 decibels up and its fundamental decaying over 1 second. With each partial losing level in proportion to its number, the composite stops resembling the pluck at 70 ms, when the pluck is still 5.2 decibels the louder. With every partial fading together it would keep the note until 543 ms. The dashed line is the balance at which the steady-state doubling changes owner, minus 20.6 decibels: the release crosses the owner long before its balance gets there, because what hands the note over is the pluck's upper partials going, not its level.

A doubled pizzicato gives its note away early

The attack turns the balance between two players on one note by a few decibels and stops. A pluck does not stop — every partial of it decays, so a pizzicato doubled by a held instrument walks the balance for the whole note, and the expectation was a handover as slow as the decay. It is fast. A one-second pizzicato over a flute loses its note in 70 milliseconds, while it is still five decibels the louder, because what hands the note over is its upper partials going first. A uniform fade would have kept it eight times as long.

7 figures
How long a doubled pizzicato keeps its note, seat by seat, in two rooms. How long a doubled violin pizzicato on 392 hertz keeps its note against the metres from the players, the pluck starting 12 dB up and decaying over 1 s with a loss exponent of 1. a concert hall, a flue pipe: 1 → 86 ms, 1.5 → 123 ms, 2 → 226 ms, 3 → 359 ms, 5 → 445 ms, 7 → 481 ms, 10 → 506 ms, 15 → 522 ms, 20 → 528 ms, 30 → 532 ms; a concert hall, an oboe: 1 → 52 ms, 1.5 → 55 ms, 2 → 59 ms, 3 → 77 ms, 5 → 149 ms, 7 → 195 ms, 10 → 224 ms, 15 → 242 ms, 20 → 248 ms, 30 → 254 ms; a concert hall, a clarinet: 1 → 44 ms, 1.5 → 45 ms, 2 → 45 ms, 3 → 47 ms, 5 → 53 ms, 7 → 65 ms, 10 → 86 ms, 15 → 105 ms, 20 → 112 ms, 30 → 118 ms; a large stone church, a flue pipe: 1 → 440 ms, 1.5 → 578 ms, 2 → 651 ms, 3 → 728 ms, 5 → 784 ms, 7 → 803 ms, 10 → 814 ms, 15 → 820 ms, 20 → 822 ms, 30 → 824 ms; a large stone church, an oboe: 1 → 56 ms, 1.5 → 65 ms, 2 → 89 ms, 3 → 207 ms, 5 → 281 ms, 7 → 303 ms, 10 → 315 ms, 15 → 322 ms, 20 → 325 ms, 30 → 326 ms; a large stone church, a clarinet: 1 → 44 ms, 1.5 → 43 ms, 2 → 43 ms, 3 → 44 ms, 5 → 53 ms, 7 → 67 ms, 10 → 80 ms, 15 → 89 ms, 20 → 92 ms, 30 → 94 ms. The mid-band critical distance is 5.3 m in a concert hall and 2.3 m in a large stone church. In none of the 60 cases does the note return to the pluck once it has left.

A room keeps a pizzicato from giving its note away

Doubled by a flute, a one-second pizzicato loses its note in 70 milliseconds dry, because its upper partials go first. The question left open was whether a room, whose reverberation keeps those partials alive, gives the note back afterwards. It does not give it back. It stops the note going: ten metres into a concert hall the pluck keeps it for 506 milliseconds, in a stone church for 814, and the room's own uneven decay takes back between a quarter and two fifths of that. In a room the loss law that decided everything dry matters a tenth as much, because the room's decay has become the clock.

6 figures
A bow shows the loss law a blow conceals. The same string on 130.8 hertz under three loss laws, drawn twice each: struck, and held by a continuous drive. The pale marks are the spectrum a blow produces, and they are identical in all three rows — a strike is the source spectrum and has no loss in it yet, which is why both endpoints of a struck note were found to carry nothing about the law. The solid marks are where each partial settles when a drive balances its own loss, at drive over loss, so the steady spectrum rolls off as the source's roll-off plus the exponent. At an exponent of 0.5 the held spectrum's centroid sits at 167 hertz, 5.7 semitones under the strike's 232; At an exponent of 1 the held spectrum's centroid sits at 144 hertz, 8.2 semitones under the strike's 232; At an exponent of 2 the held spectrum's centroid sits at 133 hertz, 9.6 semitones under the strike's 232. The quantity that is invisible at both ends of a struck note is the slope of a bowed one, for as long as the bow moves.

A bow holds the number a blow hides

Two struck notes with different loss laws are identical at the strike and identical at the end, which is why separating them at all meant looking in the middle. Drive the same two strings continuously and the loss law stops being a rate and becomes a slope: each partial settles at its drive over its own loss, so the exponent adds to the source's roll-off and sits in the spectrum for as long as the bow moves. It is 18.7 decibels of separation available from the first instant, against 41.3 that a blow delivers after four tenths of a second and then takes away.

7 figures
From partial 3 the room is the slower of the two. Decay rates in nepers a second for each partial of a note on 130.8 hertz, in a concert hall. The rising curve is the string's own loss, which grows as the partial number to the power 1. The flat-ish curve is the room's, from its reverberation time at that partial's frequency. A reverberant field is the source convolved with the room, so a partial's tail falls at the SLOWER of the two — the heavy line — and the room keeps returning energy the string has stopped making. From partial 3, at 392 hertz, the room is in charge: 6 of the note's 8 partials are held up by the room rather than let go by the string. Those are exactly the partials the string was losing fastest, which is why the room does not merely lengthen the note.

The room is the slower of the two

A reverberant field is the source convolved with the room, so a partial's tail falls at the slower of the two rates rather than at their sum — and the room is slower for exactly the partials the string is losing fastest. Half a note's colour is gone in 0.163 seconds in no room at all, 0.313 in a concert hall and 1.441 in a stone church. The destination is identical in all three, because a room cannot hold a partial up above the fundamental it is also holding. What a hall takes away is the rate, and the rate was the whole of the identity cue.

7 figures
The drain does not stop when the key does. The note does.. Semitones of colour gone, against time, for a note on 130.8 hertz left to ring and for the same note released after 0.4 seconds onto a damper of 0.15 seconds. The two curves lie on each other until the key comes up, and the damped one then ends: the note is inaudible at 0.52 seconds with 8.7 of the free note's 9.9 semitones delivered. A damper adds one loss to every partial alike, so it adds the same number to every decay rate and leaves every DIFFERENCE between rates exactly as it was — the spectrum at each instant is the ringing spectrum shifted bodily down by 400 decibels a second. The colour goes on draining at its own rate the whole time. What the damper takes away is not the drain but the seconds.

A damper changes the clock, not the colour

A damper is an extra loss on the string rather than a second decay, so it adds the same number of nepers a second to every partial — and adding a constant to every rate leaves every difference between rates exactly where it was. The damped spectrum at any instant is the ringing spectrum at that instant shifted bodily down, to machine precision. The colour goes on draining at its own rate; the note simply runs out of seconds, and how many it gets is written on the page as a note value and a tempo.

6 figures
A room pulls the compass apart rather than evening it out. How long a pizzicato entering 6 decibels above a held note keeps the composite spectrum, at eight pitches across two and a half octaves, heard 15 metres from the stage. no room: 0.07, 0.05, 0.06, 0.06, 0.02, 0.05, 0.05, 0.04 seconds; a concert hall: 0.61, 0.27, 0.50, 0.38, 0.01, 0.25, 0.27, 0.19 seconds; a large stone church: 1.03, 0.39, 0.79, 0.56, never, 0.33, 0.35, 0.23 seconds. Dry the figures barely move — a spread of 3.0 across the whole compass — because a room is the thing that varies with frequency and there is none. In a hall the spread is 44. The room does not scale the dry answer by a constant: it multiplies it by between four and nine times depending on the note, and at C6 it makes the pluck's position worse rather than better, because the two instruments' spectra nearly coincide there and the pluck starts only 4.0 decibels ahead instead of twelve.

One note in the compass loses its pizzicato

Dry, how long a pluck keeps the composite spectrum barely depends on which note it plays: three-hundredths of a second at the worst pitch and seven at the best, a spread of three. In a concert hall the same eight notes spread by a factor of forty-three, and in a stone church one of them never gets the note at all. The room does not scale the dry answer by a constant — it multiplies it by between four and nine times depending on the pitch, and at the one note where the two instruments' spectra nearly coincide it makes the pluck's position worse instead of better.

6 figures
A damper is a loss on the string, so a room can overrule it. Decay rate in nepers a second against partial number, for a note on 130.8 hertz in a room of 2 seconds. The rising line is the string's own loss, 1.15 nepers a second at the fundamental and growing as the partial number to the power 1. The line above it is that plus the damper's 46.1, which is what the string does once the key comes up. The flat line is the room. What a listener receives is the SLOWER of the damped string and the room, because a hall goes on radiating what the string has already given it — and here the room is slower on 8 of 8 partials, from the fundamental upward. The composition proposed earlier — take the slower of the string and the room, then add the damper to whichever won — would put the damper outside the minimum, where nothing can overrule it, and would predict a note 2.54 seconds shorter than ringing where the arithmetic here predicts 0.90.

A damper cannot reach into the room

The essay before this one proposed the arithmetic for a damped note in a hall: take the slower of the string's rate and the room's, then add the damper's to whichever won. The composition is wrong, and it is wrong in the one place that decides the answer. A damper is a loss on the string, so it belongs inside the minimum where a room can overrule it — and past about three seconds of reverberation it is overruled on every partial, so the damper removes no audible seconds of note at all.

5 figures
No seventh chord can be spaced to last as long as a triad. Every inversion and spacing within 2 octaves over C3, struck at 80 dB, for two triads and five seventh chords: the bar is the longest any spacing keeps every pair's partial coincidence, and the tick is the bound set by the chord's worst pitch-class distance — the longest any presentation of that distance lasts. major triad: 1.26 s over 12 voicings, bound 1.26 set by the minor third; minor triad: 1.26 s over 12 voicings, bound 1.26 set by the minor third; dominant seventh: 0.66 s over 32 voicings, bound 0.64 set by the tone; major seventh: 0.42 s over 32 voicings, bound 0.38 set by the semitone; minor seventh: 0.69 s over 32 voicings, bound 0.64 set by the tone; half-diminished seventh: 0.69 s over 32 voicings, bound 0.64 set by the tone; diminished seventh: 0.86 s over 32 voicings, bound 0.87 set by the tritone. Every seventh chord contains a distance worse than any a triad contains, except the diminished seventh, whose distances are only minor thirds and tritones.

A seventh chord cannot be spaced to last like a triad

A struck triad keeps the partial coincidences of all its intervals for at most 1.26 seconds, whichever way it is spaced. Run the same census over every inversion and spacing of five kinds of seventh chord and none gets near: the dominant, minor and half-diminished sevenths top out at about two thirds of a second, the major seventh at 0.42. The diminished seventh, which theory calls the least stable of them, lasts longest at 0.86 — because it is the only one with no tone or semitone among its pitch-class distances, and the worst distance a chord contains sets a ceiling no spacing can lift.

5 figures
A staccato is the direct sound's, and the room takes it within a fifth of the critical distance. A note on 130.8 Hz held 0.4 s and damped, in a room of 2 s reverberation, heard at distances from 0.02 to 5 times the critical distance: how long after the release the note takes to fall 10 dB and 20 dB. To fall 10 dB: 24 ms at the source, 333 ms far away; 0.02: 24 ms, 0.05: 25 ms, 0.1: 25 ms, 0.15: 26 ms, 0.2: 28 ms, 0.3: 35 ms, 0.5: 103 ms, 0.75: 187 ms, 1: 234 ms, 1.5: 281 ms, 2: 302 ms, 3: 318 ms, 5: 328 ms; doubled by 0.38 of the critical distance. To fall 20 dB: 49 ms at the source, 667 ms far away; 0.02: 49 ms, 0.05: 51 ms, 0.1: 60 ms, 0.15: 117 ms, 0.2: 198 ms, 0.3: 308 ms, 0.5: 436 ms, 0.75: 520 ms, 1: 568 ms, 1.5: 614 ms, 2: 635 ms, 3: 652 ms, 5: 661 ms; doubled by 0.14 of the critical distance. Where the direct sound and the room are equal, the damper's work is already hidden: the room's copy is only 20 dB below the direct sound at a tenth of the critical distance, and a 20 dB fall reaches it there.

Only the player hears a staccato end

A damper stops a string in a seventh of a second, and in a hall the room goes on for two. A listener hears both, mixed in proportion to how close they sit, and the question was at what distance the short part stops mattering. The answer is closer than any seat. A damped note's twenty-decibel fall has doubled in length by a seventh of a hall's critical distance — 77 centimetres in a two-second concert hall — and by a quarter of it in a jazz club. The end of a staccato is something the pianist hears and the front row does not.

5 figures

Perception and the listener

Loudness is not amplitude and a beat is not in the signal. What the ear adds, hides, merges and supplies — measured, with the numbers every other field has been assuming.

Every point on one curve sounds equally loud. The equal-loudness contours of ISO 226:2003, evaluated from the standard's own parameters. The lowest curve is the threshold of hearing. Because the curves are not parallel — they crowd together in the bass and spread apart in the middle — the same change in decibels is a different change in loudness at every frequency, and a spectrum that was balanced at one level is not balanced at another.

The quietest thing audible, and why the volume knob is a tone control

A decibel is a fact about air. A phon is a fact about a listener, and the two do not line up — the map between them bends with frequency, and it bends differently at every level. One consequence is that turning a piece of music down does not turn all of it down equally, and the amount by which it does not is a number.

8 figures
What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward.

One sound hides another, and it hides upward

A tone can be made completely inaudible by a second tone that is nowhere near it in frequency, and the region it disappears into is lopsided. Masking spreads up the spectrum and barely down it, and the reach grows with the masker's level — so which line in an arrangement vanishes is a prediction, not a matter of taste.

6 figures
The threshold before, during and after a burst. The level a brief probe needs in order to be heard, plotted against when it happens relative to a 70 dB burst that occupies the shaded band. To the right is forward masking, which decays over about 200 milliseconds. To the left is backward masking: the threshold is raised for a probe that has already finished before the masker begins.

A sound hides what came before it

Masking does not stop when the masker does. A loud sound raises the threshold for about a fifth of a second after it ends, which is unremarkable, and for several milliseconds before it begins, which is not. The auditory present is a window rather than an instant, and inside the window the order of events is not the order they arrived.

5 figures
A 7-semitone sequence at 120 ms a tone. Tones drawn as pitch against time, one bar per tone. The events are the same in both readings of this pattern; what changes is whether a listener assigns them to one line that leaps back and forth or to two lines that each stay put. Nothing in the drawing decides which, and nothing in the sound does either.

The ear builds objects, and sometimes offers a choice

What arrives at an ear is one pressure signal. What a listener gets is a set of separate things — a violin, a voice, a car outside. The assignment is a construction, and the clearest evidence is that it can be flipped by changing nothing but the speed: one sequence of tones is a single line when slow and two lines when fast, with a wide region in between where the listener may choose.

7 figures
A source 45° off centre, and the path difference it makes. A head from above with a source to one side. The near ear is reached first; the far ear's path runs round the head, and the difference between the two is 13.1 centimetres, which at 343 metres a second is 381 microseconds. That number, and the level difference the head's shadow produces, are the whole of what the ear has to work with.

Two ears, and the whole of the difference is 655 microseconds

Direction is computed from two numbers — when a sound reaches each ear and how loud it is at each — and which of the two is usable is decided by the wavelength against the width of a head. The changeover frequency is not a design choice. It falls out of 343 metres a second and 17.5 centimetres, and it is why the mechanism of hearing where something is changes halfway up the piano.

7 figures
How late a reflection has to be before it is an echo. What a single reflection does to the sound it follows, against its delay, on a logarithmic axis. Under a millisecond the two combine into one image that is pulled towards the earlier source. From there out to a few tens of milliseconds the reflection is not heard as a separate event at all and does not move the image — it only changes the timbre. Past the echo threshold it becomes a second sound, and the threshold is five times later for speech than for a click.

The first wavefront wins

A room sends a hundred copies of every note to a listener from a hundred directions, and the listener hears one note in one place. The mechanism that does it is brutal and simple: for the first few tens of milliseconds after a sound arrives, everything that follows is denied a vote on where it came from — even when it is louder than the original.

5 figures
The smallest audible difference, and what has to clear it. The difference limen for frequency, converted from Wier, Jesteadt and Green's 1977 fit into cents, against the intervals and commas the rest of these essays argue about. Anything drawn below the curve is a quantity nobody can hear as a change of pitch; anything well above it is a quantity a listener can be asked about. The limen is for pure tones, successive, with trained listeners — the most favourable case there is, and therefore the right one to test a claim against.

How small a difference is audible

Every essay here about tuning has assumed a listener who can hear the difference between two systems. The assumption has a number: about five cents in the middle of the range. It clears the two commas fourfold and it does not clear the schisma at all, which sorts the whole subject of tuning into the part that is about music and the part that is about arithmetic.

6 figures
Three octaves, and none of them is 2:1. How far above an exact doubling the upper note of an octave is set, against frequency. The listener's octave is measured with pure tones, which have no partials to beat against each other, so nothing about a stiff string can account for it. The piano's stretch is a different quantity with a different cause, and the two are drawn together only so that the difference is visible.

The octave that is not two to one

The octave is the one interval nobody argues about: two to one, exact, in every tradition that has one. Asked to set an octave by ear, listeners set it wide — and they do it with pure tones, which have no partials to beat against each other. Whatever is stretching the octave, it is not the stiffness of a piano string.

5 figures
Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.

The chord is still major, and that is why temperament works

A major third can be seventeen cents wrong and still be a major third. That tolerance is not a failure of hearing — it is the reason the whole subject of tuning is a discussion rather than a catastrophe. Every temperament ever proposed moves intervals around inside their categories, and the one thing none of them may do is push one across a boundary.

6 figures
The probe-tone profile, major key. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.

Counting produced the hierarchy

Ask listeners how well each of the twelve notes fits after a passage in C major and the answers are not a smooth gradient. They fall into four groups with no overlap at all: the tonic, then the rest of the tonic triad, then the rest of the scale, then everything else — categories the subject had names for centuries before anybody ran the experiment.

4 figures
Two models of consonance, and where they disagree. Every chord scored twice: horizontally by summed Plomp–Levelt roughness, vertically by the largest integer needed to write it as members of one harmonic series. Both are supposed to be measuring consonance and both are computed here from the chord itself. They correlate, but not tightly enough to be the same claim, and the chords furthest from the diagonal are the ones any experiment has to be run on.

Consonance is half learned, and this is the half

This site's founding claim is that consonance is small whole numbers. Two computable models say so and they disagree about which chords — which is already awkward. The cross-cultural evidence is worse: listeners with little exposure to Western music discriminate roughness exactly as anyone does, match octaves exactly as anyone does, and rate consonant and dissonant chords as equally pleasant.

5 figures
Direct and reverberant sound in a shoebox concert hall. The direct sound falls six decibels for every doubling of distance and the reverberant field does not fall at all, so they cross once — at 5.5 metres in a room of 18700 cubic metres with a 2-second decay. Both are drawn relative to their level at that crossing. Everything past the crossing is a seat at which the room is louder than the instrument.

How far away the room takes over

Direct sound falls six decibels every time the distance doubles and the reverberant field does not fall at all, so the two cross once. In a concert hall the crossing is at about five and a half metres, which is nearer than nearly every seat — so almost everybody in almost every hall is hearing the building more than the players, and the number that says so is built from two quantities already computed — a room's reverberation time and an instrument's directivity.

7 figures
The ranking is settled either side of one narrow band. Remembered repetition — each bar's best match to an earlier bar, discounted by exp(−Δt/τ) with Δt in seconds — for 6 schemes at 108 beats a minute, against the decay constant τ on a logarithmic axis. The order of the schemes changes only between 8 and 13 seconds; outside that band it is fixed, so an estimate of τ wrong by any amount that stays outside it leaves the ranking alone.

A return has to be remembered

A stripe four bars off the diagonal and a stripe twenty-four bars off it are the same ink and are not the same experience. Convert the lag axis to seconds, discount every comparison by how long ago it was, and the ranking of these six schemes by how repetitive they are changes — and the decay constant and the tempo turn out to enter the arithmetic as one number rather than two.

8 figures
Four voices, placed by the arithmetic. The rules in force are: no parallel octaves; no parallel fifths; no voice crossing; no gap over an octave above the tenor; the leading note is not doubled; the leading note resolves, outer voices; no augmented melodic interval. The four parts are drawn lowest to highest — bass, tenor, alto, soprano. I – vi – ii – V – I in C major, realised in four voices by the cheapest set of voicings obeying 7 rules, at 16 semitones of motion in all. Every gap between adjacent voices narrower than the fission boundary of 5.2 semitones is marked, and below that boundary two parts cannot be heard as two however hard a listener tries.

A voice is a stream, and the ear decides which

Seven earlier essays have assigned voices to notes. Whether a listener follows the assignment is a separate question with laboratory numbers attached, and the numbers are unkind to it: two parts closer than about five semitones cannot be heard as two at any speed, a third of the gaps in the cheapest four-part writing are inside that limit, and in a third of chord changes the ear's own rule for continuing a line does not recover the parts as written.

7 figures