The collection

Every essay — page 12

Page 12 of 21, continuing through the fields in the same order.

Pitch and tuning Intervals and chords Scales and modes Harmony and voice leading Rhythm and metre Timbre and acoustics Perception and the listener Instruments and their design Form and structure Series Objects Sounds Search

Rhythm and metre

Time divided, subdivided and deliberately misaligned — drawn as a circle rather than a line.

A cycle already known, against a cycle just arrived at. How many bits of uncertainty about position a listener has, against how long the cycle takes, for a single timeline and for a layered colotomy — each drawn twice, once as a listener arriving and once as a listener who has been hearing it long enough to settle. At 1.6 seconds a cycle the timeline goes 0.94 bits arriving and 0.00 settled, and the colotomy 0.71 and 0.00; At 16 seconds a cycle the timeline goes 1.10 bits arriving and 0.08 settled, and the colotomy 0.93 and 0.23; At 60 seconds a cycle the timeline goes 2.47 bits arriving and 2.27 settled, and the colotomy 2.14 and 1.89. The gap between each pair is what the repetitions are worth, and it narrows as the cycle slows. The two designs are drawn at their own step counts rather than at equal strokes, so the levels here are not the earlier ones and the gaps are.

Repetition buys least where it is needed most

Every locating figure so far is a listener arriving — the uncertainty averaged over the first cycle heard. Cyclic music comes round dozens of times, and the same model already carries the answer for a listener who has settled: a floor of uncertainty that nothing had read. At two seconds a cycle the repetitions close the whole gap. At sixty they close eight per cent for a single timeline and twelve for a layered code. A slow cycle is worse on the first hearing and gains less from the second, and the two disadvantages compound.

5 figures
A listener who hears only the landmarks recognises the metre sooner. The share of metres recognised within the 16 steps of a 3.5-second present after coming in at a random step, against how many metres the listener knows, for five listeners: every onset with nothing marked, every onset with the long beats accented, only the downbeats and long beats with nothing marked, every onset with the downbeat accented, and only the downbeats and long beats with the downbeat marked. Every onset, nothing marked: 5 known, 49% within the present, 3% never; 10 known, 19% within the present, 8% never; 30 known, 1% within the present, 17% never; 100 known, 0% within the present, 34% never. Every onset, long beats accented: 5 known, 68% within the present, 3% never; 10 known, 39% within the present, 8% never; 30 known, 6% within the present, 17% never; 100 known, 0% within the present, 34% never. Landmarks only, nothing marked: 5 known, 69% within the present, 0% never; 10 known, 61% within the present, 3% never; 30 known, 26% within the present, 3% never; 100 known, 9% within the present, 9% never. Every onset, downbeat accented: 5 known, 88% within the present, 0% never; 10 known, 74% within the present, 0% never; 30 known, 54% within the present, 0% never; 100 known, 14% within the present, 0% never. Landmarks only, downbeat marked: 5 known, 91% within the present, 0% never; 10 known, 82% within the present, 0% never; 30 known, 56% within the present, 0% never; 100 known, 23% within the present, 0% never.

A late dancer needs the landmarks, not the rhythm

A listener who joins an additive-metre dance part-way recognises it far more reliably with the downbeat accented. Strip the stream down to its landmarks — the onsets that begin a bar or a long beat, with every other onset removed — and the listener does as well or better: knowing a hundred metres, 23 per cent are recognised within the present against 14 with every onset. Unmarked, the landmarks still beat every onset with the long beats accented. Neither half does it alone; what identifies a metre from a late entry is where its long beats sit relative to its bar.

5 figures
Against a pulse, the bell pattern is the quickest of its orders to place. The bits of position a listener is still missing, averaged over the first cycle heard, for each cyclic order of the gaps 1 1 2 2 2 2 2 in 12 steps, heard alone, against a pulse every three steps and against a pulse every four, at perfect memory, half-life 3 steps, half-life 1.5 steps. 2 2 2 1 2 1 2: alone 1.08, 1.23, 1.73; against a pulse every 3 steps 0.69, 0.75, 0.96; against a pulse every 4 steps 0.63, 0.67, 0.85. 2 2 2 2 1 1 2: alone 1.22, 1.63, 2.02; against a pulse every 3 steps 0.60, 0.73, 0.92; against a pulse every 4 steps 0.83, 1.07, 1.36. 2 2 1 2 2 1 2 (the standard bell pattern): alone 1.25, 1.45, 1.82; against a pulse every 3 steps 0.54, 0.56, 0.67; against a pulse every 4 steps 0.55, 0.57, 0.68. Alone, the bell pattern is not the quickest order to place at any memory. Against either pulse it is the quickest at every memory.

Against a pulse the bell pattern is the easiest to place

Heard alone, the standard bell pattern is not the quickest order of its own gaps to place in its cycle, for a listener with any memory. Heard against a pulse every three steps or every four — which is how anyone hears it — it is the quickest, at every memory and at every alignment of pulse and bell, and by a wide margin: at a memory of a quarter of the cycle, 0.56 bits unplaced over the first cycle against 0.73 for either rival against a pulse in threes. Six of eight named timelines do the same. A timeline's order of gaps looks chosen for how it sits against the beat, not for how it sounds alone.

5 figures
A timed expectation would erase a slow cycle's cost, and a listener cannot time a slow cycle that well. Bits of position a listener with a 3.5-second memory is still missing over the first cycle of son clave, against how long the cycle takes, for a newcomer with no expectation, a listener timing the cycle with the Weber fraction a duration that long is judged with, and a listener timing it to ten per cent. a newcomer, no expectation: 2 s 0.94, 8 s 0.97, 24 s 1.41, 40 s 2.02, 60 s 2.47; timing as well as listeners do: 2 s 0.68 (w 0.150), 8 s 0.69 (w 0.150), 24 s 0.85 (w 0.150), 40 s 1.90 (w 0.375), 60 s 2.38 (w 0.375); timing the cycle to ten per cent: 2 s 0.47, 8 s 0.48, 24 s 0.55, 40 s 0.73, 60 s 1.02. At ten per cent even a sixty-second cycle is placed about as well as a newcomer places a two-second one. At the precision a listener actually has for durations of half a minute or more, the expectation is worth a tenth of a bit.

An expectation cannot rescue a cycle too slow to time

A listener who knows a piece arrives with an expectation of where in the cycle they are, and the size of that expectation was the number the last essay said nobody had. It can be given one: a listener who has been timing the cycle carries a spread of their Weber fraction times the cycle, which is the same number of steps at any tempo. Timed to ten per cent, a forty-second cycle would be placed better than a newcomer places a two-second one. But forty seconds is judged in the band where the Weber fraction is nearer forty per cent, and there the expectation is worth a tenth of a bit.

5 figures

Timbre and acoustics

What makes a clarinet a clarinet, and what a room does to it before it reaches an ear.

The first eight partials of a string. A string vibrating in one, two, three and more equal parts, with the frequency ratio and the nearest named note beside each. The seventh partial is a third of a semitone flat of anything on a keyboard, which is a fact about strings rather than about tuning.

A string does everything at once

A plucked string does not vibrate at one frequency. It vibrates at all the whole-number multiples of one frequency simultaneously, and nearly everything in music theory is downstream of that fact.

6 figures
Four spectra of the same note. The amplitude of each partial for 4 timbres at the same pitch — pure, string, clarinet, bell. These are the exact lists the sound buttons here synthesise from, so the picture and the sound are the same data.

The ear hears the list, not the shape

Two sounds with the same partials and different phases have completely different waveforms and sound identical. What the ear extracts is a list of frequencies and strengths, and everything else is discarded.

6 figures
Three envelopes. How loudness changes over the life of a note, for plucked, bowed and struck. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.

The shape of a note, which is most of what an instrument is

Cut the first fifty milliseconds off a recorded piano and listeners stop calling it a piano. The attack carries more identity than the steady tone it leads into, and it is the part every spectrum plot leaves out.

6 figures
A small room's lowest modes. The first few axial standing waves of a room, drawn in plan, with the frequency of every mode below 160 hertz listed underneath. The low modes are far apart in frequency, so some bass notes are loud in one corner and absent in another. The sound buttons play these two octaves above their real pitch, because a room's lowest modes are below what most speakers reproduce.

The room is part of the instrument

A room has frequencies it supports and frequencies it will not. In a small one those frequencies are far apart, so some bass notes are loud in one corner and absent in another — and no equipment fixes it.

7 figures
The vowel in "hod", sung at 110 Hz. The partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.

A vowel is two resonances

The vowel in "heed" is the same vowel sung high or low, and nothing about it is a property of the note. It is two peaks in the response of the mouth, sitting at fixed frequencies while the partials of the voice slide underneath them.

6 figures
Where a tuned piano actually sits. The departure from equal temperament of every key of a small upright, computed from the stiffness of its strings and the fact that a tuner sets octaves without beats rather than at a ratio of two to one. The treble ends up 52 cents sharp and the bass 18 cents flat, and neither is an error.

The piano is tuned wrong on purpose

Every well-tuned piano has a sharp treble and a flat bass, by up to a third of a semitone at the extremes. It is not an error, it is not a compromise about keys, and it follows from one property of a steel wire that can be computed from its diameter and its length.

6 figures
Reverberation time, by two formulas. Sixty-decibel decay time against average absorption, divided by the room's volume-to-surface ratio so that every room sits on the same pair of curves. Sabine's equation, which is the one every textbook gives, and Eyring's correction to it. They agree in the reflective rooms Sabine measured and separate above ᾱ ≈ 0.18: at 0.6 Sabine reads 53% high, and at ᾱ = 1 — a room whose walls absorb everything, which is the outdoors — it still returns a positive time for a space with no reverberation at all. Marked: a concert hall 1.59 s, a stone church 3.37 s, a studio live room 0.45 s, a carpeted bedroom 0.13 s, an anechoic chamber 0.03 s.

How long a room rings, and where the formula stops

Sabine's reverberation time is one line of arithmetic — volume over absorption — and it built the modern concert hall. It also predicts that a room whose walls absorb everything still rings, which is a room with no reverberation at all, and the error is largest in exactly the rooms most music is now made in.

6 figures
Three attacks, the first 50 ms. How loudness changes over the life of a note, for plucked, bowed and struck, drawn over the first 50 milliseconds. By the right-hand edge the plucked note is at 91%, the bowed note is at 36%, the struck note is at 97% — attack times of 4 ms, 140 ms, 2 ms, a spread of 70 to one, and the part a listener uses to tell them apart. Remove the attack from a recorded piano and it stops sounding like a piano, which is the shortest demonstration that the envelope carries as much identity as the spectrum.

The first fifty milliseconds

A spectrum is supposed to be what makes a trumpet a trumpet. Cut the first fifty milliseconds off a recorded note and listeners stop being able to name the instrument — while the spectrum they are hearing is unchanged. Identity is in the part of the sound that ends before the note has properly started.

6 figures
Partial 3, mistuned by 3%. A 10-partial tone on 220 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 3.0 per cent, which is 19.8 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune.

What makes two partials one note

A note is a stack of ten or twenty simultaneous tones and is heard as one thing. The obvious explanation is that they are whole-number multiples of a fundamental — and the obvious explanation is not sufficient. Mistune one partial by three per cent and it leaves the note; give a perfectly harmonic partial a thirty-millisecond head start and it leaves too. Shared behaviour beats arithmetic.

6 figures
A bowed string on 196 Hz, through a violin body. The source is a sawtooth at one over n; the filter is the body's measured response, with A0 at 275 Hz, B1− at 460 Hz, B1+ at 540 Hz, bridge hill at 2500 Hz. What is radiated is their product, drawn as the bars. The resonances stay where they are when the note changes, exactly as a vowel's formants do — which is why an instrument has a voice rather than a tone, and why the same argument that identifies a vowel identifies a violin. The body frequencies are measured means over full-size instruments rather than computed from a plate.

The body is the filter

A violin string radiates almost nothing. What reaches a room is the string's sawtooth multiplied by the body's response, and that response is a comb of measured resonances that stays put while the note moves. It is the same arithmetic that identifies a vowel, on wood instead of a mouth — which is why an instrument has a voice rather than a tone.

5 figures
A string mode swept through a body resonance at 460 Hz. What the string plays against what comes out. Away from the resonance the two are the same and the line is the diagonal. Near it the mode splits into a pair, and the note warbles at the difference between them — 12.9 Hz at the centre, which is slow enough to be counted and far too fast to be a tremolo. The splitting is a coupled oscillator and has nothing to do with the wolf fifth of a tuning system, which is a twenty-three cent arithmetic residue and shares only the word.

The other wolf

A cellist's wolf note is a string mode landing on a body resonance, at which point the two stop being separable and start exchanging energy — the mode splits in two and the note warbles at the difference. It is a coupled oscillator. The tuning system's wolf is twelve fifths failing to close by 23.5 cents. They share a word and nothing else.

6 figures
Where each frequency goes, from a source 18 cm across. Polar response of a circular radiator of radius 9 cm at 200 Hz (ka = 0.3), 800 Hz (ka = 1.3), 2000 Hz (ka = 3.3), 5000 Hz (ka = 8.2). Zero degrees is straight ahead. The low frequency is a circle — it goes everywhere — and the high one is a narrow lobe with nulls either side of it, so a listener off to the side hears the same note with its top missing.

An instrument points

A source radiates evenly while it is small compared with the wavelength and beams once it is not, and the crossover is one number. So the same instrument is omnidirectional in its bottom octave and a searchlight in its top one — which means its spectrum depends on where the listener is standing, and a microphone position is a choice about what the instrument sounds like.

7 figures
6 chords in a gothic cathedral. Each chord's reverberant decay in a room with a 8 second reverberation time, at 1 chord a second. Decay is linear in decibels, so each line is straight with a slope of -7.5 dB a second. When a chord arrives, 2 earlier ones are still above 20 dB down.

The room chooses the harmonic rhythm

A chord in a cathedral is still sounding, seven decibels down, when the next one arrives — and the one after that, and the one after that. Reverberation is linear in decibels, so the number of chords audible at once is one number divided by another, and it puts a hard ceiling on how fast a composer writing for that building can change harmony. The ceiling is computable, and the music written for those rooms sits under it.

7 figures
Where each room stops being a set of resonances. The Schroeder frequency of 6 rooms — a carpeted bedroom at 217 Hz, a domestic living room at 183 Hz, a rehearsal room at 120 Hz, a jazz club at 68 Hz, a shoebox concert hall at 21 Hz, a gothic cathedral at 36 Hz — marked on a logarithmic frequency axis with the ranges of 4 instruments underneath. Below the mark a room is a handful of separable modes and a note's loudness depends on where the listener is standing; above it the modes overlap and the room is described by one decay time.

Where a room stops being a room

A small room has frequencies it supports and frequencies it will not, and a hall has a reverberation time. Those are two separate accounts and they are descriptions of the same building at different frequencies. The crossover is one formula, and in a bedroom it lands at about two hundred hertz — in the middle of the bass register, and above nothing at all in a concert hall.

7 figures
Six reverberation times for one room. Sabine's arithmetic evaluated in each octave band from the published absorption coefficients of the surfaces. a large stone church runs from 6.3 seconds at 125 Hz to 2.4 at 4 kHz — a bass ratio of 1.38, where concert halls are specified between 1.1 and 1.25.

A room does not decay evenly

Sabine's arithmetic gives one number and absorption is a strong function of frequency, so a room has six reverberation times rather than one. A stone church rings for 6.3 seconds at 125 hertz and 2.4 at 4 kilohertz, which means a chord left in it does not fade — it changes shape, losing its top before it loses its bottom, and arriving at the listener as a different sonority from the one played.

7 figures
One decay, two verdicts, and the line is the listener's. The early-to-late energy ratio against reverberation time, for 50 millisecond and 80 millisecond windows. Nothing about the room differs between the curves; only where the line is drawn across its decay. Zero comes at 1.00 seconds for the 50 ms window and 1.59 seconds for the 80 ms window — which are, to two figures, the published rules of thumb for a room for speech and a room for music. The design targets were not put in; they came out.

The first eighty milliseconds are a different room

Draw a line across a room's decay and the energy on either side is two opposite verdicts about one building — clarity before it, reverberation after. The line is a property of the ear, not of the room, and putting it at 50 milliseconds and at 80 makes the two published design targets fall out — a room for speech at one second, a room for music at 1.6.

8 figures
The roughness curve for ideal bar. The partials are at 1 : 2.756 : 5.404 : 8.933 : 13.340 times the fundamental, which is not a harmonic series, so nothing about where the wells fall can be read off the small whole numbers. Sensory dissonance between two complex tones as the upper one is swept through 13.00 semitones, computed by summing the roughness between every pair of their partials. Nothing here is placed by hand. The wells this spectrum produces sit 7 cents below 3/2, 14 cents below 5/3, 13 cents above 7/4, found by scanning the curve rather than by marking them. The wells are where this spectrum's partials coincide, and for a spectrum with no even partials they are not where an octave-based scale would put them.

The spectrum that was supposed to explain the gamelan

Run the model that designed the Bohlen–Pierce scale on a bar's partials and it asks for a compressed pseudo-octave at 1166 cents and wells at 694 and 870. A measured slendro's degrees are at 231, 474, 717 and 955, and only one of the four is near anything the model wants. Sweep the partials and a spectrum wanting a 240-cent step can be found — sitting sixty per cent of the way up its own roughness curve, which is what a scale-level test says about the whole comparison.

7 figures
Long-term average spectra: an orchestra, playing forte against a trained operatic soloist. Each source's mean spectrum over a long passage, in decibels below its own strongest region, on a logarithmic frequency axis. An orchestra, playing forte peaks at 250 Hz and is 30 dB down by 3,150 Hz; a trained operatic soloist peaks at 250 Hz and is 11 dB down by 3,150 Hz. The shapes are the same until about 1 kHz and separate above it: at 3153 Hz the difference is 19.0 decibels, which is the largest anywhere in the range. Nothing here is about level. Both curves are drawn against their own peaks, so what is being compared is shape.

One voice over ninety players

A soloist heard over a full orchestra is not louder than it and could not be. What the trained voice does instead is put a peak of energy at three kilohertz, which is where the orchestra's spectrum has already fallen away and where the ear's own threshold happens to be lowest. Nineteen decibels of advantage, in a place nobody is competing for.

8 figures
What gets out of an opening, for 4 openings. The fraction of the wave's energy radiated at an open end against frequency, in the baffled-piston model — the radiation resistance of a circular piston, normalised to the tube's own impedance. Each curve runs from nothing at the bottom, where the opening is far smaller than a wavelength and the wave simply turns round, to everything above ka ≈ 2. Half the energy leaves at 6364 Hz for a flute's embouchure end (radius 10 mm), 2015 Hz for a clarinet's bell (radius 30 mm), 975 Hz for a trumpet's bell (radius 62 mm), 403 Hz for a horn's bell (radius 150 mm). The crossover goes as one over the radius, so the widest and narrowest here are 15.8 times apart in frequency. The same number decides how strongly the tube resonates and how much sound it makes, which is why a bell cannot brighten an instrument without also weakening its own resonances.

The bell decides what gets out

A tube resonates because the wave turns round at the open end, and it is audible because some of the wave does not. Those are the same number with opposite signs. One quantity — the size of the opening against a wavelength — decides how loud an instrument is, how bright it is and how directional it is, and a bell moves the boundary rather than removing it.

8 figures
The vowel in "hod", sung at 110 Hz. The partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.

The sound a listener knows best

A voice is recognisable across every vowel it says, across two octaves of pitch, down a bad telephone line and in a whisper where there is no pitch at all. Nothing that survives all of that can be a frequency. What survives is a ratio: the resonances of a vocal tract are set by its length, so a shorter tract multiplies every formant by the same factor, and identity is a scale on the spectral envelope rather than a position within it. Between an adult man and a child the whole pattern moves by a fifth, and the vowel does not change at all.

7 figures