The collection

Every essay — page 16

Page 16 of 21, continuing through the fields in the same order.

Pitch and tuning Intervals and chords Scales and modes Harmony and voice leading Rhythm and metre Timbre and acoustics Perception and the listener Instruments and their design Form and structure Series Objects Sounds Search

Perception and the listener

Loudness is not amplitude and a beat is not in the signal. What the ear adds, hides, merges and supplies — measured, with the numbers every other field has been assuming.

A turn of 0.99 degrees tells front from back. A source 45 degrees off centre and its mirror image 135 degrees off, which produce the same interaural delay and are therefore the same signal to a listener who does not move. As the head turns the two predictions separate: the front source's delay falls and the rear source's rises, because the fold at ninety degrees puts them on opposite branches of the same curve. They differ by the 15-microsecond threshold after 0.99 degrees of turn — which is exactly half the 1.97 degrees a source would have to move for the same listener to notice it moving, and it is half for a reason: a turn displaces the two hypotheses from each other by twice what it displaces either of them from where it started.

The turn is half the angle

A stationary head cannot tell a sound in front from the same sound behind, and an earlier essay said so at length. The turn that breaks the confusion is 0.84 degrees — exactly half the angle a source would have to move for the same listener to notice it moving, and half for a reason. In a hall the same turn does something else: the source swings at 8.9 microseconds a degree and the room swings at 2.7, so a listener who moves is separating the soloist from the reverberation as well as the front from the back.

7 figures
The same geometry at four sizes of head. Woodworth's interaural delay against direction, for 4 head radii from 5.8 to 9.8 centimetres. The whole range runs from 431 microseconds for a newborn to 735 for a large adult, and it scales exactly with the radius because the delay is (r/c)(θ + sin θ) and r is a multiplier. The detection threshold does not scale with the listener, so the number of distinguishable delays across the whole range falls from 98 to 57: a smaller head has the same directions in front of it and a shorter ruler to measure them with.

A smaller head in the same hall

Ten earlier essays draw one head. Every parameter belonging to the room has been varied by some figure and the one belonging to the listener never has, and it is the only one whose change the detection threshold does not follow: a six-year-old in the same seat receives the same fifty-four reflections at the same instants and reads them onto an axis with seventy distinguishable positions instead of eighty-seven. The speed of sound, swept over every temperature a hall is ever at, changes nothing at all — and the reason it cannot is the reason head size can.

7 figures
A clarinet's partials, each on its own resonance. The 8 partials of the clarinet's chalumeau D that ride an impedance peak, each building toward its steady amplitude as 1 − exp(−t/τ) with τ = Q/πf from that peak's own Q. The time constants run from 11.9 milliseconds to 64.5, so the partials do not arrive at different times — they all begin the instant the reed does — and what differs is how fast each approaches its final level. The horizontal bars are how far apart the first and last are at three criteria: 5.5 ms at 10 per cent, 36.4 ms at 50 per cent, 121.0 ms at 90 per cent. A twenty-millisecond asynchrony is the threshold for hearing a partial out of a note, and this note crosses it at 32 per cent of steady amplitude — so whether a blown note's onset cue is unanimous or divided is decided entirely by how far along a partial has to be before it counts as having started.

A blown note does not start late, it starts slowly

Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.

7 figures
Eleven partials is one partial too many. What fraction of a spectrum the harmonicity census finds fused, against how many partials it is asked to census. At ten a perfect harmonic series fuses 10 of 10 and the fundamental it finds is the right one. At eleven it fuses 5 of 11 and the fundamental jumps to exactly 2.00 — the octave above. The cause is the cap the census carries for a reason established earlier: without it a bell fuses perfectly at a fundamental nobody could hear, so the search refuses any fundamental more than about ten harmonics below the top partial. At eleven partials the first thing that cap excludes is the series' own fundamental, and the census then takes the octave and calls every odd partial inharmonic. So the number of partials and the cap are the same number, and nothing had ever said so, because every earlier figure censuses ten.

Eleven partials is one too many

Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.

7 figures
The census with the criterion moved under it. Every instrument's speaking time in milliseconds, against the fraction of the steady amplitude counted as speaking. The criterion is in the wind instruments alone: a resonance takes −ln(1−p)·Q/(πf) to reach a fraction p, so those lines rise across the whole picture, while a bow's capture and an exciter's contact contain no criterion at all and are flat. Every earlier figure sits at 0.9, where the wind instruments are the slowest things in the collection by a factor of 20.5. At 0.05 they are the fastest: a violin's G3 string is the slowest at 15.5 milliseconds and a trumpet takes 1.4. The three clusters cross at a criterion between 0.18 and 0.39, which is inside the range the perceptual measurements work in — their three named criteria are 15 decibels below peak, 6 decibels below peak, and ninety per cent — and the settling figures have only ever used the third of them.

Read at two different heights

Ten placements of these figures, one value: the settling criterion is nine tenths in every one of them, and nothing is measured behind it. It is a multiplicative constant only inside the mechanism that has it — a bow's capture and an exciter's contact contain no criterion at all — so moving it rescales one of three clusters against two that stand still. The most-quoted number here, a factor of sixty-nine between the instrument's account and the listener's, is 5.9 at the criterion the listener's own measurements use, and the ordering an earlier essay was written about does not exist below a fifth.

7 figures
An entering part is worth 0.9 phons, in the middle of its range. A texture of 5 parts at 62 decibels each, with one more part added at the same level, tried at every semitone from C2 to C7. The vertical axis is what the addition is worth in phons, and a phon is a decibel here; the shaded strip is the difference limen for loudness, so an entry inside it is not heard as a change of level. The median entry is 0.89 phons and only 29 of 61 clear the limen — the lowest of them at A♭4, 415 hertz. The best available, at B♭6, is worth 4.2. The two lines are the two loudness models to hand: they agree everywhere above the tenor register and part company below it, where the greedy critical-band grouping reports 24 entries that make the texture quieter and the excitation pattern reports none.

A part entering is not a change of level

Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.

7 figures
Articulation is worth 2.6 phons, and nobody counts it. A passage of notes at 80 decibels, 2 to the beat, at 7 tempi and 3 articulations, scored against the same level held continuously. The variable is the fraction of each inter-onset interval that is sounding — 0.95 is a legato, 0.4 a staccato — and the vertical axis is what that costs the passage's running loudness in phons. Nothing here is anybody playing harder or softer. At 40 to the beat the span from legato to staccato is 1.47 phons; at 200 it is 2.61, because a staccato note there lasts 60 milliseconds and no longer reaches its own loudness either.

A staccato is a dynamic mark

Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.

7 figures
The top voice arrives whole and the bottom one arrives as a sine. 4 parts sounding together, each of 8 partials, with every partial tested against the summed masked threshold of every component in the texture. A filled mark is a partial the listener receives and an open one is a partial the part would have had alone and does not have here. The bass at C3 keeps 1 of 8, the tenor at C4 keeps 2 of 8, the alto at E4 keeps 5 of 8, the soprano at G5 keeps 8 of 8, every one of them at 70 decibels. Every part is at the same level and the difference is entirely where each one sits: masking spreads upward, so the part at the top of the texture has nothing above it to be masked by and the part at the bottom has everything.

The listener is given the top voice, and the bass as a sine

Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.

7 figures
The roughest chord on the page is at C1 and the roughest one heard is at A♭2. One close triad at 70 decibels through 17 registers, with its roughness computed twice: over every partial in the score, and over only the partials that stand above what the chord itself masks. The written curve rises all the way down and its maximum is the lowest register drawn, C1, which is the low-interval rule as it has always been computed here. The delivered curve turns over at A♭2 and falls to nothing below A♭1: a close triad down there is not rough, because it is not arriving as a chord — 1 of its 24 partials survives at C1 and there is almost nothing left to beat against anything.

A low chord stops being rough by stopping being a chord

Every count of audible partials until now is of a chord at middle C, and the three registers it did compare span C3 to C5 — a third of the range a chord is written in. Move the same triad down and the count collapses: 79 per cent of its partials arrive at E3 and 4 per cent at C1. So the roughest chord on the page is the lowest one and the roughest chord a listener receives is at G2, and where that maximum sits moves nearly two octaves with the dynamic.

7 figures
One attack time, three shapes, 48 ms of disagreement. Three amplitude envelopes with the same 90-millisecond attack time, which is the only quantity the published tables report. a resonator from a step rises as one minus a decaying exponential; an excitation ramping rises in a straight line; a ramp through a resonator is a raised cosine. Each is normalised so that its own 10-to-90 per cent rise takes exactly 90 milliseconds, so all three are the same measurement. The horizontal rules are the three criteria a heard moment is read off. At 6 dB below peak the three shapes put the heard moment at 28.5, 56.4, 76.3 milliseconds — a spread of 48, on one attack time, from a property nothing in the table records.

An attack time is not an attack

Eight earlier essays read a heard moment off an envelope, and every one of them used the same envelope shape without saying so: the source table records one curve for all nine of its families, and the map's own arithmetic does not carry the parameter at all. A published attack time fixes a ten-to-ninety time and nothing else. Under the two other shapes the same measurement admits, every millisecond computed so far doubles — and the constructed passage called inaudible earlier becomes three times a listener's threshold.

7 figures
A head model a few millimetres out reads every azimuth but the front. The azimuth a listener reports against the azimuth a source is at, for internal head radii from 8.22 to 9.28 centimetres against a true radius of 8.75. The delay a source produces is (r/c)(θ + sin θ) and the listener inverts it with the radius they believe they have, so their answer solves θ̂ + sin θ̂ = (r/r̂)(θ + sin θ). Every curve passes exactly through the origin: on the median plane there is no delay and therefore no error, whatever the head model is. The error grows with azimuth and is largest at the side. An internal head 5.3 millimetres too small runs out of azimuth at 82 degrees: beyond that the world is delivering a delay larger than any its owner's model can produce, and every source out there collapses onto the side.

Where a wrong head gives itself away

Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.

7 figures
A 20-decibel crescendo is 27 phons on a bass note and 20 on a high one. The same change of level, from 60 to 80 decibels, converted to loudness at each register through ISO 226's equal-loudness contours rather than at one kilohertz. The heavy curve gives each note a string spectrum, so its partials are converted in their own bands and summed; the pale one is the fundamental alone. On the spectrum-aware curve the crescendo is worth 27.3 phons at C1 and 20.3 at C7. On the fundamental alone it is 77 at C1, which is not a finding but an artefact: a 60-decibel tone at 33 hertz sits 1.8 decibels above the threshold of hearing and is very nearly nothing. The honest correction is the smaller one, and it is still a difference of 7.1 phons across the compass for a mark written in the same ink.

A subito piano is four seconds longer in the bass

Every loudness figure with time in it converts level to loudness at one kilohertz, and the equal-loudness contours say that no other frequency works that way. Joining the two sorts the published numbers into those that were about the treble and those that were not. Three move a great deal — a twenty-decibel crescendo is worth 27 phons on a bass note and 20 on a high one, and the seven seconds a subito piano takes becomes eleven and a third. Three do not move at all, and the reason they do not is the same reason in every case.

7 figures
Scored the way these figures score it, a clarinet is the worst of the six. One close triad at 70 decibels through 17 registers, drawn once for each of the 6 spectra to hand. The score is the share of every partial written, which is the quantity the register figure published. pure 100 per cent at best, string 79 per cent at best, clarinet 54 per cent at best, reed 71 per cent at best, bell 52 per cent at best, organ 72 per cent at best. A clarinet's four even partials are twenty-eight decibels below its odd ones and are inaudible beside their own neighbours before any chord is built, so counting them in the denominator makes the spectrum that survives its own masking best look like the one that survives it worst.

A clarinet keeps what a string loses

Every masker, probe, chord, line and texture until now is eight partials falling as 1/n, and it was not even an option a placement could pass. Sweeping the six spectra to hand says the clarinet is the worst of them — 54 per cent of itself at best against a string's 79 — and that answer is an artefact of the score. Counted against what each note keeps on its own, the clarinet keeps 100 per cent where the string keeps 79, because its components stand a twelfth apart rather than an octave. The missing parameter was the spectrum; the second missing parameter was the denominator.

7 figures
A displaced map is displaced by the same amount everywhere. How far a listener's heard direction is displaced, in units of the smallest angular change they could detect at that azimuth, for four constant offsets added to every interaural delay. Each curve is flat. an offset of 5 microseconds is worth 0.33 just-noticeable steps at every azimuth; an offset of 10 microseconds is worth 0.67 just-noticeable steps at every azimuth, and past 88° hands the listener a delay their own head cannot produce; an offset of 20 microseconds is worth 1.33 just-noticeable steps at every azimuth, and past 86° hands the listener a delay their own head cannot produce; an offset of 40 microseconds is worth 2.67 just-noticeable steps at every azimuth, and past 82° hands the listener a delay their own head cannot produce. The reason is exact: differentiating Woodworth's curve gives a slope proportional to (1 + cos θ), so the angular displacement a fixed offset produces carries a factor of 1/(1 + cos θ) — and so does the smallest detectable angle, so the ratio has no azimuth in it. That is the opposite of a wrong head radius, whose displacement is zero on the median plane and grows toward the side.

The error that moves straight ahead

The essay before this one found that a listener whose internal head is the wrong size makes no error at all on the median plane, and has to look hard to the side to catch it. Every head drawn here has its ears at equal radii, which makes the delay curve odd and every error a factor — and a factor cannot move a zero. Real heads are not symmetric. A constant offset of twenty microseconds displaces a listener's straight ahead by two and a quarter degrees, and it displaces every other direction by the same number of just-noticeable steps, exactly.

6 figures
The arch belongs to hearing, and the spacing only moves it. The share of a close major triad's twenty-four components that stand above what the rest of the chord masks, at 70 dB, with the root from C1 to C7, for three spectra given the same amplitude law and different frequencies: the harmonic series, a founder's bell, and a stiff string with B = 0.01. harmonic series: 0.04 at C1, peaking at 0.79 on E3, 0.42 at C7; a founder's bell: 0.04 at C1, peaking at 0.75 on C4, 0.38 at C7; a stiff string: 0.04 at C1, peaking at 0.71 on E3, 0.46 at C7. Only one of the three is a harmonic series, and all three rise out of the bass, peak in the middle of the compass and fall in the treble.

The arch belongs to hearing, not to the series

A chord delivers most of its partials in the middle of the compass and loses them in the bass and the treble, and every spectrum that showed that arch was built on whole multiples of a fundamental. Give the same amplitudes to a bell's eight modes and to a stiff string's stretched partials and the arch is still there, peaking within a major third of where the harmonic series peaks. What the spacing changes is the detail: a bell crowds its tierce and quint into a quarter of a critical band in the bass and loses them, and a stiff string's stretch buys the bass back.

5 figures
Counted over what arrives, the balanced bass is not the roughest register. The mean roughness of the I – vi – IV – V – I arrivals with each chord played as loud as the written register's, relative to the written register, counted over every partial and over the partials that stand above what the rest of the chord masks. Every partial: 70 −2 octaves, 7.48 −1 octave, 1.00 as written, 0.20 +1 octave. Delivered partials only: 6e-9 −2 octaves, 2.83 −1 octave, 1.00 as written, 0.19 +1 octave. Over every partial the lowest register is 343 times rougher than the highest; over what arrives it is the smoothest of the four, and the roughest is −1 octave, 2.8 times the written register.

A bass chord low enough to balance has already hidden its tenor

Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.

6 figures

Instruments and their design

Where the sound came from before any of the above. A stopped tube has only odd partials, a hammer at one seventh silences the seventh, and a bass string is wound because a plain one would be longer than the room.

What a 60 cm tube supports, by how its ends are closed. The first 6 modes of an open cylinder and a stopped cylinder, all of the same acoustic length. An open tube supports every whole multiple of the fundamental and reaches the second mode an octave up. A cylinder stopped at one end supports only the odd multiples and reaches its second mode 1902 cents up, which is a twelfth. Its fundamental is also an octave below the others, because it fits a quarter of a wavelength where they fit a half.

A tube that skips every other partial

Stop one end of a cylinder and half its modes vanish. That single fact about where the pressure has to be decides that a clarinet sounds hollow, that it plays an octave below its length suggests, and that it must cover nineteen semitones with fingers before it can overblow — while every other woodwind covers twelve.

4 figures
What a 60 cm tube supports, by how its ends are closed. The first 6 modes of a stopped cylinder and a cone, all of the same acoustic length. A cone supports every whole multiple of the fundamental and reaches the second mode an octave up. A cylinder stopped at one end supports only the odd multiples and reaches its second mode 1902 cents up, which is a twelfth. Its fundamental is also an octave below the others, because it fits a quarter of a wavelength where they fit a half.

A cone is not a cylinder

A saxophone has a reed at a closed end, exactly as a clarinet does, and it overblows at the octave rather than the twelfth. If the previous essay's argument were about reeds that would refute it. It is about geometry, and a cone closed at its apex has the complete harmonic series for a reason that takes one line of algebra and is genuinely surprising.

5 figures
The end correction, for a bore of radius 7.5 mm. How flat a tube sounds against what its physical length alone would predict, because the wave carries on past the opening before it turns round. The correction is 4.6 mm at every note — Levine and Schwinger's 0.6133 times the radius for an unflanged end — and the error it causes is 13 cents on a 60 cm sounding length and 52 cents on 15 cm. It is the same millimetres in both cases.

The tube ends after it ends

A wave does not turn round at the opening. It carries on into the room for about six-tenths of the bore radius and reflects there, so every tube is acoustically longer than it is. The correction is a fixed number of millimetres against a wavelength that halves every octave — a rounding error at the bottom of an instrument's range and most of a semitone at the top.

6 figures
Where each family's tone-hole lattice stops reflecting. The cutoff frequency of an open tone-hole lattice, from Benade's formula, for four woodwind geometries: clarinet 1824 Hz, oboe 2990 Hz, flute 1690 Hz, bassoon 506 Hz. Below its cutoff a note's wave turns round at the first open hole and the instrument is a tube of that length; above it the wave passes through the whole lattice and radiates from the far end, so the upper part of every note's spectrum leaves the instrument from the same place whichever note is fingered. That is what gives a family one recognisable voice across its range.

Above a certain note the holes stop working

A row of open tone holes reflects the wave and makes the tube shorter — up to a frequency. Above it the wave runs straight past the whole lattice and leaves from the bell, so the top of every note's spectrum radiates from the same place whichever note is fingered. That cutoff is computable, it differs by family, and it is most of what makes an oboe sound like an oboe.

5 figures
One register vent, 12 fingerings. Where the second mode's pressure node sits for each fingering of a stopped tube, against a single register hole drilled 12 cm from the mouthpiece. The node is a third of the way along the sounding length, so it moves every time a hole is opened, and the vent's error runs from -38 to -1 cents across the range. A perfect register system would need one hole per fingering. The number of holes actually fitted is one, and the leftover is a design decision rather than a fault.

One hole doing a dozen jobs

A register key works by forcing a pressure node where the second mode already has one, which kills the fundamental and leaves the mode above. The node sits a fixed fraction along the sounding length — and the sounding length changes with every fingering, while the hole stays where it was drilled. The leftover error is computable, and it is why the throat notes are the ones players complain about.

7 figures
A string struck at one 7th of its length. The amplitude of each partial of an ideal string excited at 0.1429 of its length. The mode shape is a sine, so a partial with a node at the excitation point cannot be set moving at all: partials 7, 14 are silent here. The envelope over the rest is one over n, a struck string's.

Where the hammer lands

Strike a string at exactly one over n and the nth partial is silent, because the hammer has landed on that mode's node and cannot move it. A piano's hammers strike between a seventh and a ninth of the way along, which puts the seventh partial — the most dissonant member of the series — at or near a null. That is a design decision made in wood, and it takes one line of trigonometry.

7 figures
The same note, hit at a middling dynamic. The spectrum of a struck string with the hammer's own contact time applied as a low-pass. Contact lasts 1.60 ms at this force, against 2.26 ms at the softest and 0.95 ms at the loudest drawn — felt is a nonlinear spring, so a harder blow is a shorter contact and a brighter note. The spectral centroid moves from partial 1.5 to partial 2.2, which is a change of timbre and not of loudness.

A hammer is not an impulse

Contact lasts a couple of milliseconds, which low-passes the note — any partial whose half-period is shorter than the contact is barely excited. Piano felt is a spring that stiffens as it compresses, so a harder blow makes the contact shorter, the corner higher and the note brighter. A loud note is not a scaled-up quiet one, and no linear model gives that.

5 figures
Helmholtz motion, bowed at 9% of the way from the bridge. Above: the string at 5 instants of one period. It is two straight lines meeting at a corner, and the corner travels round the string rather than the string swinging. Below: the resulting force on the bridge, a sawtooth whose two segments are in the ratio 0.09 to 0.91 — the bow's own position. A sawtooth contains every harmonic at exactly one over n, so the spectrum barely changes with bow position even though the waveform plainly does.

The bow makes a corner

A bowed string is not a string swinging. It is two straight lines meeting at a single sharp corner that travels round the string once per period, triggering the slip that keeps it going. The force on the bridge is therefore a sawtooth, and a sawtooth is every harmonic at exactly one over n — which is why a bowed string is the most nearly perfect harmonic series in the orchestra.

5 figures