Series

The voice — the series

13 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. The flow through the larynx, over two periods of a 110 Hz note. Volume flow against time, in Rosenberg's two-half-cosine model of the glottal pulse — a slow opening, a faster closing, and a closed phase during which no air passes at all. M1 — chest is open for 50 per cent of each period and opens 2.4 times as slowly as it closes. Nothing here is a displacement: the folds are a valve on a steady stream of air, and the flat stretches are the moments they are shut. At 110 Hz each period lasts 9.1 milliseconds, of which 4.5 is silence.

    The other instrument with a reed

    The folds do not vibrate the way a string does. They open and shut across a steady stream of air, once per period, and what leaves the larynx is a train of flow pulses with a closed phase in it. Everything said about the voice's tone is a statement about the shape of that pulse — and the shape has two numbers in it.

    part 1 · instruments
  2. Two mechanisms, the notes both of them make, and the seam. The frequency range of each laryngeal mechanism for an adult male voice, on a logarithmic axis, with the band both can produce shaded. M1 — chest runs 82–349 Hz and M2 — falsetto runs 220–698 Hz, so 799 cents of the range — 8.0 semitones — can be sung either way. The two dots inside that band are the measured signature that this is a bifurcation rather than a threshold: the change upward happens at 330 Hz and the change downward at 294 Hz, 200 cents lower. A threshold is crossed at the same place in both directions and this is not.

    Two mechanisms, and the seam between them

    Every singer has a place in the range where the voice changes character, and eight semitones of it can be produced either way. The measurement that settles what kind of a place it is takes ten seconds: the change upward happens two hundred cents higher than the change downward, and a threshold cannot do that.

    part 2 · instruments
  3. Long-term average spectra: an orchestra, playing forte against a trained operatic soloist. Each source's mean spectrum over a long passage, in decibels below its own strongest region, on a logarithmic frequency axis. An orchestra, playing forte peaks at 250 Hz and is 30 dB down by 3,150 Hz; a trained operatic soloist peaks at 250 Hz and is 11 dB down by 3,150 Hz. The shapes are the same until about 1 kHz and separate above it: at 3153 Hz the difference is 19.0 decibels, which is the largest anywhere in the range. Nothing here is about level. Both curves are drawn against their own peaks, so what is being compared is shape.

    One voice over ninety players

    A soloist heard over a full orchestra is not louder than it and could not be. What the trained voice does instead is put a peak of energy at three kilohertz, which is where the orchestra's spectrum has already fallen away and where the ear's own threshold happens to be lowest. Nineteen decibels of advantage, in a place nobody is competing for.

    part 3 · timbre
  4. A note sung at 440 Hz, drawn in cents. Deviation from the notated pitch against time, for a vibrato of ±71 cents at 6.0 cycles a second — Seashore's and Prame's measured values, which agree. The note is 142 cents wide, and the shaded bands across the middle are the syntonic comma at 21.5 cents, the Pythagorean comma at 23.5 cents, the difference limen at 440 Hz at 4.0 cents. The pitch a listener reports is near the middle of the excursion rather than at either edge — and not exactly at the middle either: because cents are logarithmic and frequency is not, the mean frequency of this trace sits 0.73 cents above the centre line.

    A note that is never at its pitch

    Eleven essays in this field argue about differences of one to twenty-four cents. An ordinary operatic vibrato is a hundred and forty cents wide and completes six excursions a second, so every one of those distinctions fits inside a single sung note several times over — and the beat rate a tuner would null passes through zero eleven times a second.

    part 4 · tuning
  5. The vowel in "hod", sung at 110 Hz. The partials of a 110 Hz note, each drawn at the amplitude the vocal tract's resonances give it. The peaks of the curve are the formants — 730 Hz and 1090 Hz — and they stay where they are when the pitch changes, because they are a property of the shape of the mouth and not of the note being sung.

    The sound a listener knows best

    A voice is recognisable across every vowel it says, across two octaves of pitch, down a bad telephone line and in a whisper where there is no pitch at all. Nothing that survives all of that can be a frequency. What survives is a ratio: the resonances of a vocal tract are set by its length, so a shorter tract multiplies every formant by the same factor, and identity is a scale on the spectral envelope rather than a position within it. Between an adult man and a child the whole pattern moves by a fifth, and the vowel does not change at all.

    part 5 · timbre
  6. The fluctuation stays; the rate goes. A unison of n voices with a spread of 15 cents, averaged over 5 draws. The depth of the amplitude fluctuation does not fall as voices are added — a choir is no steadier than a duet — but the fraction of that fluctuation in any single modulation component falls from 77 per cent at two voices to 24 at 32. Two voices make one beat and it can be counted; 16 make 120 and none of them is a rate. That is why a choir cannot be tuned by nulling anything.

    What a choir does that a soloist cannot

    Two singers on one note produce one beat and it can be counted. Sixteen produce a hundred and twenty at once, and the amplitude still fluctuates by as much as it did — a choir is no steadier than a duet. What has gone is not the fluctuation but its rate: the modulation energy that sat in a single line at two voices is spread across a band at sixteen, with no line in it. That is the choral sound, and it is also why the just-intonation drift this site measured describes only ensembles that hold their pitch still.

    part 6 · timbre
  7. How far equal temperament puts each interval's coincidence out. For each interval, the pair of partials it brings together and how many cents equal temperament mistunes that coincidence by. A fifth's third-against-second is out by 2 cents and a major third's fifth-against-fourth by 14 — so the same temperament that is inaudible on a fifth produces, at 220 hertz, a beat of 8.7 per second between two sections singing a third, with every singer in both of them perfectly in tune.

    A section against another section

    The choir has been treated as a unison, and no choir sings only unisons. Two sections an interval apart beat between partials rather than between fundamentals — the third brings the fifth partial of one against the fourth of the other — and equal temperament puts that coincidence fourteen cents out. So two sections singing a tempered third beat at nearly nine per second with every singer in both of them perfectly in tune, and the same temperament is inaudible on a fifth.

    part 7 · timbre
  8. The beat rate between two sections, second by second. The 5th partial of the lower section against the 4th of the upper, over 256 pairs of voices each sweeping 100 cents 6 times a second with its own phase. The band is the tenth to the ninetieth percentile of the instantaneous rate and the line is the median. With no vibrato the whole thing would be one flat line at 8.7 hertz, which is what was computed earlier. With it, the pair is inside the beating band 20 per cent of the time and above it for the rest — so what a listener gets is neither a beat nor a roughness but an alternation between them at the vibrato rate.

    Sixteen sweeps against sixteen

    Every intonation figure about the voice treats a singer as a frequency. A singer is a frequency being swept a hundred cents wide six times a second, and two sections singing an interval are two hundred and fifty-six pairs of sweeps. The beat rate between the partials the interval brings together stops being a number and becomes a function of time — and the pair spends four fifths of its time above the rate at which beating is beating at all.

    part 8 · intervals
  9. A vibrato flattens the dissonance curve. Each interval twice: hollow is its roughness computed at the two notes' nominal frequencies, filled is the average of its roughness over a vibrato cycle of 50 cents at 6 hertz. They are not the same number, because roughness is a curved function of the frequency difference and the average of a curve is not the curve of the average. The largest effect is at octave, where the moving average is 19.3 times the still value — an interval sitting in a deep narrow minimum is smeared out of it. The smallest is at major seventh, where it is 0.95: an interval near a maximum is smeared out of that too. Vibrato pushes every interval toward the middle, and what it takes away from the consonances is much more than what it takes away from the dissonances.

    A roughness with a rate of its own

    Every roughness figure so far computes one number for a steady spectrum. Evaluate the same sum at every instant of a vibrato and there are three numbers instead — a mean, a depth and a rate — and the mean is not the roughness of the mean frequency. On an octave it is nineteen times it, because an octave sits in a deep narrow minimum and a vibrato smears it out of one.

    part 9 · intervals
  10. The same roughness, before and after the window it has to be heard through. The instantaneous roughness of an interval under a vibrato, and the same quantity after a running average of 59 milliseconds — the time a dissonance has to last to be heard as one, which is 4 cycles of this interval's own 68-hertz fluctuation rather than a number chosen for the figure. The mean is identical to every digit, 0.1487 against 0.1487, because a running average cannot change an average — so the earlier Jensen factor of 1.0 survives the window untouched and its prediction that the window would shrink it is wrong. What the window destroys is the depth: 0.30 of the mean becomes 0.23, which is 77 per cent. The roughness a vibrato adds is heard; the fact that it is moving is mostly not.

    The mean survives the window

    A roughness that moves has a mean, a depth and a rate — all three of which a listener could only have through a temporal window. Applying the window already to hand settles which of the three survives, and the answer refutes the guess: a running average cannot change an average, so the octave's factor of nineteen stands and the movement is what goes.

    part 10 · intervals
  11. The period is still there, and it is wider. The autocorrelation of a 12-partial complex on 220 hertz, drawn twice: steady, and averaged over one cycle of a 71-cent vibrato. A vibrato moves every partial by the same number of cents, so the complex is exactly harmonic at every instant and nothing is mistuned — what moves is the period the extractor is looking for. The peak survives. It loses 6 per cent of its height above the surrounding lags and gains 11 per cent in width, because the vibrato swings the period by 0.37 milliseconds against a peak 0.90 wide. Its maximum also moves, to 2.4 cents sharp of the still tone's, which is a prediction with a sign in it.

    The pitch that does not wobble

    Three earlier essays have treated a vibrato as a modulation of roughness. The reason singers use one is what it does to the note, and there is an extractor here that turns a set of partials into a pitch and has never been asked what it does with partials that will not hold still. The period survives, at a cost that rises with the extent — and the practice stops within a hair of where the cost becomes total.

    part 11 · instruments
  12. Partials 3, 4, 12 of "hod", over one vibrato cycle. The level of three partials of a 220 hertz note on the vowel in "hod", each about its own mean, over one cycle of a vibrato of ±71 cents at 6.0 hertz. The pale curve is the frequency deviation itself, for phase reference. Partial 3 at 660 hertz swings 4.22 decibels and peaks with the frequency; Partial 4 at 880 hertz swings 0.41 decibels and peaks twice a cycle; Partial 12 at 2640 hertz swings 8.62 decibels and peaks against it. The formants of this vowel are at 730, 1090, 2440 hertz and do not move; a partial below one rises as the frequency rises and one above it falls, so the modulations of a single note run in opposite directions at the same instant.

    The partial that gets louder as it goes sharp

    Eleven earlier essays sweep a set of partials and hold their amplitudes still, and a real tract does not move with the fundamental. Put the formant account and the sweeping account together and every partial acquires an amplitude modulation at the vibrato rate: 0.07 decibels on the fundamental and 8.6 on the twelfth partial of the same note. They are in phase below a formant and anti-phase above one — not ninety degrees apart — and the whole note swings 0.82 decibels, because they cancel.

    part 12 · instruments
  13. Where a section's fluctuation stops being a beat, on 220 hertz. Two rates up the spectrum of a 220 hertz note sung by a section whose voices are spread by 15 cents. The rising line is the beat rate between a typical pair of them, which grows with the partial because a mistuning in cents is a difference in hertz that scales with frequency; it reaches the 15 hertz at which a beat stops being a beat by partial 5.5, at 1217 hertz. The flat line is the amplitude modulation the vibrato imposes through the formants, which is 6.0 hertz at every partial because the vibrato modulates every partial by the same number of cents at the same rate. The two are equal at 487 hertz. Above 1217 hertz the beating has become roughness and the only fluctuation left is the vibrato's — and that frequency is the same one an octave up, where it is partial 2.8 instead.

    The rate that does not rise with the partial

    Twelve earlier essays give every vibrato the same six hertz, and the measured spread is 5.5 to 7.5. Putting the two fluctuations a choir contains on one axis shows why the rate matters: the beating between mistuned voices rises with the partial and leaves the range a listener follows as fluctuation at 1,217 hertz, while the vibrato's own modulation is six hertz at every partial. Above that frequency a section fluctuates by vibrato alone — and if every singer had the same rate, it would barely fluctuate at all.

    part 13 · instruments

All series