The ear is not a microphone
Two notes and a ratio, which is the whole of consonance
Sound two tones together and the pair either settles or does not. What decides it is the ratio of their frequencies, and the rule is that simpler ratios settle — which is two and a half thousand years old and still not quite an explanation.
Beats are arithmetic that anybody can hear
Two tones a few hertz apart swell and fade at exactly their difference. It is the most direct evidence available that the ear does sums on what reaches it, and it is how every instrument in the world gets tuned.
The note that is not there
A telephone reproduces nothing below about three hundred hertz, and a bass line down at eighty comes through it perfectly. The pitch that is heard is not a frequency present in the sound, and one nineteenth-century experiment settles what it is instead.
A third is rougher in the bass
Consonance is usually presented as a property an interval has. It is not. The same major third is muddy two octaves below middle C and clean two octaves above it, the ratio never changed, and the frequency where it stops being muddy can be solved for.
A string does everything at once
A plucked string does not vibrate at one frequency. It vibrates at all the whole-number multiples of one frequency simultaneously, and nearly everything in music theory is downstream of that fact.
The ear hears the list, not the shape
Two sounds with the same partials and different phases have completely different waveforms and sound identical. What the ear extracts is a list of frequencies and strengths, and everything else is discarded.
The shape of a note, which is most of what an instrument is
Cut the first fifty milliseconds off a recorded piano and listeners stop calling it a piano. The attack carries more identity than the steady tone it leads into, and it is the part every spectrum plot leaves out.
The ear sorts into boxes, and the boxes are the theory
Slide one note slowly upward against another and the interval between them changes continuously. What a listener reports does not. It stays a minor third, stays a minor third, and then in the space of about twenty cents becomes a major third — and nothing in the sound corresponds to the moment of the change.
Two voices that stop being two
The ban on parallel fifths is the most famous rule in Western music and it is usually taught as taste. It is not taste. At an octave the upper voice contributes no frequency the lower one did not already have, at a fifth it contributes half of them, and the number can be counted — which turns a prohibition into a measurement.
A rhythm fast enough to be a chord
Three against two is a polyrhythm. Speed the same pattern up until the events arrive faster than about twenty a second and it is a perfect fifth. The ratio never changed, the figure never changed, and the only thing that moved is the rate — which makes rhythm and pitch one continuum with a perceptual boundary across the middle of it.
The beat is inferred, and sometimes wrongly
A bar line is not in the sound. One sequence of onsets supports a reading in four, in three and in six, and which one a listener settles on comes from a small set of preference rules rather than from anything in the signal — which is why the same recording can be heard two ways and why nobody can be talked out of either.
The first fifty milliseconds
A spectrum is supposed to be what makes a trumpet a trumpet. Cut the first fifty milliseconds off a recorded note and listeners stop being able to name the instrument — while the spectrum they are hearing is unchanged. Identity is in the part of the sound that ends before the note has properly started.
The quietest thing audible, and why the volume knob is a tone control
A decibel is a fact about air. A phon is a fact about a listener, and the two do not line up — the map between them bends with frequency, and it bends differently at every level. One consequence is that turning a piece of music down does not turn all of it down equally, and the amount by which it does not is a number.
One sound hides another, and it hides upward
A tone can be made completely inaudible by a second tone that is nowhere near it in frequency, and the region it disappears into is lopsided. Masking spreads up the spectrum and barely down it, and the reach grows with the masker's level — so which line in an arrangement vanishes is a prediction, not a matter of taste.
A sound hides what came before it
Masking does not stop when the masker does. A loud sound raises the threshold for about a fifth of a second after it ends, which is unremarkable, and for several milliseconds before it begins, which is not. The auditory present is a window rather than an instant, and inside the window the order of events is not the order they arrived.
What makes two partials one note
A note is a stack of ten or twenty simultaneous tones and is heard as one thing. The obvious explanation is that they are whole-number multiples of a fundamental — and the obvious explanation is not sufficient. Mistune one partial by three per cent and it leaves the note; give a perfectly harmonic partial a thirty-millisecond head start and it leaves too. Shared behaviour beats arithmetic.
The ear makes its own sound, and it is not the missing fundamental
Play two loud tones and a third pitch appears that is in neither of them. The ear is not a passive analyser: it is nonlinear, it generates frequencies of its own, and it emits sound back out of the ear canal. None of which explains the missing fundamental — the products land in the wrong place, and finding out where they land is the experiment that made the residue theory necessary.
A dissonance has to last
Roughness is a fluctuation, and a fluctuation needs cycles. A minor second at the bottom of a cello fluctuates thirty-three times a second, so a semiquaver holds four of them and a demisemiquaver holds two — which turns the counterpoint rule that a dissonance may pass if it is short into a number, and puts that number at about seventy milliseconds through most of the range, once the pairs beating too slowly to be roughness at all are taken out of the average.
A voice is a stream, and the ear decides which
Seven earlier essays have assigned voices to notes. Whether a listener follows the assignment is a separate question with laboratory numbers attached, and the numbers are unkind to it: two parts closer than about five semitones cannot be heard as two at any speed, a third of the gaps in the cheapest four-part writing are inside that limit, and in a third of chord changes the ear's own rule for continuing a line does not recover the parts as written.
Which harmonics carry the pitch
A missing fundamental is inferred from a pattern, and a pattern has to be legible before it can be matched. Counting how many harmonics of a note land in separate auditory filters prices the inference — and the answer at the bottom of a bass guitar's range is none of them.
A pitch with nothing to match
Filter a click train into a band where no partial is separable from its neighbours and it still has a pitch at its repetition rate. Displace each click by a fraction of a millisecond, leaving the average rate and the long-term spectrum exactly where they were, and the pitch goes. The mechanism is reading the timing — which bounds the account endorsed here from the start.
The same chord is harsher when it is louder
Every roughness number so far was computed at a level nobody stated. Roughness is the product of two partial amplitudes, so it is quadratic in pressure, while loudness is compressive — which makes a minor third at middle C thirty-two thousand times rougher at fortissimo than at pianissimo and only twenty-five times louder. A chord has no single consonance to report.
How late is a different note
A deviation of thirty milliseconds is expression and a deviation of two hundred is a wrong note, so there is an edge. The edges in time are arithmetic — the midpoints between the simple ratios — and the swing ratio crosses two of them as the tempo rises, at 171 and at 240 beats a minute, while the notation says triplet feel throughout.
The first eighty milliseconds are a different room
Draw a line across a room's decay and the energy on either side is two opposite verdicts about one building — clarity before it, reverberation after. The line is a property of the ear, not of the room, and putting it at 50 milliseconds and at 80 makes the two published design targets fall out — a room for speech at one second, a room for music at 1.6.
A melody is a walk, not a set
Nine essays here are about which seven of the twelve a scale takes, and every one of them describes a set. A tune is not a set; it is a path across one, and the path is nearly all small steps. That is not a matter of taste. Above about eight notes a second the ear stops being able to hold a large interval and a small one in the same line, and at sixteen the choice disappears altogether — so a fast passage is scalar because a fast passage that leaps is two pieces of music.
The sound a listener knows best
A voice is recognisable across every vowel it says, across two octaves of pitch, down a bad telephone line and in a whisper where there is no pitch at all. Nothing that survives all of that can be a frequency. What survives is a ratio: the resonances of a vocal tract are set by its length, so a shorter tract multiplies every formant by the same factor, and identity is a scale on the spectral envelope rather than a position within it. Between an adult man and a child the whole pattern moves by a fifth, and the vowel does not change at all.
What a choir does that a soloist cannot
Two singers on one note produce one beat and it can be counted. Sixteen produce a hundred and twenty at once, and the amplitude still fluctuates by as much as it did — a choir is no steadier than a duet. What has gone is not the fluctuation but its rate: the modulation energy that sat in a single line at two voices is spread across a band at sixteen, with no line in it. That is the choral sound, and it is also why the just-intonation drift this site measured describes only ensembles that hold their pitch still.
The same distance, under two names
Four hundred cents is a major third or a diminished fourth, and on a keyboard nothing in the sound distinguishes them. An earlier essay was about the boundary between two categories; this is about two categories at one acoustic value, and the surprise is where the ambiguity comes from. In quarter-comma meantone a major third is 386 cents and a diminished fourth is 427 — two names, two pitches, forty-one cents apart. Equal temperament collapsed them, and what a listener now supplies from context used to be in the sound.
The mark that is not a level
There are six of them, they carry no units, and a performer has to turn one into a number before it means anything. What they instruct is not loudness. On a struck string a harder blow shortens the hammer's contact from 2.26 milliseconds to 0.95, which moves the first null of its own pulse from the third partial to the sixth: the partials between those are not quieter at pianissimo, they are gone. A fortissimo is a different sound, and the page has one word for both things it changes.
The pitch that moves the wrong distance
Take three partials two hundred hertz apart and move every one of them up by forty. The spacing has not changed, so anything reading the pitch off how often the waveform repeats must give the same answer as before. The pitch moves to 204 — the shift divided by the harmonic number — which is what a harmonic template predicts and what listeners report. Push the shift to a hundred and a second reading overtakes the first, so there are two pitches and neither is the spacing. This is the measurement that closes the question, and it closes it by ruling out one mechanism rather than by choosing between the two that are left.
The series has three tops
How far up the harmonic series can an ear go? The question has three answers and they are an order of magnitude apart. Consecutive partials stop being separately resolvable somewhere around the eighth, and where exactly depends on the fundamental. They stop being a semitone apart at the seventeenth, at every fundamental, because the ratio does not know what it is measured in. And they stop being distinguishable in pitch at all between the thirty-fourth and the hundred and fortieth. Every claim about where the series runs out is a claim about which of the three was meant.
How long a note has to be
Every difference limen quoted so far is for a tone that lasts as long as the listener needs, and no note in music does. A tone of duration T occupies a band about 1/2T wide whatever the ear does with it, so at 440 hertz the quoted five-cent limen is the right number only for notes longer than 486 milliseconds. A tenth of a second gives 19.6 cents, which does not clear the syntonic comma. Most of the tuning arguments in this collection are about a quantity that only exists in long notes, and the essays that made them said so about the listener and not about the note.
The beat that is not in the air
Every sound this site synthesises reaches both ears identically, and that is the assumption none of its figures ever varied. Put 500 hertz in one ear and 504 in the other and nothing sums anywhere: each eardrum sees a steady sinusoid with no modulation on it at all. A listener still hears a four-per-second beat, which means the arithmetic is being done behind the ears rather than in the room. And it stops working above about a kilohertz — not where phase locking gives out at five, but where a head 17.5 centimetres across stops being able to name a direction, which is 762 hertz.
A chord is not as loud as its notes
The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.
A note is heard after it starts
Every rhythm essay until now has treated a note's onset as the moment it happens. It is not: the instant a listener aligns a note with a beat is later than its physical start by an amount the note's own attack decides, and for a sung or bowed note that amount is about thirty milliseconds — the size of the whole quantity six essays on microtiming set out to measure.
Playing louder is playing earlier
An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.
Loud is relative, and it comes down slowly
The account of loudness had a model of a moment and the account of closure asked it for a model of a form. The published one exists and its content is a pair of numbers that are not the same: a listener's running impression of how loud the music is rises to meet a step in a fifth of a second and takes seven seconds to come back down. A twenty-decibel crescendo spread over eight seconds therefore buys almost no contrast at all, and the same twenty decibels taken as a step buys a factor of two.
Every partial beats at its own rate
Five earlier essays have drawn one beat rate per figure, and every one of them is the rate between two fundamentals. Two real notes beat between all of their partials at once, the k-th pair beats k times as fast, and somewhere up the spectrum the rate passes the point at which a beat stops being a beat — so a chorused note is a beat at the bottom of itself and a roughness at the top, simultaneously, with a crossover partial that is arithmetic.
A section against another section
The choir has been treated as a unison, and no choir sings only unisons. Two sections an interval apart beat between partials rather than between fundamentals — the third brings the fifth partial of one against the fourth of the other — and equal temperament puts that coincidence fourteen cents out. So two sections singing a tempered third beat at nearly nine per second with every singer in both of them perfectly in tune, and the same temperament is inaudible on a fifth.
Where the two ears stop agreeing
A room sends both ears versions of the same sound, alike at low frequencies and increasingly unlike at high ones. Where they stop resembling each other is 980 hertz, and it is set by the 17.5 centimetres between the ears rather than by anything about the room — which is within a quarter of a frequency found earlier for a completely different reason. One minus that correlation is spaciousness, and it is computable from a room's own reverberation.
Sixteen sweeps against sixteen
Every intonation figure about the voice treats a singer as a frequency. A singer is a frequency being swept a hundred cents wide six times a second, and two sections singing an interval are two hundred and fifty-six pairs of sweeps. The beat rate between the partials the interval brings together stops being a number and becomes a function of time — and the pair spends four fifths of its time above the rate at which beating is beating at all.
A room with directions in it
Every room until now has been a reservoir of energy that drains at a rate. That model has no directions in it at all, so it cannot say the one thing every published measure of spaciousness is about: how much of what arrives comes from the side. Mirror the source in six walls and every reflection acquires an angle and a time — and the answer to why a concert hall is narrow falls out at eighteen metres.
The spectrum that will not fuse
A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.
The dynamics are in the score already
Count the parts in each bar, realise them in their ranges, put every partial in its critical band, sum the loudnesses and run the result through the two smoothers built earlier. What comes out is a dynamic curve for a piece with no performance in it anywhere — and it says that doubling the number of parts inside a fixed register adds three decibels of power and about one of loudness, because the extra parts land in bands that were already occupied. Let the register widen with the parts and the same arithmetic gives eight phon, which is what a tutti actually is.
The ranking survives the dynamic and the chord does not
Roughness is quadratic in pressure and loudness is compressive, so sixty decibels multiply a chord's roughness by a million and its loudness by ninety-six. Roughness per sone therefore rises ten thousandfold between a pianissimo and a fortissimo of the same four notes — and yet the ranking of which doubling is smoothest, over four hundred and eighty voicings, does not move by a single place.
An echo is prevented by the crowd around it
The echogram says when every reflection arrives and how loud it is; the published echo threshold says when a reflection that late and that quiet is heard separately. Put one against the other and the rear wall of every hall anybody builds is past the threshold — a 45-metre hall puts it 210 milliseconds late and 25 decibels down against a threshold of 114. It is not heard as an echo, and what saves it is not the geometry. It is everything else arriving at the same time.
The partials that do not die together
The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.
A roughness with a rate of its own
Every roughness figure so far computes one number for a steady spectrum. Evaluate the same sum at every instant of a vibrato and there are three numbers instead — a mean, a depth and a rate — and the mean is not the roughness of the mean frequency. On an octave it is nineteen times it, because an octave sits in a deep narrow minimum and a vibrato smears it out of one.
A beat has a depth, and six essays held it at one
Every beat figure so far adds two tones of equal amplitude, which is the single ratio at which the trough of a beat is a true null — and a null is what a tuner is actually listening for. Vary the ratio and the picture changes: at two to one the dip is nine and a half decibels, at ten to one it is under two, and the interval that gives the shallowest null of all is the octave.
The mean survives the window
A roughness that moves has a mean, a depth and a rate — all three of which a listener could only have through a temporal window. Applying the window already to hand settles which of the three survives, and the answer refutes the guess: a running average cannot change an average, so the octave's factor of nineteen stands and the movement is what goes.
A position and a width
Sorting a hall's arrivals into echoes and everything else settled that fusion is not a yes or a no. Everything else is not nothing: a reflection too early to be heard separately still moves the apparent source, widens it and colours it. All three come out of the same list of times, levels and angles, and none of them needed a new input.
The exchange rate nobody has
There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.
A loud chord is a smaller chord
Two earlier essays hold the level fixed, and the level decides how much of a chord a listener is given. At thirty decibels twenty of a triad's twenty-four partials stand above what the rest of it masks; at a hundred, nine do. Every roughness figure until now counts partials that are in the score, and a partial the chord masks is not a partial the listener has.
The pitch that does not wobble
Three earlier essays have treated a vibrato as a modulation of roughness. The reason singers use one is what it does to the note, and there is an extractor here that turns a set of partials into a pitch and has never been asked what it does with partials that will not hold still. The period survives, at a cost that rises with the extent — and the practice stops within a hair of where the cost becomes total.
Two players on one note
Six essays have put one instrument on each note of a chord, and the commonest thing an orchestrator actually does is put two on the same note. Two independent sources add in power, so the composite is neither of them — except that it nearly always is one of them, because the level at which ownership changes hands is rarely at zero. And a unison ten cents out is rougher than a major third dead in tune.
The dissonance arrives and the dynamic does not
A scoring decides two things at once and both of them have to be integrated by a listener before they exist. The loudness smoother's release is two seconds and the roughness window is thirty-seven milliseconds, and that ratio of fifty decides which of the two survives at the pace music is actually played. Nothing anybody performs is fast enough to blur a dissonance, and a great deal of it is fast enough to average a dynamic.
The hall through a head
Every direction computed so far is a vector from a seat to an image source, and a listener has no vectors. Run the echogram through the head computed six essays ago and two things happen: the position and width become microseconds, and a third of the room disappears — because both cues fold at ninety degrees and a reflection from behind is identical to one in front.
An equal note cannot be masked
Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.
The cue that settles it
Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.
A page has two decibels
The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.
The turn is half the angle
A stationary head cannot tell a sound in front from the same sound behind, and an earlier essay said so at length. The turn that breaks the confusion is 0.84 degrees — exactly half the angle a source would have to move for the same listener to notice it moving, and half for a reason. In a hall the same turn does something else: the source swings at 8.9 microseconds a degree and the room swings at 2.7, so a listener who moves is separating the soloist from the reverberation as well as the front from the back.
A smaller head in the same hall
Ten earlier essays draw one head. Every parameter belonging to the room has been varied by some figure and the one belonging to the listener never has, and it is the only one whose change the detection threshold does not follow: a six-year-old in the same seat receives the same fifty-four reflections at the same instants and reads them onto an axis with seventy distinguishable positions instead of eighty-seven. The speed of sound, swept over every temperature a hall is ever at, changes nothing at all — and the reason it cannot is the reason head size can.
A blown note does not start late, it starts slowly
Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.
Eleven partials is one too many
Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.
A part entering is not a change of level
Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.
A staccato is a dynamic mark
Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.
The listener is given the top voice, and the bass as a sine
Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.
A low chord stops being rough by stopping being a chord
Every count of audible partials until now is of a chord at middle C, and the three registers it did compare span C3 to C5 — a third of the range a chord is written in. Move the same triad down and the count collapses: 79 per cent of its partials arrive at E3 and 4 per cent at C1. So the roughest chord on the page is the lowest one and the roughest chord a listener receives is at G2, and where that maximum sits moves nearly two octaves with the dynamic.
A section has a loudest member, not a colour
Two players on one note have a balance at which the composite belongs to neither, and that is what blending means. Three should have three such balances and no reason for them to agree — a trio with a rock-paper-scissors ownership would have no strongest member at all. Twenty trios, sixty pairwise comparisons, and not one disagreement: the possibility is real, arbitrary spectra do it once in twenty, and instruments never do.
The blend table has a row for every note
Eight earlier essays sound their instruments at one note, and one of them says why that cannot be innocent: every filter here is fixed in frequency and the fundamental is not. Swept over four octaves, the number of pairs that blend doubles from six to twelve, the ranking turns over rather than shuffles — seventy of a hundred and five comparisons swap — and a clarinet with an oboe goes from the best pair in the collection to the eleventh.
A soft chord has to fade in
Forward masking sits between the two integration times already in play — two hundred milliseconds against a thirty-seven millisecond roughness window and a two-second loudness release — and it was owed as the term that might eat the dissonance contrast. It does not. It widens it, by three per cent at a chorale's pace and fifty-nine at four chords a second, because it can only ever remove partials and it can only reach the chord a dynamic has already made quiet. What it does instead is stranger: one chord in the passage is entirely inaudible for its first twelve milliseconds and takes a quarter of a second to arrive whole.
Three beats at most, and only in the middle of the keyboard
A mistuned octave on a real piano makes eight beats at once, and a listener attending to one of them is doing something that has a threshold. Two thresholds, in fact — a rate and a place — and once both are applied the eight become four at A3, one at A1 and one at A5. Every interval a tuner sets goes to zero at both ends of the compass and peaks at eight countable beats in the octave the bearing is laid in.
The top that falls while the note lasts
Four earlier essays have asked where the harmonic series stops, and all four answered with a number computed for a tone that never ends. A struck C3 has 152 audible partials at the strike and eight after two-thirds of a second, so the ear's own resolution limit governs the first ten per cent of the note and the decay governs the rest. Playing ten decibels louder buys the ear a further tenth of a second, and doubling its reign would take fifty-eight.
The partial that gets louder as it goes sharp
Eleven earlier essays sweep a set of partials and hold their amplitudes still, and a real tract does not move with the fundamental. Put the formant account and the sweeping account together and every partial acquires an amplitude modulation at the vibrato rate: 0.07 decibels on the fundamental and 8.6 on the twelfth partial of the same note. They are in phase below a formant and anti-phase above one — not ninety degrees apart — and the whole note swings 0.82 decibels, because they cancel.
The rate that does not rise with the partial
Twelve earlier essays give every vibrato the same six hertz, and the measured spread is 5.5 to 7.5. Putting the two fluctuations a choir contains on one axis shows why the rate matters: the beating between mistuned voices rises with the partial and leaves the range a listener follows as fluctuation at 1,217 hertz, while the vibrato's own modulation is six hertz at every partial. Above that frequency a section fluctuates by vibrato alone — and if every singer had the same rate, it would barely fluctuate at all.
Twelve violins are more punctual than one
Every essay until now treats a part as one player, and an orchestral part is a dozen. Sectioning does two things at once and only one of them was expected: it pulls the part's heard moment forward, by four milliseconds against a map spanning twenty-six, and it makes the part's arrival more accurate by very nearly the root of the number of players. So the map of required leads applies to an orchestra better than it applies to a quartet, and the case where it fails is three trumpets rather than fourteen violins.
Where a wrong head gives itself away
Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.
A subito piano is four seconds longer in the bass
Every loudness figure with time in it converts level to loudness at one kilohertz, and the equal-loudness contours say that no other frequency works that way. Joining the two sorts the published numbers into those that were about the treble and those that were not. Three move a great deal — a twenty-decibel crescendo is worth 27 phons on a bass note and 20 on a high one, and the seven seconds a subito piano takes becomes eleven and a third. Three do not move at all, and the reason they do not is the same reason in every case.
A clarinet keeps what a string loses
Every masker, probe, chord, line and texture until now is eight partials falling as 1/n, and it was not even an option a placement could pass. Sweeping the six spectra to hand says the clarinet is the worst of them — 54 per cent of itself at best against a string's 79 — and that answer is an artefact of the score. Counted against what each note keeps on its own, the clarinet keeps 100 per cent where the string keeps 79, because its components stand a twelfth apart rather than an octave. The missing parameter was the spectrum; the second missing parameter was the denominator.
A fifth on a piano is not a fifth a second later
Nine essays draw one note in silence, and a listener is given a texture. Put two struck notes an interval apart and consonance comes apart into two quantities that had agreed while nothing moved: the partial coincidence that names the interval outlives the strike in exactly the order common-practice theory ranks its intervals, and the roughness that scores it reorders itself inside a third of a second, with the fifth overtaken by four intervals every treatise calls harsher.
The blend arrives before the note does
Nine essays on spectrum draw a steady state, and the strongest cue that two instruments are two instruments is that they do not start together. Two envelopes rising at different rates turn out to be a balance — the same dial an earlier essay swept — so the attack is that dial moved by the clock, and its whole travel is fixed at twenty times the log of the two attack times. It is six decibels for a clarinet with a violin against a crossing twelve to twenty-two decibels out, so one pair in ten changes hands during its own attack, and which one depends on a convention rather than on the instruments.
Every member of a beat family is the same depth
A mistuned octave's beats have been counted on their rate and their place, with the depth recorded as the thing left out and a prediction that the shallow upper members would take the count from four to two. The depth turns out not to fall at all: on any power-law spectrum every member of a family has exactly the modulation index its interval's own ratio gives, at every register, on every wire. The count does fall to two, and the thing that takes it there is the criterion that essay was already using.
An entrance is a change of colour
Eight essays on orchestration move the assignment and hold the ensemble still, and a score does the opposite: it brings players in and takes them out. Loudness is a sum over parts and roughness is a sum over pairs, so the player who joins adds one term to the first and one to the second for everybody already there. What the entrance is worth in phons falls by a factor of forty-eight across the range an ensemble spans and crosses the difference limen at five players; what it is worth in roughness rises by twelve, and by a further factor of ten for every ten decibels the passage is played at.
The release is on the wrong side
Whether the loudness model's two-second release makes an entrance inaudible has the answer no, for a reason the question did not anticipate. The smoother is asymmetric — ninety-nine milliseconds going up and two seconds coming down — so a rise is tracked twenty times faster than a fall, and an entrance is received promptly by every one of a listener's three readings. The colour of it arrives first, at twenty-five milliseconds against ninety, and the reading that moves with the ensemble is the one nobody would have picked.
The chord that has room for an entrance
Three essays have made the ensemble something a score can change and none of them has asked when. The ensemble already sounding puts a masked threshold over whatever register an entering part takes, and that threshold is set by the voicing rather than by the dynamic — so the five chords of one passage differ by 8.7 decibels in how much of an entering oboe survives them, and the quietest chord of the five is the worst place in the passage to bring somebody in. Swept over the entrant's own pitch, the choice of moment is worth as much as the choice of register.
One fluctuation or two
Whether two members of a beat family are one thing or two was decided by asking whether their rates differ by a factor of two, and the factor was written down as a stand-in for a modulation filterbank nobody had run. Run, the bank gives 1.618 — the golden section, and not by accident, since a channel of quality one has its half-power points there. The stand-in was conservative rather than optimistic, and the recomputed counts do not change at a single register, because the criterion was never what was binding.
How hard the note was struck
The auditory filter is not a fixed shape: its lower skirt shallows by about 38 per cent of its reference value every ten decibels, so a loud note is analysed through a wider filter than a quiet one. Every share computed so far was quoted at a moderate level, and a tuner does not strike moderately. Recomputed, the mean share of a mistuned octave's filter falls from 0.55 at forty decibels to 0.07 at seventy, and the count of separable beats goes from two to none — which is a prediction too strong to be right, and the way it fails is the useful part.
Counted in the decay, or not at all
A mistuned octave struck hard delivers no countable beat at the strike, and the reconciliation offered for that was that a tuner listens to the decay. Computed through a real decay it fails on its own terms: the partials fall silent before the filter has narrowed enough to separate them. It succeeds only when the filter is broadened by the level inside it, which is what the published parameterisation was fitted against — and then the window opens at a twelfth of the note's life and shuts at a quarter.
The third sound magnifies cents, not hertz
Tartini's third sound is said to be a few cents off on a tempered interval. It is sixty-seven cents off on a major third and ninety-six on a minor third, because a difference tone moves p/(p − q) times as many cents as the interval p:q that made it. In hertz it moves exactly as far as the note that moved, and no further — so what the magnifier is worth is the ear's finer resolution at the low frequency where the product lands, which is a factor of two for a long note and nothing at all for a short one.
The tone on the root changes hands at the fifth
Every combination tone of a just interval is a harmonic of a fundamental neither note contains, and which harmonic is fixed by the ratio. The difference tone lands on that fundamental for every interval up to the fifth; the cubic product lands on it for the fifth and every interval above except the minor sixth. So the loud product names the root of a narrow interval and the quiet one names the root of a wide one — and a just major seventh's difference tone is a note seven harmonics up that no keyboard has.
A major triad's combination tones are its own notes
Play a just major triad of pure tones and two of the ear's cubic products land exactly on its root and its fifth. The reason is a condition rather than a coincidence — a chord's cubic products fall on its own notes when its middle note is the mean of the outer two in hertz — and it holds for the major triad in root position and in the six-four, and for no minor triad in any position or tuning. Equal temperament misses the landing by one number, 5.6 hertz on middle C, which is a beat that belongs to no pair of notes in the chord.
The bass line under a passage in thirds
A major scale harmonised in parallel thirds gives the ear a difference tone under every pair, and in five-limit just intonation those tones are a diatonic bass line — C, A, C, F, G, F, G, C — made of the scale's own notes. Tempered, the same line moves only by whole tones, a neutral third and a fourth stretched to 650 cents, and wobbles by up to 84 cents from note to note. In sixths the bass is drawn by the other product, because the cubic product of a pair is the difference tone of the same pair inverted.
A doubled pizzicato gives its note away early
The attack turns the balance between two players on one note by a few decibels and stops. A pluck does not stop — every partial of it decays, so a pizzicato doubled by a held instrument walks the balance for the whole note, and the expectation was a handover as slow as the decay. It is fast. A one-second pizzicato over a flute loses its note in 70 milliseconds, while it is still five decibels the louder, because what hands the note over is its upper partials going first. A uniform fade would have kept it eight times as long.
Room is used up by whoever enters first
The chord with the most room for a part entering alone is a fact about that chord. It stops being a fact the moment two parts want it, because each part that comes in raises the mask over everybody after it. Given six parts waiting to enter a five-chord passage, choosing each part's moment the way one part's moment is chosen puts three of them into the same chord and lands in the bottom fifth of all 15,625 schedules. Placing them one at a time does no better. The schedule under which the least audible entrance is heard best is unique, and it brings the low and middle parts in while the texture is thin and holds the three highest back for the last three chords — because a high part keeps its room over a full texture and a middle part does not.
A rough arrival is rough because of its spacing
The pair the expectation essays report for every chord — how surprising it was, how rough its voicing is — has no level in it. Putting level back in answers the question it left open, and not the way it was framed. At one written dynamic the arrivals keep their order from 40 to 90 dB at three registers of four, and the bass stays 8.6 times rougher than the treble. Made equally loud, the bass has to be played 12.8 dB harder, and it is 343 times rougher: level does not explain the register's roughness away, it multiplies it.
A combination-tone bass needs a forte
A scale in just thirds draws a diatonic bass line through its difference tones, and in sixths the cubic product draws one. Given the two published level laws, with their constants swept, the thirds' bass is not heard at all below primaries of about 66 dB and is heard whole only from 71 to 81. The cubic products are a different kind of object: the primaries mask them decibel for decibel as they rise, so no dynamic changes whether they are heard. Most of the thirds' inner line never is, and the sixths' bass needs a forte and a gentle law.
A string that decays twice is counted early
A mistuned octave's beats were found countable only between a twelfth and a quarter of a note's life, on a note decaying once. A piano string decays twice, a fast prompt sound over a slow aftersound, and the prediction was that this would open the count sooner and close it later. It opens sooner — at a seventh of a second rather than a second — and closes exactly where a single decay struck twenty decibels softer closes, so at 80 dB it holds 8.8 beats instead of 11.8. The count now rises with the strike to 90 dB, and a tuner who strikes hard is right.
A room keeps a pizzicato from giving its note away
Doubled by a flute, a one-second pizzicato loses its note in 70 milliseconds dry, because its upper partials go first. The question left open was whether a room, whose reverberation keeps those partials alive, gives the note back afterwards. It does not give it back. It stops the note going: ten metres into a concert hall the pluck keeps it for 506 milliseconds, in a stone church for 814, and the room's own uneven decay takes back between a quarter and two fifths of that. In a room the loss law that decided everything dry matters a tenth as much, because the room's decay has become the clock.
A bow holds the number a blow hides
Two struck notes with different loss laws are identical at the strike and identical at the end, which is why separating them at all meant looking in the middle. Drive the same two strings continuously and the loss law stops being a rate and becomes a slope: each partial settles at its drive over its own loss, so the exponent adds to the source's roll-off and sits in the spectrum for as long as the bow moves. It is 18.7 decibels of separation available from the first instant, against 41.3 that a blow delivers after four tenths of a second and then takes away.
Three harmonics of the bass arrive before the bass
Every product priced until now was between two pure tones, and nothing that plays thirds is pure. Give each note a spectrum and the ear receives every pair of partials — and for a just interval p:q every one of their products is an exact multiple of the same absent fundamental. That crowd lands where the threshold of hearing is tens of decibels cheaper, so it names the bass at 71 decibels where the component at the bass's own frequency needs 74, and at 75 against 85 an octave lower. Tempered, the crowd still forms and names a note seventy cents flat.
The ghost bass drops a twelfth at a forte
Both crowds arrive at once and every member of both is a multiple of the same absent fundamental, so a listener is never given a choice between them — only a different subset of one harmonic series at every dynamic. Softly, the subset is an exact gapless series on three times the fundamental. Loudly, the difference tones fill in the low harmonics and no template on the higher note survives them. Between 62 and 72 decibels, depending on the interval, the note the crowd names falls by an octave or a twelfth, and the two qualities of third cross at different levels.
The error that moves straight ahead
The essay before this one found that a listener whose internal head is the wrong size makes no error at all on the median plane, and has to look hard to the side to catch it. Every head drawn here has its ears at equal radii, which makes the delay curve odd and every error a factor — and a factor cannot move a zero. Real heads are not symmetric. A constant offset of twenty microseconds displaces a listener's straight ahead by two and a quarter degrees, and it displaces every other direction by the same number of just-noticeable steps, exactly.
Counting beats moves the price of a chord, not the tuning
The search that found a guitar piece's best tuning counted every cent of error alike, and the obvious objection is that a cent of a third beats faster than a cent of an octave. Counted in beats instead, every piece gets exactly the same tuning back. What changes is how much of the piece the abandoned chord has to hold before it is rescued — thirty beats instead of forty-four, arriving in three steps instead of one.
A louder final chord is a deeper silence and a brighter sound
A final chord marked a step louder than the passage was supposed to stand above a listener's running impression for as long as it sounded, since the impression can climb no higher than the chord. It climbs exactly that high, and the stand closes in half a second as it always did. What a louder mark actually buys is depth — about seven phons, the same depth a second of silence buys — and a spectrum whose balance point sits most of a whole tone higher, which, unlike the stand, lasts for the whole chord.
An exit is worth nothing until the tutti is given up
Six parts entering a five-chord passage have a best schedule when each enters once and stays, and letting parts leave and come back was supposed to improve it. Searched over every set of sounding parts at every chord, it improves it by exactly nothing, with or without the chord before still masking — as long as all six must be playing at the end. Let one part be missing from the final chord and the weakest entrance gains 1.4 decibels; let two be missing and it gains 1.9, by a relay in which the parts with least room come in, are heard for one chord, and give way.
The played notes already name the ghost bass
The products of a just third's partials, fitted on their own, name a note a twelfth above the bass when the interval is soft and drop to the bass when it is loud. Put the two played notes back beside them and the drop disappears: the notes and their products name the bass at every dynamic, because the notes' own partials are harmonics of it already. What the dynamic changes is not which note is implied but how complete its harmonic series is — nine holes from the notes alone, four when soft, none when loud.
Only the player hears a staccato end
A damper stops a string in a seventh of a second, and in a hall the room goes on for two. A listener hears both, mixed in proportion to how close they sit, and the question was at what distance the short part stops mattering. The answer is closer than any seat. A damped note's twenty-decibel fall has doubled in length by a seventh of a hall's critical distance — 77 centimetres in a two-second concert hall — and by a quarter of it in a jazz club. The end of a staccato is something the pianist hears and the front row does not.
The arch belongs to hearing, not to the series
A chord delivers most of its partials in the middle of the compass and loses them in the bass and the treble, and every spectrum that showed that arch was built on whole multiples of a fundamental. Give the same amplitudes to a bell's eight modes and to a stiff string's stretched partials and the arch is still there, peaking within a major third of where the harmonic series peaks. What the spacing changes is the detail: a bell crowds its tierce and quint into a quarter of a critical band in the bass and loses them, and a stiff string's stretch buys the bass back.
A bass chord low enough to balance has already hidden its tenor
Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.