Perception and the listener
The quietest thing audible, and why the volume knob is a tone control
A decibel is a fact about air. A phon is a fact about a listener, and the two do not line up — the map between them bends with frequency, and it bends differently at every level. One consequence is that turning a piece of music down does not turn all of it down equally, and the amount by which it does not is a number.
One sound hides another, and it hides upward
A tone can be made completely inaudible by a second tone that is nowhere near it in frequency, and the region it disappears into is lopsided. Masking spreads up the spectrum and barely down it, and the reach grows with the masker's level — so which line in an arrangement vanishes is a prediction, not a matter of taste.
A sound hides what came before it
Masking does not stop when the masker does. A loud sound raises the threshold for about a fifth of a second after it ends, which is unremarkable, and for several milliseconds before it begins, which is not. The auditory present is a window rather than an instant, and inside the window the order of events is not the order they arrived.
The ear builds objects, and sometimes offers a choice
What arrives at an ear is one pressure signal. What a listener gets is a set of separate things — a violin, a voice, a car outside. The assignment is a construction, and the clearest evidence is that it can be flipped by changing nothing but the speed: one sequence of tones is a single line when slow and two lines when fast, with a wide region in between where the listener may choose.
Two ears, and the whole of the difference is 655 microseconds
Direction is computed from two numbers — when a sound reaches each ear and how loud it is at each — and which of the two is usable is decided by the wavelength against the width of a head. The changeover frequency is not a design choice. It falls out of 343 metres a second and 17.5 centimetres, and it is why the mechanism of hearing where something is changes halfway up the piano.
The first wavefront wins
A room sends a hundred copies of every note to a listener from a hundred directions, and the listener hears one note in one place. The mechanism that does it is brutal and simple: for the first few tens of milliseconds after a sound arrives, everything that follows is denied a vote on where it came from — even when it is louder than the original.
How small a difference is audible
Every essay here about tuning has assumed a listener who can hear the difference between two systems. The assumption has a number: about five cents in the middle of the range. It clears the two commas fourfold and it does not clear the schisma at all, which sorts the whole subject of tuning into the part that is about music and the part that is about arithmetic.
The octave that is not two to one
The octave is the one interval nobody argues about: two to one, exact, in every tradition that has one. Asked to set an octave by ear, listeners set it wide — and they do it with pure tones, which have no partials to beat against each other. Whatever is stretching the octave, it is not the stiffness of a piano string.
The chord is still major, and that is why temperament works
A major third can be seventeen cents wrong and still be a major third. That tolerance is not a failure of hearing — it is the reason the whole subject of tuning is a discussion rather than a catastrophe. Every temperament ever proposed moves intervals around inside their categories, and the one thing none of them may do is push one across a boundary.
Counting produced the hierarchy
Ask listeners how well each of the twelve notes fits after a passage in C major and the answers are not a smooth gradient. They fall into four groups with no overlap at all: the tonic, then the rest of the tonic triad, then the rest of the scale, then everything else — categories the subject had names for centuries before anybody ran the experiment.
Consonance is half learned, and this is the half
This site's founding claim is that consonance is small whole numbers. Two computable models say so and they disagree about which chords — which is already awkward. The cross-cultural evidence is worse: listeners with little exposure to Western music discriminate roughness exactly as anyone does, match octaves exactly as anyone does, and rate consonant and dissonant chords as equally pleasant.
How far away the room takes over
Direct sound falls six decibels every time the distance doubles and the reverberant field does not fall at all, so the two cross once. In a concert hall the crossing is at about five and a half metres, which is nearer than nearly every seat — so almost everybody in almost every hall is hearing the building more than the players, and the number that says so is built from two quantities already computed — a room's reverberation time and an instrument's directivity.
A return has to be remembered
A stripe four bars off the diagonal and a stripe twenty-four bars off it are the same ink and are not the same experience. Convert the lag axis to seconds, discount every comparison by how long ago it was, and the ranking of these six schemes by how repetitive they are changes — and the decay constant and the tempo turn out to enter the arithmetic as one number rather than two.
A voice is a stream, and the ear decides which
Seven earlier essays have assigned voices to notes. Whether a listener follows the assignment is a separate question with laboratory numbers attached, and the numbers are unkind to it: two parts closer than about five semitones cannot be heard as two at any speed, a third of the gaps in the cheapest four-part writing are inside that limit, and in a third of chord changes the ear's own rule for continuing a line does not recover the parts as written.
A pitch with nothing to match
Filter a click train into a band where no partial is separable from its neighbours and it still has a pitch at its repetition rate. Displace each click by a fraction of a millisecond, leaving the average rate and the long-term spectrum exactly where they were, and the pitch goes. The mechanism is reading the timing — which bounds the account endorsed here from the start.
The beat that is never sounded
A listener who has heard four bars of a groove and then hears two bars with the downbeats taken out does not move the downbeat. This site's rule set does, every time, on every pattern tried — and the direction it moves in says exactly what kind of model would be needed instead.
How late is a different note
A deviation of thirty milliseconds is expression and a deviation of two hundred is a wrong note, so there is an edge. The edges in time are arithmetic — the midpoints between the simple ratios — and the swing ratio crosses two of them as the tempo rises, at 171 and at 240 beats a minute, while the notation says triplet feel throughout.
What a tonic costs in seconds
The standard key-finding algorithm cannot be run on a pitch-class set at all — a flat histogram has no variance and the correlation is undefined. Give it durations and it answers with the parent key for all seven modes identically, and it takes between 15.8 and 30.0 per cent of the total time spent on one note before it names that note instead.
Three answers to how finely a pitch can be heard
Two notes one after the other are told apart at about four cents at A440. Whether a melodic interval is in tune is a judgement an order of magnitude coarser. And two notes held together are heard to beat at a third of a cent, because the question is answered by counting rather than by hearing pitch at all. Every equal division ever built sits between the coarsest and the finest.
How much evidence a modulation needs
Run a key-finder bar by bar over a progression that moves to the dominant at bar six. With four bars of history the answer becomes the new key at bar seven and holds. With three bars it never gets there at all, and reports E minor and B minor on the way. The window decides the lag as much as the music does.
The part of the tune that is kept
Contour survives transposition, retuning, a change of instrument and a doubling of every interval, and the usual explanation is that it is what a listener retains. That can be counted rather than assumed. A six-note melody over eight degrees carries eighteen bits; its contour carries 6.59 — not the 7.92 the number of distinct shapes suggests, because the shapes are wildly unequal — and the effective alphabet is ninety-six out of two hundred and forty-three. Each further note adds 1.28 bits of shape against three of melody, and at about nine notes a contour is specific enough to pick one tune out of a thousand.
The pitch that moves the wrong distance
Take three partials two hundred hertz apart and move every one of them up by forty. The spacing has not changed, so anything reading the pitch off how often the waveform repeats must give the same answer as before. The pitch moves to 204 — the shift divided by the harmonic number — which is what a harmonic template predicts and what listeners report. Push the shift to a hundred and a second reading overtakes the first, so there are two pitches and neither is the spacing. This is the measurement that closes the question, and it closes it by ruling out one mechanism rather than by choosing between the two that are left.
How long a note has to be
Every difference limen quoted so far is for a tone that lasts as long as the listener needs, and no note in music does. A tone of duration T occupies a band about 1/2T wide whatever the ear does with it, so at 440 hertz the quoted five-cent limen is the right number only for notes longer than 486 milliseconds. A tenth of a second gives 19.6 cents, which does not clear the syntonic comma. Most of the tuning arguments in this collection are about a quantity that only exists in long notes, and the essays that made them said so about the listener and not about the note.
The beat that is not in the air
Every sound this site synthesises reaches both ears identically, and that is the assumption none of its figures ever varied. Put 500 hertz in one ear and 504 in the other and nothing sums anywhere: each eardrum sees a steady sinusoid with no modulation on it at all. A listener still hears a four-per-second beat, which means the arithmetic is being done behind the ears rather than in the room. And it stops working above about a kilohertz — not where phase locking gives out at five, but where a head 17.5 centimetres across stops being able to name a direction, which is 762 hertz.
A chord is not as loud as its notes
The first essay on loudness said what a tone's loudness is and recorded that it had said nothing about a chord's. Here is the missing rule, and it has a musical consequence nobody would predict from it: the same three notes, at the same power, are twice as loud in the treble as in the bass — because the critical band that makes a low triad five times rougher also makes it one sound instead of three.
The boundary that barely moves
Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.
Playing louder is playing earlier
An accent has two effects on when its note is heard and neither is a timing decision. A harder-driven instrument has a shorter attack, and a criterion set by the surrounding music is crossed sooner by a bigger rise — so a twelve-decibel accent on a bowed note is heard twenty-four milliseconds early with no change whatever in when the bow was put down. It is also the measurement that tells the two competing models apart.
Where the two ears stop agreeing
A room sends both ears versions of the same sound, alike at low frequencies and increasingly unlike at high ones. Where they stop resembling each other is 980 hertz, and it is set by the 17.5 centimetres between the ears rather than by anything about the room — which is within a quarter of a frequency found earlier for a completely different reason. One minus that correlation is spaciousness, and it is computable from a room's own reverberation.
A room with directions in it
Every room until now has been a reservoir of energy that drains at a rate. That model has no directions in it at all, so it cannot say the one thing every published measure of spaciousness is about: how much of what arrives comes from the side. Mirror the source in six walls and every reflection acquires an angle and a time — and the answer to why a concert hall is narrow falls out at eighteen metres.
The spectrum that will not fuse
A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.
An echo is prevented by the crowd around it
The echogram says when every reflection arrives and how loud it is; the published echo threshold says when a reflection that late and that quiet is heard separately. Put one against the other and the rear wall of every hall anybody builds is past the threshold — a 45-metre hall puts it 210 milliseconds late and 25 decibels down against a threshold of 114. It is not heard as an echo, and what saves it is not the geometry. It is everything else arriving at the same time.
The partials that do not die together
The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.
A position and a width
Sorting a hall's arrivals into echoes and everything else settled that fusion is not a yes or a no. Everything else is not nothing: a reflection too early to be heard separately still moves the apparent source, widens it and colours it. All three come out of the same list of times, levels and angles, and none of them needed a new input.
The exchange rate nobody has
There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.
The hall through a head
Every direction computed so far is a vector from a seat to an image source, and a listener has no vectors. Run the echogram through the head computed six essays ago and two things happen: the position and width become microseconds, and a third of the room disappears — because both cues fold at ninety degrees and a reflection from behind is identical to one in front.
An equal note cannot be masked
Three earlier essays are about one instant, and forward masking lasts two hundred milliseconds — longer than a note at any brisk tempo. So a fast line should be a sequence of events hiding each other, and it is not: a note masks itself ten decibels harder than its predecessor can, at any speed. What does hide a line is dynamic contrast, and the boundary is fifteen decibels.
The cue that settles it
Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.
A page has two decibels
The account of closure called its dynamic component a corpus debt: it had no model of dynamics in a form. Loudness supplies one, and applying it answers the debt by refusing it. Adding parts to a final chord one at a time, each at the same level, moves the loudness by under two decibels and not monotonically — while a player has sixty. A texture that thins is not a diminuendo.
The turn is half the angle
A stationary head cannot tell a sound in front from the same sound behind, and an earlier essay said so at length. The turn that breaks the confusion is 0.84 degrees — exactly half the angle a source would have to move for the same listener to notice it moving, and half for a reason. In a hall the same turn does something else: the source swings at 8.9 microseconds a degree and the room swings at 2.7, so a listener who moves is separating the soloist from the reverberation as well as the front from the back.
A smaller head in the same hall
Ten earlier essays draw one head. Every parameter belonging to the room has been varied by some figure and the one belonging to the listener never has, and it is the only one whose change the detection threshold does not follow: a six-year-old in the same seat receives the same fifty-four reflections at the same instants and reads them onto an axis with seventy distinguishable positions instead of eighty-seven. The speed of sound, swept over every temperature a hall is ever at, changes nothing at all — and the reason it cannot is the reason head size can.
A blown note does not start late, it starts slowly
Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.
Eleven partials is one too many
Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.
Read at two different heights
Ten placements of these figures, one value: the settling criterion is nine tenths in every one of them, and nothing is measured behind it. It is a multiplicative constant only inside the mechanism that has it — a bow's capture and an exciter's contact contain no criterion at all — so moving it rescales one of three clusters against two that stand still. The most-quoted number here, a factor of sixty-nine between the instrument's account and the listener's, is 5.9 at the criterion the listener's own measurements use, and the ordering an earlier essay was written about does not exist below a fifth.
A part entering is not a change of level
Five earlier essays measure a sonority that is already sounding. The one thing an orchestrator actually controls is the entry, and priced at every semitone it comes to 0.89 phons — under the difference limen — with only 29 of 61 available entries clearing it and none below A♭4. What an entry is instead is an object arriving, and the reason is that its own rise time is faster than the listener's.
A staccato is a dynamic mark
Every loudness figure in this collection is of a sound that has been going on long enough, and no note in music has. Run the running-loudness model on notes with lengths in them and an articulation turns out to command 1.5 phons at a slow tempo and 8.1 at a fast one — more than the 1.8 decibels a whole texture commands, on the same page, written down in the same ink, and counted by nobody.
The listener is given the top voice, and the bass as a sine
Four earlier essays put the masker and the probe in the same voice. Put them in different voices — a four-part texture at one level — and the soprano arrives with all eight of its partials, the alto with five, the tenor with two and the bass with one. Balancing the loudness, which is the constraint a scoring is solved under, changes none of that: equal loudness is not equal spectrum and cannot be made so.
A low chord stops being rough by stopping being a chord
Every count of audible partials until now is of a chord at middle C, and the three registers it did compare span C3 to C5 — a third of the range a chord is written in. Move the same triad down and the count collapses: 79 per cent of its partials arrive at E3 and 4 per cent at C1. So the roughest chord on the page is the lowest one and the roughest chord a listener receives is at G2, and where that maximum sits moves nearly two octaves with the dynamic.
An attack time is not an attack
Eight earlier essays read a heard moment off an envelope, and every one of them used the same envelope shape without saying so: the source table records one curve for all nine of its families, and the map's own arithmetic does not carry the parameter at all. A published attack time fixes a ten-to-ninety time and nothing else. Under the two other shapes the same measurement admits, every millisecond computed so far doubles — and the constructed passage called inaudible earlier becomes three times a listener's threshold.
Where a wrong head gives itself away
Every claim so far maps a delay to a direction through one fixed geometry, and the listener acquires that map while the geometry grows under them by seventy per cent. So the map can be wrong — and the essay before this one said the error would be largest on the median plane, where the delay curve is steepest. It is exactly zero there. The steepness is in the error and in the threshold and cancels between them, which leaves a listener whose internal head is 1.3 millimetres out with one place to catch it: hard to the side, where nobody localises well.
A subito piano is four seconds longer in the bass
Every loudness figure with time in it converts level to loudness at one kilohertz, and the equal-loudness contours say that no other frequency works that way. Joining the two sorts the published numbers into those that were about the treble and those that were not. Three move a great deal — a twenty-decibel crescendo is worth 27 phons on a bass note and 20 on a high one, and the seven seconds a subito piano takes becomes eleven and a third. Three do not move at all, and the reason they do not is the same reason in every case.
A clarinet keeps what a string loses
Every masker, probe, chord, line and texture until now is eight partials falling as 1/n, and it was not even an option a placement could pass. Sweeping the six spectra to hand says the clarinet is the worst of them — 54 per cent of itself at best against a string's 79 — and that answer is an artefact of the score. Counted against what each note keeps on its own, the clarinet keeps 100 per cent where the string keeps 79, because its components stand a twelfth apart rather than an octave. The missing parameter was the spectrum; the second missing parameter was the denominator.
The error that moves straight ahead
The essay before this one found that a listener whose internal head is the wrong size makes no error at all on the median plane, and has to look hard to the side to catch it. Every head drawn here has its ears at equal radii, which makes the delay curve odd and every error a factor — and a factor cannot move a zero. Real heads are not symmetric. A constant offset of twenty microseconds displaces a listener's straight ahead by two and a quarter degrees, and it displaces every other direction by the same number of just-noticeable steps, exactly.
The arch belongs to hearing, not to the series
A chord delivers most of its partials in the middle of the compass and loses them in the bass and the treble, and every spectrum that showed that arch was built on whole multiples of a fundamental. Give the same amplitudes to a bell's eight modes and to a stiff string's stretched partials and the arch is still there, peaking within a major third of where the harmonic series peaks. What the spacing changes is the detail: a bell crowds its tierce and quint into a quarter of a critical band in the bass and loses them, and a stiff string's stretch buys the bass back.
A bass chord low enough to balance has already hidden its tenor
Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.