Concept

Auditory scene analysis — where it appears

The process by which a listener sorts one pressure signal into separate sources, and the account of the cues that process uses. Its cues are proximity in pitch and time, common onset and common fate, and they can be set against each other in the laboratory.

Named by 10 essays across 2 fields — each of them below, with the objects they name alongside it.

A 7-semitone sequence at 120 ms a tone. Tones drawn as pitch against time, one bar per tone. The events are the same in both readings of this pattern; what changes is whether a listener assigns them to one line that leaps back and forth or to two lines that each stay put. Nothing in the drawing decides which, and nothing in the sound does either.

The ear builds objects, and sometimes offers a choice

What arrives at an ear is one pressure signal. What a listener gets is a set of separate things — a violin, a voice, a car outside. The assignment is a construction, and the clearest evidence is that it can be flipped by changing nothing but the speed: one sequence of tones is a single line when slow and two lines when fast, with a wide region in between where the listener may choose.

perception · Auditory scene
Four voices, placed by the arithmetic. The rules in force are: no parallel octaves; no parallel fifths; no voice crossing; no gap over an octave above the tenor; the leading note is not doubled; the leading note resolves, outer voices; no augmented melodic interval. The four parts are drawn lowest to highest — bass, tenor, alto, soprano. I – vi – ii – V – I in C major, realised in four voices by the cheapest set of voicings obeying 7 rules, at 16 semitones of motion in all. Every gap between adjacent voices narrower than the fission boundary of 5.2 semitones is marked, and below that boundary two parts cannot be heard as two however hard a listener tries.

A voice is a stream, and the ear decides which

Seven earlier essays have assigned voices to notes. Whether a listener follows the assignment is a separate question with laboratory numbers attached, and the numbers are unkind to it: two parts closer than about five semitones cannot be heard as two at any speed, a third of the gaps in the cheapest four-part writing are inside that limit, and in a third of chord changes the ear's own rule for continuing a line does not recover the parts as written.

perception · Voice-leading
Ode to Joy as a path. Ode to Joy plotted as 30 notes against the 8 scale degrees it uses, one column per note. Its largest melodic interval is 2 semitones and it spans 7; the mean absolute step is 1.24 semitones. Beethoven, Ninth Symphony, finale, 1824 — the theme as first stated, eight bars.

A melody is a walk, not a set

Nine essays here are about which seven of the twelve a scale takes, and every one of them describes a set. A tune is not a set; it is a path across one, and the path is nearly all small steps. That is not a matter of taste. Above about eight notes a second the ear stops being able to hold a large interval and a small one in the same line, and at sixteen the choice disappears altogether — so a fast passage is scalar because a fast passage that leaps is two pieces of music.

scales · Melody
500 Hz in one ear, 504 in the other. Two tones 4 hertz apart, one to each ear. They never meet in the air, so neither eardrum sees any modulation at all and there is no acoustic beat to hear. What changes is the phase between the ears, which advances a whole cycle every 250 milliseconds — and the direction that phase implies sweeps with it, drawn here as azimuth against time. The sweep is clipped at the edges, because the implied delay leaves the range a head can produce. A head 17.5 cm across gives at most 656 microseconds, so the phase stops naming a direction above 762 Hz.

The beat that is not in the air

Every sound this site synthesises reaches both ears identically, and that is the assumption none of its figures ever varied. Put 500 hertz in one ear and 504 in the other and nothing sums anywhere: each eardrum sees a steady sinusoid with no modulation on it at all. A listener still hears a four-per-second beat, which means the arithmetic is being done behind the ears rather than in the room. And it stops working above about a kilohertz — not where phase locking gives out at five, but where a head 17.5 centimetres across stops being able to name a direction, which is 762 hertz.

perception · Localisation
How much of each spectrum a listener can assemble into one note. Each partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there.

The spectrum that will not fuse

A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.

perception · Auditory scene
Two fusion cues, and they do not agree about a single spectrum. Each spectrum twice. Hollow is the harmonicity census — the fraction of partials near enough a whole multiple of one fundamental to fuse, which is harmonicity. Filled is the same fraction under common fate: how many partials decay at a rate within a factor of 2 of the strongest partial's. Ranked by harmonicity the order is an ideal string, a piano string, a bell, a kettledrum, a bar; ranked by common fate it is a kettledrum, a bar, an ideal string, a piano string, a bell. The two orderings are nearly reversed. An ideal string is perfect on the first cue and 20 per cent on the second, and a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose partials all die together.

The partials that do not die together

The fusion census is a still photograph: it asks whether a set of partials fits one harmonic series and has no term for time. Put time in and the ranking reverses. An ideal string is perfect on harmonicity and holds a fifth of its partials together by decay; a kettledrum — the worst spectrum in this collection for fitting a series — is the only one whose modes all die at one rate.

perception · Auditory scene
What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged.

The exchange rate nobody has

There are now two cues that disagree about how a spectrum divides, and every figure so far reports them separately because there is no principled way here to weigh one against the other. The published apparatus is a competition between grouping hypotheses with a cost per cue — and the useful thing it produces is not the winner but how much of the exchange rate the winner survives.

perception · Auditory scene
The cue that settles it. Every spectrum to hand, arbitrated by the earlier competition and then again with the onset cue added at equal weight. 3 of the 5 change their verdict, and all 3 change the same way — from splitting into two streams to staying as one: a piano string, a bell, a bar. Nothing changes the other way, because the onset cue on a struck source votes for fusion on every partial and can only ever push toward one stream. The bell is the case worth naming: its partials are wildly inharmonic and it is heard as one sound, which is a fact the harmonicity cue alone cannot produce.

The cue that settles it

Arbitrating between two grouping cues meant sweeping an exchange rate nobody could supply. The cue it had no term for at all is the one every account calls strongest, and its strength is computable: a struck string's partials start together to within a tenth of a millisecond against a threshold of twenty. Put that into the competition and three of five verdicts change, all the same way — and a bell becomes one sound.

perception · Auditory scene
A clarinet's partials, each on its own resonance. The 8 partials of the clarinet's chalumeau D that ride an impedance peak, each building toward its steady amplitude as 1 − exp(−t/τ) with τ = Q/πf from that peak's own Q. The time constants run from 11.9 milliseconds to 64.5, so the partials do not arrive at different times — they all begin the instant the reed does — and what differs is how fast each approaches its final level. The horizontal bars are how far apart the first and last are at three criteria: 5.5 ms at 10 per cent, 36.4 ms at 50 per cent, 121.0 ms at 90 per cent. A twenty-millisecond asynchrony is the threshold for hearing a partial out of a note, and this note crosses it at 32 per cent of steady amplitude — so whether a blown note's onset cue is unanimous or divided is decided entirely by how far along a partial has to be before it counts as having started.

A blown note does not start late, it starts slowly

Computing the onset cue removed a free parameter and turned out to be unanimous, and it predicted that a wind instrument would put it back, because a blown note's partials arrive over tens of milliseconds. They do — 121 on a clarinet — and it is not an asynchrony: every partial begins the instant the reed does and they differ in rate, not in time. Read at a tenth of the steady amplitude the spread is 5.5 milliseconds against a threshold of twenty, so the cue is still unanimous, and the missing number is no longer the exchange rate but the criterion.

perception · Auditory scene
Eleven partials is one partial too many. What fraction of a spectrum the harmonicity census finds fused, against how many partials it is asked to census. At ten a perfect harmonic series fuses 10 of 10 and the fundamental it finds is the right one. At eleven it fuses 5 of 11 and the fundamental jumps to exactly 2.00 — the octave above. The cause is the cap the census carries for a reason established earlier: without it a bell fuses perfectly at a fundamental nobody could hear, so the search refuses any fundamental more than about ten harmonics below the top partial. At eleven partials the first thing that cap excludes is the series' own fundamental, and the census then takes the octave and calls every odd partial inharmonic. So the number of partials and the cap are the same number, and nothing had ever said so, because every earlier figure censuses ten.

Eleven partials is one too many

Six earlier essays census exactly ten partials and no figure has ever passed another number. At eleven, the harmonicity census stops finding a perfect harmonic series' own fundamental, takes the octave above it, calls every odd partial inharmonic, and the competition cuts an ideal string in two. It is not the arbitration — the cost of a second stream was swept over a factor of fifty and every verdict came back identical — it is a cap that exists for a good reason and turns out to be the same number as the count.

perception · Auditory scene

Named alongside it

The objects these essays reach for when they reach for this one.

FusionPartialStreamingCommon fateInharmonicityFission boundaryTemporal coherenceHarmonicityOnsetTimbreAir columnAttack transient

All concepts