The ear builds objects, and sometimes offers a choice
A room with a violin, a voice and a car outside contains one sound. The air at any point in it has one pressure at any instant, and an eardrum can report exactly one number. That a listener in that room experiences three separate things is not a description of the signal — it is the output of a process, and the process makes mistakes, has boundaries, and can be caught in the act.
The easiest way to catch it is a sequence of alternating tones played at increasing speed.
Play it slowly and it is one line with a limp. Speed it up and, at some point, it stops being one line. Two lines appear: a steady high pulse and a steady low one, each perfectly regular, running at the same time. The gallop is gone, and it is not merely hard to hear — it is unavailable, and reconstructing it requires effort or fails altogether.
Two boundaries, not one
The obvious hypothesis is that there is a speed at which the switch happens, and that it depends on how far apart the two tones are. Half right. There are two speeds, and the gap between them is the finding.
Leon van Noorden measured them in 1975 by asking listeners to do something unusual. Rather than asking which percept arrived, he asked them to try to hold one — to hear it as one stream if they could, and separately to hear it as two if they could — and to report when they could no longer manage it.
The two boundaries behave differently, and the difference is the argument.
The fission boundary is nearly flat. Below about five semitones, the sequence is one stream at any speed. Separation, not rate, decides it, and the value barely moves across the whole range of tempi.
The coherence boundary moves with rate, and steeply. At a slow rate, tones fifteen semitones apart can still be held together as one line. At a fast rate, four semitones is enough to break them apart whether the listener likes it or not.
If streaming were a determinate function of the stimulus, these two would be one line. They are not, and the region between them is several semitones wide over the entire musical range. In that region the percept is genuinely bistable: the same input, the same listener, two available organisations, and a switch that is partly voluntary.
What the composers did with it
The bistable region is not a laboratory curiosity. It is a compositional device with a name, a repertoire and a three-hundred-year history.
Compound melody. A single melodic line, written for a single instrument, that alternates between two registers fast enough that the listener hears two voices. It is the entire basis of the unaccompanied string and keyboard writing of the Baroque: the solo violin partitas, the cello suites, the keyboard preludes. A single violin plays one note at a time and the listener hears a bass line and a melody, because the notes are far enough apart and fast enough to cross the coherence boundary.
This is a real compositional constraint with real numbers behind it. The device fails if the registers are too close — the fission boundary at five semitones is a lower bound on how far apart the two implied voices must sit — and it fails if the tempo is too slow, because at slow rates even a wide separation stays coherent. Whether written-out compound melody in the repertoire sits above both boundaries is therefore a testable claim about a body of music, and the test is run below. It half survives, and the half that fails is the more interesting one.
The hocket, and interlocking parts. Two players alternating single notes so that the result is heard as one line neither of them is playing. This is the opposite use of the same mechanism: keep the two parts close in register so that they fall below the fission boundary and fuse into one stream, and the individual players disappear. Balinese kotekan and medieval European hocket are both built on it, and both keep the interlocking parts within a few semitones for exactly the reason the flat boundary predicts.
The illusion that comes free. A melody whose notes alternate between two registers can be made to be heard two ways, and the switch is available to the listener at will. Anything sitting in the middle band is a passage a listener can hear more than one way — and unlike an ambiguous metre, where the reading is chosen early and is expensive to revise, a stream assignment can be flipped back and forth during listening.
A real passage, measured
The claim that a repertoire sits above these boundaries can be checked rather than asserted.
Take the standard case: the arpeggio figuration of an unaccompanied Baroque prelude, in which a held bass note alternates with a moving upper line. In the first prelude of Bach’s first cello suite the pattern is semiquavers at roughly seventy crotchets a minute, which is about four and a half notes a second, or 215 milliseconds between onsets. The bass note and the upper line are typically an octave to a twelfth apart — twelve to nineteen semitones.
Read those two numbers onto the boundary plane, and the first thing to say is that 215 ms is off the end of it. Van Noorden’s slowest condition is 160 milliseconds, where the coherence boundary stands at 14.5 semitones and is still climbing; a prelude at four and a half notes a second is slower than anything he measured, and the boundary there is unknown and at least that high.
Take it at its last measured value and the passage straddles it. A twelfth — nineteen semitones — is above the boundary, so those bars are obligatorily two streams and a listener who wanted one leaping line could not have it. An octave is below it, in the bistable band, where the two-voice reading is available and so is the one-voice reading. So the device does not work by putting the passage beyond choice. It works by writing at a separation where the two-voice reading is the easier one, repeatedly, for two minutes, so that the listener who could in principle hold it together as one line has no reason to and every bar’s worth of encouragement not to.
That is a weaker claim than the one usually made for compound melody, and it is the one the numbers support.
It also has a tempo in it, which is worth extracting because it is the one variable a performer holds. The coherence boundary falls as the rate rises, so every separation crosses it at some speed, and the plane can be read backwards to say which:
| leap | crosses the boundary at | which is a prelude at |
|---|---|---|
| a twelfth | 160 ms between notes | 94 crotchets a minute |
| an octave | 146 ms | 103 |
| a fifth | 104 ms | 144 |
Play the prelude at 103 crotchets a minute and its octaves stop being a choice. That is not a fast tempo — it is an ordinary one, inside the range of recorded performances, and the essay’s seventy is at the slow end of them. So the passage the paragraph above describes as bistable is bistable at that tempo, and the same notes played a third faster are obligatorily two voices with nothing left for a listener to decide.
Which reverses what the device asks of a performer. It is not that a slow tempo lets the two voices be heard; a slow tempo is what leaves the two-voice reading optional, and speed is what makes it compulsory. The one thing the numbers cannot say is whether the boundary keeps climbing past 160 milliseconds, so a tempo slower than ninety-four is off the measured plane in the direction where nobody has looked.
What makes even the weaker result worth having is that nothing was fitted to anything: the boundaries were measured in the 1970s on pure tones by people who were not thinking about Bach, and the repertoire was written in the 1720s by somebody who was not thinking about psychoacoustics. What the figure cannot show is the boundary at the rate the music actually moves at, because nobody has measured it there — every number on this plane beyond 160 milliseconds is the last measured one held flat, and the honest reading of the passage is that it sits at or past the edge of the evidence.
Whose music, and how far it goes. The claim above is about European unaccompanied instrumental writing between roughly 1700 and 1750, where the device is deliberate and consistent enough to be checked. The mechanism is not European and neither are its uses — West African bell patterns, Balinese interlocking figuration and the two-register alternation of much banjo and mbira playing all exploit the same boundaries — but the numbers have only been checked against one repertoire, and a claim about all music from one corpus is the failure mode this site is most exposed to. What generalises is the mechanism. What has been measured is one shelf of one library.
What decides the assignment
Separation and rate are the two dimensions van Noorden measured, and they are not the only cues. Albert Bregman’s account, which is the standard one, lists several and the list is a list of regularities of the physical world rather than of properties of sound.
Proximity in frequency. Sources tend to stay put in frequency; a sudden leap is more likely to be a different source than a large excursion by one source. This is the cue the figures above are about.
Proximity in time. Events close together are more likely to belong together than events far apart, which is why the effect depends on rate at all.
Common onset. Components that start together probably came from one event. This is powerful enough to have its own rung on this ladder and it can override frequency proximity entirely.
Common fate. Components that change together — in level, in frequency, in position — belong together. A vibrato applied to some partials of a complex and not to others splits the complex in two.
Continuity. A sound interrupted by a louder one is heard as continuing through the interruption, provided the interrupting sound would have masked it. This is the continuity illusion, and it is the same masked threshold doing a constructive job: if the ear can prove that the missing part would have been inaudible anyway, it fills it in.
The continuity cue deserves a second look, because it is the one that most clearly shows the system asserting something rather than reporting it. Interrupt a tone with a burst of noise and the tone is heard as continuing behind the noise, unbroken. Remove the tone entirely during the noise and it is still heard as continuing — the ear inserts it. The condition is that the noise must be loud enough to have masked it: if the noise is too quiet, the gap is heard as a gap and the illusion collapses.
That is a striking piece of reasoning to find in a sensory system. It amounts to: this evidence is consistent with the tone having continued, and it is also consistent with the tone having stopped exactly when it became inaudible and resumed exactly when it became audible again — and the first is far more likely, so that is what happened. The alternative would require a coincidence. The ear does not believe in coincidences, and the masked threshold is the calculation it uses to decide whether one is being proposed.
Every one of these is a bet about how the world usually works, and every one can be made to fail by an experiment designed to make it fail. That is the strongest evidence that they are heuristics rather than measurements.
The middle band, then, is where the interesting music is not. Composers who want a result reach past the boundary rather than up to it. What the middle band is good for is a different thing entirely — passages that reward being listened to more than once, and reward being listened to differently.
The Baroque figure is the case where the assignment has a compositional purpose, and it is worth seeing what happens to it at a tempo nobody would play it at.
The device stops being a device. Written to be heard as one line implying two, played fast enough it is simply two lines, and the composer’s control over the ambiguity is gone — which is the practical form of everything above: the boundary is not a fact about the notes, so a tempo mark can move a passage across it without a single pitch changing.
Which computation produced the numbers
Almost nothing in this essay is computed. The two boundaries are van Noorden’s measurements, interpolated between his published points, and the figure says so. Nothing about them falls out of a model of the cochlea.
That is a different epistemic position from most of this site, and it is worth being explicit about the difference. The comma is arithmetic and can be derived here from nothing at all. The critical band is a fit to masking data but a tight one, reproduced by several laboratories with agreement to a factor of well under two. The streaming boundaries are looser than either: the fission boundary is reasonably stable across studies, and the coherence boundary depends on how the question is asked, on how long the listener has been listening — it drifts for tens of seconds after a sequence starts, always towards splitting — and on the listener’s musical training.
What is robust is the qualitative structure, and it is what the essay’s argument rests on: there are two boundaries, they are not the same, and the region between them is wide.
What the picture cannot show
Build-up. A sequence heard for ten seconds is far more likely to split than the same sequence heard for one. The coherence boundary is a moving target, and the figure draws it as a line. Any experiment that presents a sequence briefly measures a different thing from one that presents it for half a minute, and the literature contains both.
Attention. The bistable region is under voluntary control, which means the measurement depends on what the listener was asked to do, which means “the boundary” is partly an instruction. Van Noorden’s design confronts that honestly by measuring both instructions separately; a design that simply asks “how many streams?” is measuring the listener’s default, and the two numbers are not the same number.
Real music is not tones. Every figure here is pure tones of one duration alternating regularly. Real instruments have spectra, and spectral difference is itself a strong streaming cue — two instruments in the same register segregate readily if their lists of partials differ, which is why a flute and a clarinet in unison can still be heard as two things. None of that is on these axes, and it may be the strongest cue of all in practice. The reverse case is the one counterpoint has a rule about: two voices moving in parallel fifths have partials that mostly coincide, and they stop being heard as two — a fusion failure that the streaming literature and the eighteenth-century treatises describe in different vocabularies and agree about.
The generalisation
Streaming looks like a special-purpose mechanism for tones and is not. It is one instance of a general problem, which is that the number of sources in a scene is not given and must be inferred, and every sense that receives a mixture has to solve it.
The auditory version is harder than the visual one in a specific way. Two objects in front of the eyes occupy different places on the retina; two objects making sound occupy the same eardrum, and the mixture is complete before any analysis begins. There is no auditory equivalent of occlusion, no place where one source simply covers another and the rest is visible. Everything is superimposed, and the only route to separation is through regularities like the ones above.
That is also the reason the failures are so specific. A machine that separates a mixture by betting on statistical regularities will fail exactly when the regularities do not hold, and the compositional devices in this essay are a catalogue of composers finding those cases and using them.
Where the ladder goes next
This essay has treated tones as the atoms — the things that get assigned to streams. They are not atoms. A note is a stack of partials, and the ear has already had to decide that they belong together before there is a note to assign anywhere. That decision uses different cues and it is the next rung: a single partial mistuned by three per cent leaves the note it belongs to, and a perfectly harmonic partial given a thirty-millisecond head start leaves it too.
Sideways from here, the same word — grouping — is doing the same work in time. Metre is an assignment too, it is inferred rather than received, and it is chosen early and defended. The two are the same kind of process operating on different axes, and neither is in the signal.
Part 1 of 8
One essay in the series on auditory scene. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 21.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Auditory scene analysisBistabilityCompound melodyFission boundaryGroupingStreamingTemporal coherence
- A melody is a walk, not a set auditory scene analysis, fission boundary, streaming, temporal coherence
- The note that has a length streaming, temporal coherence