A voice is a stream, and the ear decides which
Assumes: The shortest move, which is what a chord change is · The ear builds objects, and sometimes offers a choice
Everything on this ladder has taken a voice to be a thing that exists. The solver assigns four of them, the rules constrain how they may move, and the whole apparatus is built on the assumption that a listener is following four lines and would notice if two of them merged.
That assumption is testable and it has been tested, on pure tones, in a laboratory, by people who were not thinking about part-writing. What a listener follows is a stream, and whether a sequence of notes is one stream or two is decided by how far apart they are and how fast they arrive — not by what the score says they belong to.
Run the site’s own cheapest four-part realisation past those thresholds and it does not come off well.
Two boundaries, and only one of them is about speed
Leon van Noorden measured the thresholds in 1975 by asking listeners to try — to hold a repeating sequence as one line if they could, and separately to hold it as two — and to report when they could no longer manage it. The result is two boundaries rather than one, and the two behave completely differently. The essay that introduces them is about the wide bistable region between them; this one is about where the edges fall.
The fission boundary is flat. Below about five semitones of separation, a sequence is one stream at any rate whatever: 4.5 semitones at the fastest tempo measured and 5.2 at the slowest, which across the whole musical range is no change at all. Separation decides it; speed does not come into it.
The temporal coherence boundary is steep. Slow sequences hold together across wide leaps and fast ones do not: at 160 milliseconds between onsets a leap of fourteen semitones can still be heard as one line, and at 60 milliseconds four semitones is enough to break a line in two whether the listener likes it or not. The everyday version is a trill — wide and slow, one wavering line; the same two notes fast, two steady tones.
Both are properties of a repeating alternation of pure tones, which is what makes them numbers rather than impressions, and it is also the first thing that limits them.
The parts are too close together, at every speed
The fission boundary is the one that lands on four-part writing, and it lands hard.
Across every melody the progression above admits — 416 of them realisable under all seven rules — the gaps between adjacent voices in the cheapest realisations average 6.4 semitones between alto and soprano, 6.8 between tenor and alto, and 10.5 between bass and tenor. In 35 per cent of chords the alto and soprano are closer than the fission boundary, and in 18 per cent the tenor and alto are.
A gap below the fission boundary is not a difficult case. It is a case in which the two parts are unavailable as two, by pitch proximity, at every tempo a piece could be taken at. No amount of slowing down helps, because the boundary does not move with rate.
The obvious next thought is that the minimisation caused it — that squeezing the voices to move as little as possible squeezes them together. It did not. Over every complete voicing of those same chords, chosen without regard to cost, 39 per cent of adjacent gaps fall below the boundary, which is slightly worse than the cheapest realisations manage. The crowding belongs to four voices in a choral compass, not to the objective. Four parts inside about three octaves, with the bass conventionally set well below the rest, leaves the upper three sharing something like an octave and a half between them.
The fix is available and its price can be read off. Counting every complete four-part voicing of a triad inside the SATB ranges:
| major triad | minor triad | |
|---|---|---|
| voicings | 480 | 392 |
| adjacent gaps below the boundary | 39% | 43% |
| voicings with every gap clear | 63, or 13% | 33, or 8% |
| mean span, all voicings | 20.2 semitones | 17.9 |
| mean span, the clear ones | 31.5 | 28.8 |
| narrowest clear voicing | 20 semitones | 20 |
One voicing in eight is fully separable by pitch, and the ones that are spread over two and a half octaves rather than under two. The narrowest arrangement in which four voices are all at least five semitones apart spans twenty semitones — an octave and a major sixth — which is the floor a texture of four separable lines has to clear before anything else is considered.
What is striking is what the fix does not cost. The clear voicings are smoother, not rougher: 0.444 mean roughness against 0.477 across all of them for the major triad, and 0.444 against 0.495 for the minor. So spacing the parts far enough apart to be heard as parts also spaces them far enough apart to stop beating, and the two requirements point the same way.
The whole cost is in the span, and the span is what a choral texture is. Four voices spread over thirty-one semitones is not close harmony; it is an arrangement with a hole in the middle, and it is not what any part-writing exercise asks for. So the rules and the ear want the same thing and the compass will not give it to them, which is a sharper statement of this section’s finding than “the crowding belongs to the compass”: the compass is not merely where the crowding comes from, it is the one constraint of the three that cannot be relaxed without changing what is being written.
That is the first thing this rung was expected to find and did not. The proposal was that the smoothest voice leading stops being heard as voices above some computable tempo. There is no such tempo for the upper parts: the boundary that catches them is rate-independent, so the failure is a spacing failure and it is there from the first note.
Four voices on one chord at the spacing the cheapest realisation gives them are, on a keyboard, a hand — and the gaps between them are the quantity this essay is about rather than the notes.
What speed does decide
The coherence boundary is where rate matters, and it applies to a different failure: not two parts merging, but one part splitting.
A single line that leaps far enough, fast enough, stops being one line. This is not a defect — it is compound melody, the device that lets one unaccompanied instrument state two parts — and it is a hazard for anybody wanting the part to stay one thing.
Now compare that with what the parts actually do. In the cheapest realisation of the progression above, with the tune free to be whatever costs least, the largest leap anywhere in four voices is two semitones. Minimisation immunises the parts against this failure completely: a line that never moves more than a tone is one line at any rate in the measured range.
The immunity ends the moment the tune is given, and it ends unevenly. Over all 416 melodies, the three lower voices move by an average of 0.9, 1.1 and 1.3 semitones a chord and never exceed five; the soprano averages 4.3 and reaches twelve, with 45 per cent of its moves above the coherence boundary at a rate of ten notes a second. Minimising total motion with one voice fixed does not spread the work — it pushes every semitone it can into the voice it is not allowed to choose.
The same picture at a slower rate is the control the claim needs, because a threshold that catches everything is not a threshold.
The ear has its own solver, and it is the same one
The reason the two accounts keep nearly agreeing is that they are nearly the same computation.
Auditory grouping continues a line to the nearest available note in pitch. Voice-leading minimisation assigns each voice to the nearest available note in pitch, subject to using every note once. One is a perceptual heuristic and the other is an assignment problem, and on most chord changes they return the same answer — which is why the smooth writing of the tradition is followable at all, and why it did not have to be designed for followability.
They are not identical, and the difference is measurable. Taking every chord change in every one of those 416 realisations — 1,664 changes — and asking what a listener grouping by nearest pitch alone would make of it: in 35 per cent of them the nearest-note reading does not recover the written parts. Two notes of the new chord claim the same note of the old one, which leaves one written part without a continuation and splits another in two.
Where the two disagree, the ear does not defer. The score’s assignment is a fact about the score; the listener’s is what is heard, and no instruction in the part-writing is available to overrule it. That is the honest form of the claim this ladder has been building towards: the arithmetic describes what was written, and the thresholds describe what arrives.
The measurement this ladder opened with has no register in it at all: two chords are sets of pitch classes and the distance between them is a sum over an assignment, which is exactly the information a stream needs and does not have.
Fusion and fission are opposite failures
The third rung of this ladder was about two voices becoming one. This rung is about one voice becoming two, and about two becoming one again for a completely different reason. They are worth setting side by side, because the cures pull against each other.
Fusion is simultaneous and spectral. Two notes an octave apart share every partial, and two a fifth apart share half, so a pair moving in parallel at those intervals is heard as one sound with a brighter tone. The cure is contrary motion and rhythmic independence: make the two parts do different things at different moments.
Fission and coherence are successive and positional. A part is one line if its own notes are close together in pitch relative to the rate, and it is separable from a neighbouring part if the two are far enough apart in register. The cure is separation and small steps.
Now put the cures together. Contrary motion — the remedy for fusion — carries voices through each other’s registers, which is where the fission boundary is waiting. Register separation — the remedy for confusion between parts — sets voices at wide fixed intervals, which is where fusion is waiting: two parts an octave apart in register are exactly the pair that stops sounding like two. A texture cannot maximise both, and four-part writing in a choral compass is the compromise the tradition arrived at without either measurement.
The other kind of grouping is counted rather than timed: an octave shares all of the upper note’s partials with the lower, a fifth four of twelve, a third two — so two voices at a fixed interval fuse in proportion to what they share, which is the mechanism the parallel-motion prohibitions are about and a different one from streaming.
What actually keeps parts apart
Pitch proximity and rate are two cues among several, and in real music they are not the strongest.
Timbre. Two instruments with different spectra segregate readily in the same register — the ear hears the list of partials, and two different lists are two different sources. A flute and a clarinet a third apart stay two parts where two flutes would not, and none of that is on van Noorden’s axes, which are pure tones throughout.
Onset. Parts that begin at different instants stay apart: thirty milliseconds of head start is enough to pull a partial out of a note it belongs to, and rhythmic independence between voices supplies exactly that cue several times a bar. It is also the cue a round depends on and a roughness scan cannot see: the traditional two-bar entry exists so that each voice enters against material the others are not singing.
Vibrato and other common fate. Two singers with independent vibrato have partials wandering independently, which is a strong separation cue and one reason a choir is easier to follow than an organ.
So the crowding this essay measures is real and it is routinely overcome, by cues the measurement does not contain. What the measurement establishes is that pitch alone will not do it — that a texture relying on register separation to keep four parts distinct is relying on something that a third of its chords do not supply.
The same rate with two of the tones an octave apart splits into two streams immediately, which is the control the previous figure needs: nothing about the rate changed and everything about the grouping did.
Whose texture, and when
The realisations measured here are four voices in choral compasses, which is the eighteenth-century chorale and the exercise descended from it. That is a narrow target and it is chosen because it is the texture the rules of this ladder were written for.
It is also a texture designed around a specific solution to the problem. Choral part-writing separates the bass and packs the upper three, then keeps the three apart by rhythm and by text rather than by register: they enter together and move together, and what makes an alto line followable in a chorale is largely that a listener knows the tune is on top and stops trying to follow anything else. The measurement here says the pitch cue is unavailable a third of the time; the practice never depended on the pitch cue.
Other traditions solve it differently and their solutions are legible as solutions. A Baroque trio sonata separates two upper parts by giving them different instruments and staggered entries. A string quartet spreads four parts over four octaves, which puts every adjacent pair far above the fission boundary — at the cost of the octaves and twelfths between them, which is where the other failure lives. Orchestration by register and by colour is the practical form of everything in this essay, worked out several centuries before there was a threshold to name.
What the picture cannot show
These are measurements, not computations. Every number on the boundaries is van Noorden’s, interpolated between his published points. Nothing about them falls out of a model of the cochlea, and the coherence boundary in particular depends on how the question is put — a listener told to hold one stream reports a different threshold from one asked how many they hear, and the value drifts for tens of seconds after a sequence begins, always towards splitting.
Pure tones, one duration, no music. The sequences the thresholds were measured on alternate regularly and do not stop. A chorale has four simultaneous parts with a harmonic context, a text and a metre, and nothing in the experiment resembles it.
Nothing here is measured on a listener. Not one of the claims about the realisations has been tested by asking anybody what they heard. The gaps and the leaps are computed exactly; the thresholds are published; the join between them is an inference, and the inference is the part that could be wrong.
Attention is not in it. The bistable region means the assignment can be chosen, and a trained listener can hold parts apart that a naive one hears as a block. A threshold measured on attentive listeners in a quiet room is an upper bound on what a concert audience does.
And simultaneity is not sequence. The gaps counted here are vertical intervals between parts sounding together, while van Noorden’s separations are between successive tones. Applying a sequential threshold to a simultaneous interval is the weakest step in the argument, and it is defensible only because the parts in question are moving: each voice’s next note is the tone that follows, and the neighbouring voice’s next note is another, and the ear has to choose between them.
Where this ladder ends
Eight rungs, and the object has changed under the argument at every one of them. It began as an assignment between two sets of pitch classes, became a set of chords in a geometry, then a spectral question about what fuses, then a scan over delays, then four voices in a register with prohibitions priced in semitones, then a pair of criteria that turned out to be one, then a resolution that no distance could find — and it ends with a listener for whom none of it is given.
What is left undone is worth naming. Nothing here has measured a real texture: the realisations are the site’s own, generated to a compass, and a corpus of actual four-part writing would settle in an afternoon what fraction of real chords sit inside the fission boundary. Nothing here handles the bass as the structurally different voice it plainly is. And the join between a sequential threshold and a simultaneous interval deserves an experiment rather than an argument.
The last of those is the honest summary of the whole ladder. Voice leading is a good description of what was written and an approximation of what is heard, and the approximation gets worse exactly where the writing gets smoothest.
Part 8 of 9
One essay in the series on Voice-leading. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Auditory scene analysisFission boundaryFusionPart-writingStreamingTemporal coherenceVoice-leading
- The cue that settles it auditory scene analysis, fusion, streaming
- A blown note does not start late, it starts slowly auditory scene analysis, fusion
- Eleven partials is one too many auditory scene analysis, fusion
- Four parts are easier to read than two part-writing, voice-leading
- The note that has a length streaming, temporal coherence
- The partials that do not die together auditory scene analysis, fusion