The ear sorts into boxes, and the boxes are the theory
Assumes: Two notes and a ratio, which is the whole of consonance
There is a demonstration that takes about fifteen seconds and rearranges what a listener thinks hearing is. Play two tones, hold the lower one still, and raise the upper one smoothly from three semitones above it to four. Ask what happened.
The physical event is a smooth glide of one hundred cents. The reported event is a minor third that lasted, went briefly uncertain, and turned into a major third.
This is categorical perception, and it was first characterised for speech sounds in the 1950s before anybody looked for it in music. The finding for music is due principally to Burns and Ward, whose 1978 experiments in the Journal of the Acoustical Society of America had trained musicians label melodic intervals across a continuum, and to Siegel and Siegel, who a year earlier had established the more uncomfortable half of it.
What is continuous and what is not
The acoustics contain no boundary. Sweeping the upper tone produces a smoothly changing ratio, and every sensory quantity that depends on it changes smoothly too.
Nor is the boundary at a place the arithmetic marks. The minor third’s just ratio is 6:5, at 316 cents; the major third’s is 5:4, at 386. The equal-tempered versions sit at 300 and 400. The reported boundary falls near the midpoint of the tempered pair, which is to say near neither ratio and near nothing acoustically distinguished at all.
That the boundary follows the tuning system rather than the ratios is the first evidence that the categories are learned. It is not the strongest.
It is worth noticing how far the boundary is from either ratio, because the distance is larger than the quantities this collection usually argues about. The reported boundary is near 350 cents; the just minor third is 34 cents below it and the just major third 36 above. So a listener’s dividing line sits almost exactly midway between the two ratios — which is where a tempered system puts it and is not where a system built on ratios would. If the boundary tracked the ratios instead it would fall midway between them, at 351 cents, which is the same number to within a cent. So the thirds cannot separate the two accounts at all: a boundary set by the tempered pair and a boundary set by the just pair predict the same place, and the experiment that reports 350 is consistent with either.
Separating them needs a pair whose just and tempered versions are not symmetric about their midpoint, and the differences elsewhere are small — four cents at the fifth against the tritone. So the claim that the boundary follows the tuning system rather than the ratios is not established by the thirds, which is the case the demonstration uses, and the evidence for it has to be the training and repertoire results below rather than the position of any one boundary.
What happens inside a category
Siegel and Siegel’s 1977 paper is titled, with some relish, “Categorical perception of tonal intervals: musicians can’t tell sharp from flat”. The finding is that trained musicians presented with intervals mistuned by substantial amounts within a category did not report them as mistuned. They reported the category.
The amounts involved are not small. A major third twenty cents sharp is a fifth of a semitone out — well above any threshold for detecting a pitch change when the comparison is available — and it was labelled a major third with no note taken of its size.
This is not a deficiency in trained listeners. It is what a category system is for: it discards the variation that does not change the identity, and interval size within a category is exactly such variation. Two centuries of European music were played with major thirds anywhere between 386 and 408 cents depending on where the temperament put the comma, and the whole enterprise depends on those all counting as major thirds.
But mistuning is not inaudible
Here is the part that makes the phenomenon coherent rather than merely strange. A mistuned interval is perfectly audible. It is simply not audible as a different interval.
When two notes are held together, mistuning shows up as beating between their coincident partials, which is precisely the quantity a tuner counts. That is a continuous, gradual, unmistakable cue, and it lives on a completely different perceptual axis from the interval’s name.
So the system has two channels. Which interval is this is categorical, sharp-edged and learned. How well is it tuned is continuous, sensory and immediate. A musician can tell that a chord is out to within a couple of cents while being unable to say whether the third in it was 390 or 405.
Melodic intervals — one note after another — lose the beating cue entirely, because the two notes are never present at the same time. And melodic mistuning is exactly where Siegel and Siegel found the effect strongest.
Who has the categories
The categories are not universal, and the evidence for that is two kinds.
The first is training. Burns and Ward found the effect strongly in trained musicians and weakly or not at all in listeners without musical training, who behaved much more like continuous judges of a continuous quantity. Whatever the boundaries are, they are acquired.
The second is repertoire. A category system inherited from twelve-tone equal temperament has a box for a 300-cent interval and a box for a 400-cent one and nothing between them, and there are living traditions in which the interval between them is a category of its own.
For a musician raised in that tradition the neutral third is not an out-of-tune major third or a sharp minor third: it is an interval, with a name, a function and a range of acceptable sizes of its own. A scale is not a set of pitches, and this is the perceptual half of the same point — what a listener has is a set of boxes, and which boxes depend on what they grew up sorting into.
The size of the tolerance, and what it buys
It is worth putting a number on how much variation a category absorbs, because the number is what makes the whole system useful rather than merely interesting.
The just major third is 386 cents. The equal-tempered one is 400. The Pythagorean one — four pure fifths reduced by two octaves — is 408. In quarter-comma meantone the major third is pure at 386 by construction, and in an irregular scheme like Werckmeister III it varies from key to key across a range of about fourteen cents. Every one of those is a major third to a listener, and the total spread is over twenty cents.
Twenty cents is about a fifth of a semitone. It is also the order of the deviations every practical tuning system has to accept, because twelve fifths do not close against seven octaves and the 23.46-cent discrepancy has to be distributed somewhere. The tolerance of the categories and the size of the comma are close to each other, and the relation between them is worth being exact about, because the direction of the near-miss is what makes the whole subject exist.
The comma is slightly larger than the tolerance, not slightly smaller. 23.46 against about 20 is a miss of three and a half cents, seventeen per cent — and that is why the comma has to be distributed rather than ignored. Put the whole of it in one interval and it is outside the box; spread it over four and each share is inside. If the comma had come out at eighteen cents instead, any single interval could have carried the lot with no audible consequence and there would be no literature.
The near-agreement is also a fact about twelve and not about tuning in general. A category’s tolerance scales with the division’s step, and the comma does not:
| division | step | the comma as a fraction of a step |
|---|---|---|
| 7 | 171.4 | 0.14 |
| 12 | 100.0 | 0.23 |
| 19 | 63.2 | 0.37 |
| 31 | 38.7 | 0.61 |
| 53 | 22.6 | 1.04 |
At fifty-three the comma is one whole step — a category of its own, which is exactly why the fifty-three-tone systems have a name for it and can write it down. At seven it is a seventh of a step and could be spread anywhere without anybody noticing. Twelve is near the crossing, and solving for where the comma exactly fills a fifth of a step gives a division of ten. So twelve is the smallest common division at which the comma has to be shared out, and the largest at which sharing it out is enough.
That is a genuine coincidence worth stating carefully, because it is easy to overclaim. It is not that the ear evolved a tolerance to accommodate the comma; the comma is arithmetic and the ear is older. What can be said is that the category system is roughly the right size to make tempered music possible, and that had the tolerance been three times narrower no keyboard would have worked at all.
What sharpens a boundary
The width of the transition is the quantity most worth measuring, because it is what distinguishes a category system from a mere labelling habit.
Comparing the two figures makes the claim concrete. With sharp boundaries the great majority of the range is unambiguous and only a narrow band is contested; with blunt ones there is no range that is not contested. The first is a system of categories; the second is a graded judgement wearing category names.
The evidence that trained listeners are genuinely doing the first rather than a very good version of the second comes from discrimination rather than identification. In the classic categorical-perception result, pairs of stimuli that straddle a boundary are easier to tell apart than pairs the same distance apart within a category — a discrimination peak sitting exactly where the identification function is steepest. For musical intervals this peak is real but weaker than in speech, which is the reason Burns and Ward’s title asks whether the phenomenon is a phenomenon or an epiphenomenon, and the reason the honest summary is that musical categorical perception is a strong effect in identification and a modest one in discrimination.
The parallel with speech, and where it breaks
Categorical perception was found in speech first, and the parallel is close enough to be illuminating and different enough to be worth spelling out.
The speech case is the voiced–voiceless distinction: a continuum of stimuli between ba and pa, varying the delay between the release of the lips and the onset of the voice, is reported as ba up to a boundary and pa after it, with a very narrow transition and a large discrimination peak sitting on the boundary. Listeners are markedly better at telling apart two stimuli that straddle it than two the same distance apart on the same side.
Three things are the same in the musical case. The identification function is steep. The boundary is learned rather than acoustic — speakers of different languages put it in different places, exactly as musicians from different traditions do. And within a category, variation that would be easily detectable in isolation goes unreported.
One thing is different, and it is the reason the musical result took two decades longer to establish. The discrimination peak is much weaker for intervals. A musician can hear the difference between a 390-cent third and a 405-cent third when the two are presented in immediate succession, and will report both as major thirds if asked to name them. In speech the corresponding discrimination is genuinely impaired. So music’s categories sit somewhere between speech’s — where the category almost replaces the sensory information — and a pure labelling scheme, where the sensory information is intact and merely summarised.
That intermediate position is what the two-channel account above describes, and it is why a musician is simultaneously the sort of listener who cannot tell sharp from flat and the sort who complains loudly about intonation.
Where the model in the figure stops
The curves above are a model, and it is worth being exact about what that means. The logistic shape is the standard fit for identification data and the parameter controlling its sharpness is taken from the literature; the boundaries are placed at the category midpoints, which is where trained listeners put them. What is drawn is therefore the shape of the finding, not a set of measurements.
Drawing a plausible-looking scatter of individual listeners instead would have been more persuasive and would have been an invention. The figure states its parameters, and the value of any of them can be argued with.
Three real complications the model leaves out. Boundaries move with context: the same interval is identified differently depending on the key it is heard in and what preceded it. They move with register and timbre, though less. And they are not perfectly stable within one listener across a session, which is why identification data are pooled over many trials and why an individual crossing point is not a well-defined number.
What this does to the theory
If the categories are learned and the tolerance is wide, then a good deal of what music theory calls a fact about intervals is a fact about a category system, and the two are worth keeping apart.
The vocabulary survives the substitution intact. A major third is still a major third whatever its exact size, chord names still work, and every argument about which chords are smooth or how a progression moves is stated in categories and unaffected. That is what makes the theory usable: it is written at a level of description the ear actually delivers.
What does not survive is any claim that a particular size is the interval. The just major third at 386 cents has a strong claim to being the acoustically privileged version — it is where the partials coincide, and where the roughness model puts a well for a spectrum with a strong fifth partial — and it is not what most listeners have heard most of the time. Twelve-tone equal temperament has been the default for two centuries and its third is fourteen cents away from that, which is a real distance and inside the category.
The consequence is that “the major third” names three different things depending on the sentence: a ratio, a number of semitones, and a perceptual category. They coincide closely enough that the ambiguity almost never causes trouble, and the places where it does are exactly the places this site keeps returning to — tuning, where the ratio and the semitone count come apart, and traditions in which the categories themselves are drawn differently.
What the picture cannot show
An identification curve is silent about the thing that makes the phenomenon matter, which is that both channels operate at once. A listener hearing a slightly wrong major third gets major third and and it is out simultaneously, from the same signal, on axes that do not talk to each other — and no single plot with interval size on the horizontal shows two answers at the same point.
It also cannot show the asymmetry between melody and harmony. Held together, a mistuned interval beats and the mistuning is obvious; played in sequence, the same mistuning is largely lost, because nothing is available to beat against. The categories are the same in both cases and the tolerance around them is not.
And it says nothing about how the boxes were acquired. Every result here is from adults who already have them, and the developmental question — at what age, from what exposure, and whether the boundaries can still move — is a different literature with a much thinner evidence base.
The ladder from here goes towards what the boxes are made of. If a listener has a category for a 350-cent interval only where a tradition supplies one, then the interesting question is which selections of intervals a tradition can supply — which is what makes seven notes out of twelve a workable set rather than an arbitrary one.
Part 1 of 11
One essay in the series on Categorical-hearing. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 23.
- How much an anchor would have to be worth
- The boundary that barely moves
- A note that is never at its pitch
- How many boxes an octave holds
- How small a difference is audible
- The number every claim here has been quoting
- The same distance, under two names
- Three answers to how finely a pitch can be heard
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Categorical perceptionCentsIntervalJust-noticeable differenceModal practice
- A boundary beside a fifth categorical perception, cents
- A comma under the threshold cents, just-noticeable difference
- A guitar cannot be in tune cents, interval
- An interval is two errors cents, just-noticeable difference
- The best seven of the twelve categorical perception, cents
- The notes in between interval, just-noticeable difference