Intervals and chords

The ear sorts into boxes, and the boxes are the theory

Slide one note slowly upward against another and the interval between them changes continuously. What a listener reports does not. It stays a minor third, stays a minor third, and then in the space of about twenty cents becomes a major third — and nothing in the sound corresponds to the moment of the change.

Assumes: Two notes and a ratio, which is the whole of consonance

There is a demonstration that takes about fifteen seconds and rearranges what a listener thinks hearing is. Play two tones, hold the lower one still, and raise the upper one smoothly from three semitones above it to four. Ask what happened.

The physical event is a smooth glide of one hundred cents. The reported event is a minor third that lasted, went briefly uncertain, and turned into a major third.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 1 Identification against interval size: the probability that a trained listener names each category as the interval is swept upward, modelled as a logistic with the sharpness such experiments report. The shaded band is where one category’s share falls from three-quarters to one-quarter. It is a quarter of the width of the category itself, so for most of the gap between two named intervals there is nothing to be uncertain about.

This is categorical perception, and it was first characterised for speech sounds in the 1950s before anybody looked for it in music. The finding for music is due principally to Burns and Ward, whose 1978 experiments in the Journal of the Acoustical Society of America had trained musicians label melodic intervals across a continuum, and to Siegel and Siegel, who a year earlier had established the more uncomfortable half of it.

What is continuous and what is not

The acoustics contain no boundary. Sweeping the upper tone produces a smoothly changing ratio, and every sensory quantity that depends on it changes smoothly too.

Roughness across an octave. Sensory dissonance between two complex tones as the upper one is swept through an octave, computed by summing the roughness between every pair of their partials. Nothing here is placed by hand. The wells this spectrum produces sit on 4/3, on 3/2, on 5/3, found by scanning the curve rather than by marking them. And the tempered minor third sits 198 cents below the nearest well, and the tempered major third sits 98 cents below the nearest well, which is a real feature of this model and not a defect of the drawing.
Fig. 2 Roughness across an octave for a string spectrum, with the minor third, major third and fifth marked. The curve is smooth: there is no discontinuity at 350 cents, no feature at any category boundary, and the thirds are not even minima. Whatever produces the boundary is not in this.

Nor is the boundary at a place the arithmetic marks. The minor third’s just ratio is 6:5, at 316 cents; the major third’s is 5:4, at 386. The equal-tempered versions sit at 300 and 400. The reported boundary falls near the midpoint of the tempered pair, which is to say near neither ratio and near nothing acoustically distinguished at all.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 70 that separate adjacent categories — so the change of mind happens in 34% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 3 The rival account drawn in the same terms: categories centred on the ratios — 6:5 at 316 cents and 5:4 at 386 — rather than on the tempered pair. Its crossing lands at 351 cents, and the tempered version’s lands at 350. The two accounts predict the same boundary to within a cent, so the thirds cannot separate them, and the demonstration everybody uses is the one case where the question is undecidable. What does change is the ratio of doubt to certainty: the just pair is 70 cents apart against the tempered pair’s 100, so the same 24-cent crossing occupies 34 per cent of the gap instead of 24.

That the boundary follows the tuning system rather than the ratios is the first evidence that the categories are learned. It is not the strongest.

It is worth noticing how far the boundary is from either ratio, because the distance is larger than the quantities this collection usually argues about. The reported boundary is near 350 cents; the just minor third is 34 cents below it and the just major third 36 above. So a listener’s dividing line sits almost exactly midway between the two ratios — which is where a tempered system puts it and is not where a system built on ratios would. If the boundary tracked the ratios instead it would fall midway between them, at 351 cents, which is the same number to within a cent. So the thirds cannot separate the two accounts at all: a boundary set by the tempered pair and a boundary set by the just pair predict the same place, and the experiment that reports 350 is consistent with either.

Separating them needs a pair whose just and tempered versions are not symmetric about their midpoint, and the differences elsewhere are small — four cents at the fifth against the tritone. So the claim that the boundary follows the tuning system rather than the ratios is not established by the thirds, which is the case the demonstration uses, and the evidence for it has to be the training and repertoire results below rather than the position of any one boundary.

What happens inside a category

Siegel and Siegel’s 1977 paper is titled, with some relish, “Categorical perception of tonal intervals: musicians can’t tell sharp from flat”. The finding is that trained musicians presented with intervals mistuned by substantial amounts within a category did not report them as mistuned. They reported the category.

The amounts involved are not small. A major third twenty cents sharp is a fifth of a semitone out — well above any threshold for detecting a pitch change when the comparison is available — and it was labelled a major third with no note taken of its size.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 4 Where the two centuries of thirds actually sat, on the plateau they sat on. Just intonation’s major third is 386 cents and Pythagorean’s is 408 — a spread of twenty-two, and both are inside the flat region where identification is at or near certainty and the crossing to the fourth is still fifty cents away. A fifth of a semitone of variation, and the identity does not move at all. That is what a category system is for rather than a defect of one: it discards the variation that does not change the name, and interval size within a category is exactly such variation.

This is not a deficiency in trained listeners. It is what a category system is for: it discards the variation that does not change the identity, and interval size within a category is exactly such variation. Two centuries of European music were played with major thirds anywhere between 386 and 408 cents depending on where the temperament put the comma, and the whole enterprise depends on those all counting as major thirds.

But mistuning is not inaudible

Here is the part that makes the phenomenon coherent rather than merely strange. A mistuned interval is perfectly audible. It is simply not audible as a different interval.

329.6 Hz against 327 Hz. Two tones 2.6 hertz apart, added. The rapid oscillation is their average; the slow swelling is their difference, heard as 2.6 beats a second and used by every tuner who has ever worked by ear.
Fig. 5 Two tones a little over two and a half hertz apart, which is roughly what the coincident partials of a tempered major third do at this register. What a listener hears when a third is out is this — a pulsing that gets faster as the interval gets further out — and not a change of name.

When two notes are held together, mistuning shows up as beating between their coincident partials, which is precisely the quantity a tuner counts. That is a continuous, gradual, unmistakable cue, and it lives on a completely different perceptual axis from the interval’s name.

So the system has two channels. Which interval is this is categorical, sharp-edged and learned. How well is it tuned is continuous, sensory and immediate. A musician can tell that a chord is out to within a couple of cents while being unable to say whether the third in it was 390 or 405.

Melodic intervals — one note after another — lose the beating cue entirely, because the two notes are never present at the same time. And melodic mistuning is exactly where Siegel and Siegel found the effect strongest.

Who has the categories

The categories are not universal, and the evidence for that is two kinds.

The first is training. Burns and Ward found the effect strongly in trained musicians and weakly or not at all in listeners without musical training, who behaved much more like continuous judges of a continuous quantity. Whatever the boundaries are, they are acquired.

The second is repertoire. A category system inherited from twelve-tone equal temperament has a box for a 300-cent interval and a box for a 400-cent one and nothing between them, and there are living traditions in which the interval between them is a category of its own.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 50 that separate adjacent categories — so the change of mind happens in 48% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 6 The same model with a third box inserted where a maqam tradition puts one. The neutral third at 350 cents is exactly where a listener trained on twelve equal steps reports a boundary, and here it is a category with a name and a function of its own. The arithmetic of that is unforgiving: the centres are now 50 cents apart rather than 100, so the crossing occupies 53 per cent of the gap instead of 24 — over half of the interval’s range is a region of genuine doubt. A finer set of boxes is not a free improvement; it is a system in which more of the continuum is ambiguous, and one that has to be learned more precisely to be usable at all.

For a musician raised in that tradition the neutral third is not an out-of-tune major third or a sharp minor third: it is an interval, with a name, a function and a range of acceptable sizes of its own. A scale is not a set of pitches, and this is the perceptual half of the same point — what a listener has is a set of boxes, and which boxes depend on what they grew up sorting into.

The size of the tolerance, and what it buys

It is worth putting a number on how much variation a category absorbs, because the number is what makes the whole system useful rather than merely interesting.

The just major third is 386 cents. The equal-tempered one is 400. The Pythagorean one — four pure fifths reduced by two octaves — is 408. In quarter-comma meantone the major third is pure at 386 by construction, and in an irregular scheme like Werckmeister III it varies from key to key across a range of about fourteen cents. Every one of those is a major third to a listener, and the total spread is over twenty cents.

Twenty cents is about a fifth of a semitone. It is also the order of the deviations every practical tuning system has to accept, because twelve fifths do not close against seven octaves and the 23.46-cent discrepancy has to be distributed somewhere. The tolerance of the categories and the size of the comma are close to each other, and the relation between them is worth being exact about, because the direction of the near-miss is what makes the whole subject exist.

The comma is slightly larger than the tolerance, not slightly smaller. 23.46 against about 20 is a miss of three and a half cents, seventeen per cent — and that is why the comma has to be distributed rather than ignored. Put the whole of it in one interval and it is outside the box; spread it over four and each share is inside. If the comma had come out at eighteen cents instead, any single interval could have carried the lot with no audible consequence and there would be no literature.

The near-agreement is also a fact about twelve and not about tuning in general. A category’s tolerance scales with the division’s step, and the comma does not:

division step the comma as a fraction of a step
7 171.4 0.14
12 100.0 0.23
19 63.2 0.37
31 38.7 0.61
53 22.6 1.04

At fifty-three the comma is one whole step — a category of its own, which is exactly why the fifty-three-tone systems have a name for it and can write it down. At seven it is a seventh of a step and could be spread anywhere without anybody noticing. Twelve is near the crossing, and solving for where the comma exactly fills a fifth of a step gives a division of ten. So twelve is the smallest common division at which the comma has to be shared out, and the largest at which sharing it out is enough.

That is a genuine coincidence worth stating carefully, because it is easy to overclaim. It is not that the ear evolved a tolerance to accommodate the comma; the comma is arithmetic and the ear is older. What can be said is that the category system is roughly the right size to make tempered music possible, and that had the tolerance been three times narrower no keyboard would have worked at all.

What sharpens a boundary

The width of the transition is the quantity most worth measuring, because it is what distinguishes a category system from a mere labelling habit.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 81 cents, against the 100 that separate adjacent categories — so the change of mind happens in 81% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 7 The same model with the transitions three times wider — which is roughly the behaviour of an untrained listener. The categories are still there in the sense that the labels are still used, but a listener like this is reporting a position on a continuum rather than an identity, and there is a large range over which two answers are equally likely.

Comparing the two figures makes the claim concrete. With sharp boundaries the great majority of the range is unambiguous and only a narrow band is contested; with blunt ones there is no range that is not contested. The first is a system of categories; the second is a graded judgement wearing category names.

The evidence that trained listeners are genuinely doing the first rather than a very good version of the second comes from discrimination rather than identification. In the classic categorical-perception result, pairs of stimuli that straddle a boundary are easier to tell apart than pairs the same distance apart within a category — a discrimination peak sitting exactly where the identification function is steepest. For musical intervals this peak is real but weaker than in speech, which is the reason Burns and Ward’s title asks whether the phenomenon is a phenomenon or an epiphenomenon, and the reason the honest summary is that musical categorical perception is a strong effect in identification and a modest one in discrimination.

The parallel with speech, and where it breaks

Categorical perception was found in speech first, and the parallel is close enough to be illuminating and different enough to be worth spelling out.

The speech case is the voiced–voiceless distinction: a continuum of stimuli between ba and pa, varying the delay between the release of the lips and the onset of the voice, is reported as ba up to a boundary and pa after it, with a very narrow transition and a large discrimination peak sitting on the boundary. Listeners are markedly better at telling apart two stimuli that straddle it than two the same distance apart on the same side.

Three things are the same in the musical case. The identification function is steep. The boundary is learned rather than acoustic — speakers of different languages put it in different places, exactly as musicians from different traditions do. And within a category, variation that would be easily detectable in isolation goes unreported.

One thing is different, and it is the reason the musical result took two decades longer to establish. The discrimination peak is much weaker for intervals. A musician can hear the difference between a 390-cent third and a 405-cent third when the two are presented in immediate succession, and will report both as major thirds if asked to name them. In speech the corresponding discrimination is genuinely impaired. So music’s categories sit somewhere between speech’s — where the category almost replaces the sensory information — and a pure labelling scheme, where the sensory information is intact and merely summarised.

That intermediate position is what the two-channel account above describes, and it is why a musician is simultaneously the sort of listener who cannot tell sharp from flat and the sort who complains loudly about intonation.

Where the model in the figure stops

The curves above are a model, and it is worth being exact about what that means. The logistic shape is the standard fit for identification data and the parameter controlling its sharpness is taken from the literature; the boundaries are placed at the category midpoints, which is where trained listeners put them. What is drawn is therefore the shape of the finding, not a set of measurements.

Drawing a plausible-looking scatter of individual listeners instead would have been more persuasive and would have been an invention. The figure states its parameters, and the value of any of them can be argued with.

Three real complications the model leaves out. Boundaries move with context: the same interval is identified differently depending on the key it is heard in and what preceded it. They move with register and timbre, though less. And they are not perfectly stable within one listener across a session, which is why identification data are pooled over many trials and why an individual crossing point is not a well-defined number.

What this does to the theory

If the categories are learned and the tolerance is wide, then a good deal of what music theory calls a fact about intervals is a fact about a category system, and the two are worth keeping apart.

The vocabulary survives the substitution intact. A major third is still a major third whatever its exact size, chord names still work, and every argument about which chords are smooth or how a progression moves is stated in categories and unaffected. That is what makes the theory usable: it is written at a level of description the ear actually delivers.

What does not survive is any claim that a particular size is the interval. The just major third at 386 cents has a strong claim to being the acoustically privileged version — it is where the partials coincide, and where the roughness model puts a well for a spectrum with a strong fifth partial — and it is not what most listeners have heard most of the time. Twelve-tone equal temperament has been the default for two centuries and its third is fourteen cents away from that, which is a real distance and inside the category.

The consequence is that “the major third” names three different things depending on the sentence: a ratio, a number of semitones, and a perceptual category. They coincide closely enough that the ambiguity almost never causes trouble, and the places where it does are exactly the places this site keeps returning to — tuning, where the ratio and the semitone count come apart, and traditions in which the categories themselves are drawn differently.

What the picture cannot show

An identification curve is silent about the thing that makes the phenomenon matter, which is that both channels operate at once. A listener hearing a slightly wrong major third gets major third and and it is out simultaneously, from the same signal, on axes that do not talk to each other — and no single plot with interval size on the horizontal shows two answers at the same point.

It also cannot show the asymmetry between melody and harmony. Held together, a mistuned interval beats and the mistuning is obvious; played in sequence, the same mistuning is largely lost, because nothing is available to beat against. The categories are the same in both cases and the tolerance around them is not.

And it says nothing about how the boxes were acquired. Every result here is from adults who already have them, and the developmental question — at what age, from what exposure, and whether the boundaries can still move — is a different literature with a much thinner evidence base.

The ladder from here goes towards what the boxes are made of. If a listener has a category for a 350-cent interval only where a tradition supplies one, then the interesting question is which selections of intervals a tradition can supply — which is what makes seven notes out of twelve a workable set rather than an arbitrary one.

Part 1 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 23.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionCentsIntervalJust-noticeable differenceModal practice