Perception and the listener

The chord is still major, and that is why temperament works

A major third can be seventeen cents wrong and still be a major third. That tolerance is not a failure of hearing — it is the reason the whole subject of tuning is a discussion rather than a catastrophe. Every temperament ever proposed moves intervals around inside their categories, and the one thing none of them may do is push one across a boundary.

Assumes: The ear sorts into boxes, and the boxes are the theory

Equal temperament’s major third is 400 cents. The just major third is 386. Pythagorean tuning’s is 408. Quarter-comma meantone’s is 386 exactly, and in Werckmeister III the thirds range from 390 to 408 depending on the key.

That is a spread of twenty-two cents across systems in ordinary use, and it is four times what a listener can discriminate. Every one of them is a major third.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 1 Identification of an interval as a listener sweeps it continuously through the region containing the minor third, the major third and the fourth. Between the transitions there are flat regions where the answer does not change at all, and within one of those regions a difference of twenty cents changes what an interval sounds like without changing what it is. Everything tuning theory argues about happens inside one of these plateaux.

The tolerance is the enabling condition

State the argument in one line: temperament is possible because interval categories are about a hundred cents wide and the errors temperament introduces are about twenty.

That is not a small observation. Imagine the alternative. If the categories were as narrow as the discrimination limit — five cents — then equal temperament’s major third would be fourteen cents outside the just third’s category, and a tempered chord would not be a major chord at all. It would be a different, nameless interval, and the entire apparatus of naming, notation and harmonic function would collapse the moment a keyboard was tuned.

Instead the category absorbs it. A listener hearing a tempered major third hears a major third that is slightly bright, or slightly hard, or in the historical vocabulary slightly sharp — a modified member of a category, not a different object. That the modification is audible is what makes temperament a subject with a literature. That it does not change the category is what makes temperament work at all.

Two quantities, and they must not be confused

Three numbers have now appeared and keeping them apart is the whole discipline of this essay.

The discrimination limen: about five cents. The smallest change a listener can detect between two tones. Established on the previous ladder.

The category width: about a hundred cents. The distance from one interval name to the next. The categories are contiguous — there is no gap between the major third’s territory and the fourth’s — so the width is set by the twelve-tone system rather than by the ear.

The tolerance in practice: perhaps thirty or forty cents. How far an interval can be moved before a listener starts to report it as out of tune rather than as a coloured version of itself. This is smaller than the category width and much larger than the limen, and it is the number that actually governs.

Confusing the first with the third produces the standard misconception in both directions: either that any tuning difference is inaudible (false — the limen is five cents), or that any tuning difference is a wrong note (also false — the tolerance is thirty). Both errors are common in writing about temperament, and both come from treating one number as though it answered every question about hearing a pitch difference. There are at least three questions and they have three different answers.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 2 The territory the argument actually lives in, drawn alone: the minor third’s category and the major third’s, and the crossing between them. Just, Pythagorean and meantone put their thirds at 386, 408 and 386 cents respectively, against equal temperament’s 400 — a spread of twenty-two cents, all of it in the flat middle of the major third’s plateau where identification is at or near certainty. The change of mind happens in 24 cents out of the 100 that separate the two names, and every historical system keeps every interval out of that band. So the differences between the systems are differences of shade inside one identity.

The exception, and why it is one

Every historical temperament stays inside the categories. One thing in the historical repertoire does not, and it is famous precisely because it does not.

The wolf fifth is what happens when the whole accumulated error of a chain of fifths is dumped into a single interval. In quarter-comma meantone the wolf is about 738 cents, against a pure fifth’s 702 — thirty-six cents wide, and closing on the boundary between the fifth and the minor sixth.

That is why it is unusable while a twenty-cent error is merely a colour. The wolf is not a bad fifth; it is on its way to being a different interval, and a listener does not hear it as a rough fifth but as wrong — a category error rather than a shading. The name is apt: a wolf is not a large dog.

The same test explains why the wolf’s position was negotiable and its existence was not. A tuner could choose which key the wolf landed in, moving it to a key the repertoire did not use, and that choice is the whole practical content of a temperament’s design. What no tuner could do was distribute the error so thinly that it vanished — the comma is 23.5 cents and it has to go somewhere — so every system is an answer to the question how many intervals shall be shaded, so that none is broken.

Equal temperament’s answer, stated in these terms, is: shade all twelve fifths by two cents each, which is under the limen and therefore invisible, and pay for it with thirds shaded by fourteen, which is well inside the tolerance. Read as a categorical problem rather than as an arithmetic one, that is not a compromise. It is exactly optimal — and the section below, which measures it, also explains why that optimality is less impressive than it sounds.

It is worth stating the constraint the whole family is working under as one number, because it makes the arithmetic of every temperament feel tighter than it reads. An interval has at most fifty cents of room in either direction before it reaches a boundary, and the comma that has to be disposed of is 23.5. So the comma is nearly half of one interval’s entire budget, and a system that puts all of it in one place has spent three quarters of that interval’s room in a single step — which is what a wolf is, and it is why there was never a version of the trick that hid the comma in one fifth without breaking it. The choice was always between twelve small shadings and one break, because the budget only stretches to one of the two, and a fifth 36 cents wide is 12 cents from a boundary while a fifth 36 cents narrow is 16 from the one below it. Neither direction is available.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 3 The wolf’s own neighbourhood, four categories wide. A pure fifth sits at 702 and meantone’s eleven good fifths at 696.6 — both deep inside the fifth’s plateau, where nothing is in doubt. The wolf is at 738, which is on the rising edge toward the minor sixth: not a bad fifth but an interval on its way to being a different one. That is the whole difference between a shading and a break, and it is why the wolf’s position was negotiable and its existence was not — a tuner could move it to a key nobody played, and no tuner could spread 23.5 cents so thinly that it vanished.
Key character in Werckmeister III. The twelve major keys in circle-of-fifths order, each with a bar as long as its major third is sharp of a pure 5:4. The bars run from 3.9 to 21.5 cents, so the keys genuinely differ.
Fig. 4 The size of the major third in every key of Werckmeister III, drawn as a wheel. The spread is about eighteen cents, from the smoothest key to the roughest, and every one of the twelve is still a major third. What the figure draws is key character as a measurable quantity — and what this essay adds is why that quantity is a character rather than a defect: eighteen cents is a fifth of a category.

This is where categorical hearing stops being a fact about laboratories and becomes the mechanism behind an entire aesthetic. An irregular temperament works because its keys are audibly different and recognisably the same. If the spread pushed one key’s third across a boundary, that key would not have a character; it would be broken, and the eighteenth-century literature would describe it as such rather than as grave or brilliant.

Where the boundary actually is, and what happens there

The transitions in the identification figure are drawn about thirty cents wide, which is what identification studies of trained listeners report. Two things about that region are worth stating.

Identification becomes unreliable before it changes. In the transition a listener’s answers become inconsistent across trials rather than consistently reporting the other category. That is the signature of a boundary rather than of a shift.

Discrimination is better near the boundary. This is the classic categorical-perception result, and it holds for intervals as it does for speech sounds: two stimuli straddling a boundary are easier to tell apart than two stimuli the same distance apart in the middle of a category. The ear is not applying a uniform ruler. It is applying labels, and the resolution is highest where the label changes.

Which means the tuning question can be stated in an unusual and rather precise form. A temperament’s audible cost is not the size of its errors; it is how close its errors bring any interval to a boundary, because that is where the ear’s resolution is best and where the character of the interval changes fastest. An error of twenty cents in the middle of the major third’s territory is a shading. The same twenty cents applied to an interval already fifteen cents from a boundary is a much larger perceptual event.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 5 The boundary the wolf is approaching. Identification of an interval swept from the fifth up towards the minor sixth: the fifth’s plateau is broad and the transition is centred near 750 cents. A pure fifth at 702 is deep in the plateau; equal temperament’s at 700 is beside it; meantone’s narrowed fifths at 697 are still there. The wolf at 738 is in the transition, which is why it is not heard as a fifth with a fault but as an interval whose identity is in doubt.

The distance from the wolf to the boundary is 12.4 cents and the distance from a pure fifth to the boundary is 48.0. The wolf is four times nearer the edge, and that ratio is the quantitative content of every complaint ever made about it.

Running the proposed measure, which turns out to be one already known

That is a metric and it has never been applied to anything but the wolf. Scoring every temperament the site carries by how near its nearest interval comes to a boundary, over all twelve keys and all eleven interval classes:

nearest boundary worst error from just
equal 50.0 15.6
Werckmeister III 38.3 21.5
Young II 38.3 21.5
Vallotti 38.3 21.5
Kirnberger III 36.3 21.5
Pythagorean 28.5 23.5

The two columns rank them differently, which is what a new measure is for. Kirnberger III is second-best by error and fifth by boundary distance, and the interval responsible is its selling point: the pure major third at 386.3 cents, which is 36.3 from the boundary at 350 where every other system’s worst is 38.3 or better. Being just means being off the hundred-cent grid, and being off the grid is being nearer an edge.

And equal temperament scores exactly fifty, which is the maximum the measure allows — a number that should be suspicious rather than satisfying. It is: the distance from an interval to the nearest boundary and its distance to the nearest hundred-cent centre must add to fifty, so the measure is exactly fifty minus the largest departure from equal temperament anywhere in the tuning. It is not an independent perceptual criterion; it is a restatement of closeness to twelve-tone equal temperament, under which equal temperament wins by construction and just intonation loses.

That is worth stating plainly rather than quietly dropping, because the trap is one this essay half-warns about elsewhere: the categories are the twelve-tone system’s, so any measure of distance-to-boundary inherits that system’s grid. What the exercise does establish is the ordering among the irregular temperaments, which is genuinely new — the boundary measure and the just-error measure disagree about Kirnberger III, and the disagreement is a real fact about where its error sits rather than how large it is.

Which computation produced the numbers

The interval sizes are computed and the category is not.

Every cents value here comes from the site’s own tuning machinery: the just third from 5:4, equal temperament’s from 2^(4/12), meantone’s from the fourth power of the tempered fifth, Werckmeister’s from its published narrowings, and the wolf from what is left of the chain. None of them is quoted.

The identification curve is a logistic with a slope chosen to produce the thirty-cent transition width that studies report. That slope is the one number in the figure taken from the literature, and the figure’s own history is worth recording: an earlier version of it carried boundary positions a few cents off the category midpoints in order to support a claim about asymmetry, the figure computed the departure and reported it as two cents, and the claim was withdrawn because the numbers had been invented to support it. What the figure asserts now is transition width, which is what the studies actually measure.

What a performer does with the tolerance

The tolerance is not merely something to be survived. Musicians with flexible pitch spend it deliberately, and the direction they spend it in is consistent enough to have been measured.

Melodic leading notes are played sharp. String players and singers raise the seventh degree approaching the tonic, often by fifteen or twenty cents beyond equal temperament and thirty beyond just intonation. Nobody hears a different note; they hear a leading note that is leading harder. The whole effect lives inside the category, and it would be unavailable if the category were narrow.

Harmonic thirds are played flat. The same performers, sustaining a chord, drop the third towards its just position, because in that situation beating is the operative measurement and the just third is where the beating stops.

Those two pull in opposite directions on the same note, by twenty or thirty cents, and a performer switches between them according to whether the note is a melodic destination or a harmonic member. That is the practical content of “playing in tune”, it is why fixed-pitch instruments cannot do it, and it is why the question of which temperament an orchestra plays in is malformed.

Whose practice. Measured chiefly on Western string and vocal performance in the twentieth century, and it does not generalise cleanly. Traditions with fixed reference intervals and much finer categories — where a twenty-cent adjustment would cross a boundary rather than shade a note — do not have this freedom in the same form, and maqam practice, whose theory disputes a third’s position by thirty-five cents, is the clearest case of a system where twenty cents is a claim rather than an inflection.

What the picture cannot show

Context, which moves the boundaries. An interval heard in a key, after a cadence, in a familiar style, is categorised differently from one heard in isolation — and the shift can be several cents. Every figure here is an isolated interval.

Training, which sharpens them. Musicians show narrower transitions than non-musicians, and the difference is large. The category structure is at least partly learned, which is the connection this rung has to the rest of the perception ladder.

Simultaneity, which changes the task entirely. Everything above is melodic. A tempered third sounded against its root beats, and beating is detectable far below any of these thresholds. A listener asked about a sustained chord is not doing the categorisation task at all; they are doing a beat-detection task, and the numbers do not transfer.

Nor does it show a chord. Every identification study behind this figure presents two notes. A major triad is three, and whether its category behaves like the union of its intervals’ categories or like something of its own is a question nobody has answered cleanly. There is a reason to expect it does not: the tempered triad’s third is fourteen cents sharp while its fifth is two cents flat, and a listener judging the chord as a whole has two disagreeing pieces of evidence about how far from just it is.

And the categories are the twelve-tone system’s. A listener raised in a tradition with different intervals has different categories, which is not a hypothetical: the same maqam has a third that two national theories place thirty-five cents apart, and both are categories with names.

Whose music, and what changes when the system does

This is a claim about listeners trained in twelve-tone Western music, and it has an obvious test: what happens to a system with more categories?

Nineteen, thirty-one and fifty-three divisions of the octave make the categories narrower — 63, 39 and 23 cents respectively. At 53 divisions the category is 23 cents wide, which is smaller than the errors ordinary temperaments introduce, and the whole tolerance argument of this essay collapses. That is not a defect of 53-tone systems; it is a statement that they are doing something different. A listener in a 53-tone system is not being asked to tolerate a shading, because the system has a name for the shading.

The interesting consequence is that the number of categories and the tolerance for error trade off directly. A system with few, wide categories is robust to mistuning and coarse in what it can express. A system with many, narrow ones is expressive and unforgiving. Twelve is not a compromise anybody negotiated, but it does sit at a workable point on that trade: wide enough that a keyboard’s compromises fit inside, narrow enough that the compromises are audible as colour.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 25 cents, against the 63 that separate adjacent categories — so the change of mind happens in 39% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 6 The same crossing arithmetic in a nineteen-tone system, which is the test this section proposes. The categories are 63 cents apart rather than 100, the crossing itself is the same width — it is a property of the listener, not of the system — and so the uncertain band grows from 24 per cent of the gap to 39. Nineteen’s major third is 378.9 cents against just’s 386.3, an error of seven where twelve’s is fourteen; the thirds get better and the room to hide anything gets smaller in the same move. That is the trade-off, priced.

Where the ladder goes next

Categories are learned, and this rung has taken their existence for granted. What learns them, and from what, is the next question — and the answer turns out to be available in the same form as everything else on this ladder: the tonal hierarchy is a statistic, accumulated from exposure, and it separates into exactly the nested groups the theory already had names for.

Sideways, the tolerance argument meets its own limit case. If categories are what make consonance judgements robust, what happens to a listener who has different categories, or none? That is where consonance turns out to be half learned and where this site’s founding claim gets a boundary drawn round it from outside.

Part 2 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 20.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionCategory widthTemperamentToleranceWolf fifth