Concept

Categorical perception — where it appears

Hearing a continuous physical difference as one of a few discrete categories, with fine discrimination at the boundaries and little within them. It is what makes a slow glide between two intervals report as a jump rather than as a slide.

Named by 15 essays across 4 fields — each of them below, with the objects they name alongside it.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.

The ear sorts into boxes, and the boxes are the theory

Slide one note slowly upward against another and the interval between them changes continuously. What a listener reports does not. It stays a minor third, stays a minor third, and then in the space of about twenty cents becomes a major third — and nothing in the sound corresponds to the moment of the change.

intervals · Categorical-hearing
Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.

The chord is still major, and that is why temperament works

A major third can be seventeen cents wrong and still be a major third. That tolerance is not a failure of hearing — it is the reason the whole subject of tuning is a discussion rather than a catastrophe. Every temperament ever proposed moves intervals around inside their categories, and the one thing none of them may do is push one across a boundary.

perception · Categorical-hearing
The swing ratio, against the categories it passes through. The same swing curve read against the boundaries between duration categories rather than against notated values. A category's centre is a simple ratio — 1:1, 2:1, 3:1 — and the boundary between two of them is the midpoint, which is arithmetic. The curve crosses 2 of them: out of 2:1 and into 1:1 at 240 beats a minute, out of 3:1 and into 2:1 at 171 beats a minute. So the same notated figure is, by the categorical criterion, a different rhythm at each end of an ordinary tempo range, and the notation says triplet feel throughout.

How late is a different note

A deviation of thirty milliseconds is expression and a deviation of two hundred is a wrong note, so there is an edge. The edges in time are arithmetic — the midpoints between the simple ratios — and the swing ratio crosses two of them as the tempo rises, at 171 and at 240 beats a minute, while the notation says triplet feel throughout.

perception · Microtiming
Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.

Three answers to how finely a pitch can be heard

Two notes one after the other are told apart at about four cents at A440. Whether a melodic interval is in tune is a judgement an order of magnitude coarser. And two notes held together are heard to beat at a third of a cent, because the question is answered by counting rather than by hearing pitch at all. Every equal division ever built sits between the coarsest and the finest.

perception · Beyond twelve
The chain of fifths in quarter-comma meantone. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and one of them — G♯ to D♯, the 12th link, where the chain is forced to close — is the wolf, at 35.7 cents.

The same distance, under two names

Four hundred cents is a major third or a diminished fourth, and on a keyboard nothing in the sound distinguishes them. An earlier essay was about the boundary between two categories; this is about two categories at one acoustic value, and the surprise is where the ambiguity comes from. In quarter-comma meantone a major third is 386 cents and a diminished fourth is 427 — two names, two pitches, forty-one cents apart. Equal temperament collapsed them, and what a listener now supplies from context used to be in the sound.

intervals · Categorical-hearing
How many notes an octave can hold, asked twice. The share of trials on which a category is named correctly, against how many equal categories the octave is cut into, for a listener whose internal estimate carries 11 cents of noise — the logistic scale this site's identification figures already use, which is the thirty-cent transition the studies report. At 95 per cent accuracy the ceiling is 6 categories, and seven scores 94.9 per cent — on the line. The other ceiling is resolution: 151 to 356 difference limens fit in an octave depending on register, which is a factor of forty larger. Every system marked below sits between the two, and the marks separate: the number of degrees a mode uses clears the criterion, and the size of the gamut it chooses them from does not.

How many boxes an octave holds

An identification model of pitch is usually handed twelve categories, and nothing ever asked how many an octave can hold. There are two answers and they are a factor of forty apart. Resolution allows between 151 and 356 — a listener can tell that many pitches apart in a direct comparison. Naming one of them without a comparison is a different faculty and it runs out at six or seven, which is where every mode in every tradition compared here sits. Turkish theory names fifty-three commas to the octave and a makam uses seven of them, and the gap between those two numbers is the whole of the argument.

intervals · Categorical-hearing
Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 2.8 cents at ten to one and 3.6 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it.

The boundary that barely moves

Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.

perception · Categorical-hearing
A duration category has a tempo range of its own. Each simple ratio's short note is the beat divided by one more than the ratio, so at a high enough tempo it falls under the fastest interval that can be a beat at all — 100 milliseconds. Each bar here runs from the slowest tempo at which the ratio's long note still belongs to a beat to the fastest at which its short note is still a note: 1:1 ends at 300 bpm, 2:1 ends at 200 bpm, 3:1 ends at 150 bpm, 4:1 ends at 120 bpm. The line is this site's swing curve, and where it crosses a category boundary the category it is leaving has already ceased to exist.

The short note is sitting on the floor

Swing is modelled here as a short note of constant absolute length, which was measured from drummers and left as a fitted parameter. The tempo window's fast edge — the shortest interval a series of events can be a beat at — was measured from listeners tapping. Both are a hundred milliseconds, and if that is not a coincidence then the swing ratio has no free parameter in it at all: the short note is not held constant, it is resting on the floor.

rhythm · Microtiming
How much correlation it would take to matter. The limen of a 7-semitone interval at a note length of 0.25 seconds, against the correlation between the two notes' errors. The independent model at the left gives 9.44 cents. Halving that needs a correlation of 0.75; a fifth off it needs 0.31. The curve is √(1 − ρ) and nothing else, so the correlation required for a stated improvement is arithmetic — which turns the question from “does a key help?” into “by how much, and here is the number it must reach”.

How much an anchor would have to be worth

Two pitch errors added in quadrature assume an independence nobody measured — a listener inside a key hears a note as a scale degree, and a shared reference is exactly a correlated error. Turning the dial is not evidence. What is evidence is that the dial is not free: a shared error cancels out of a difference completely, so a listener's single-note limen and their interval limen give the two components with nothing left over, and halving the interval limen needs a correlation of exactly 0.75.

intervals · Pitch-acuity
Every result so far, against the listener's own noise. Three findings drawn against the one parameter all of them assume: how finely the listener resolves a pitch. At 11 cents — a trained listener, and the value every earlier essay used — the octave holds 6 nameable categories, twelve equal ones are named right 91 per cent of the time, and a 20-to-one expectation moves a boundary by 3.6 cents. At 35 cents it is 2 categories, 72 per cent, and 37 cents. The capacity falls roughly as one over sigma and the shift rises as its square, so the three curves separate rather than moving together.

The listener the model was never run for

Five earlier essays rest on one number — how finely a listener resolves a pitch — and every one of them used a trained listener's eleven cents. The model's dependence on it is not gentle: the capacity goes as its reciprocal and the expectation shift as its square, so an untrained listener at thirty-five cents has two nameable categories per octave rather than six, and a foreign tuning system is not mis-transcribed by them but absorbed.

scales · Categorical-hearing
Every resolution claim here, against the number it rests on. How many equal steps of the octave can be named at 95 per cent accuracy, against the internal noise the model gives a listener. The laboratory value this collection quotes everywhere is 11 cents, which gives 6 nameable categories — the "about seven per octave" every claim here has been repeating. The laboratory measures it on isolated intervals and music never presents one, so the effective value inside a piece is smaller by an amount nobody has measured: at 4 cents it is 18, which is the chromatic scale, and the conclusion changes from "the ear has fewer boxes than the notation" to "it has exactly as many". Nothing here measures it. This is what it is worth if it moves.

The number every claim here has been quoting

Sigma is the internal noise a listener's pitch judgements carry, it is measured in a laboratory on isolated intervals, and music never presents an isolated interval. Every resolution claim here rests on the laboratory value. Sweep it and the headline finding moves: at eleven cents the ear has about six nameable categories per octave, and at six it has twelve — which is the chromatic scale, and turns 'the ear has fewer boxes than the notation' into 'it has exactly as many'.

scales · Categorical-hearing
6 step patterns, 2 category counts, and only the count matters. Each scale drawn as its own steps across the octave, with the share of trials a listener whose internal noise is 11 cents names correctly — and, beside it, the same figure for a scale of equally spaced degrees of the same size. The two agree to three decimal places on every row, though the narrowest step here is 100 cents on the diatonic major, tempered and the widest is 300. Naming is lost at boundaries, an n-degree scale has n of them wherever they are put, and each costs the same as long as no category is narrow enough for the noise to carry an estimate clean across it — 9.1 standard deviations, at the worst here.

A boundary costs the same wherever it is put

Every capacity figure drawn until now cuts the octave into equal categories, and no scale in the world is equal. Putting the real step patterns through the same model returns exactly the same accuracy to four decimal places — because naming is lost at boundaries, an n-degree scale has n of them wherever they are, and each costs the mean absolute value of the noise. That turns the most-quoted result about hearing from a search into one line: 1504 times one minus the criterion, over sigma.

scales · Categorical-hearing
One noise everywhere, and the noise each boundary actually gets. The seven boundaries of the tempered diatonic scale, with the noise the earlier model gives each of them — 11 cents, the same everywhere — and the noise the harmonicity model gives them instead, which runs from 11.0 cents to 35.6. The error rate is the sum of these over the octave rather than seven copies of one, so it rises from 5.1 per cent to 11.8. And the moment the boundaries differ, where they are put matters: a scale that moved its degrees would move its boundaries onto different intervals and would pay a different sum. That is the earlier null broken by the assumption its own last section named.

A boundary beside a fifth

A closed form established earlier says an n-category division costs n·σ·√(2/π) whatever the widths are, so a boundary costs the same wherever it is put and the step pattern cannot matter. Its own last section named the assumption that produces the null: one σ, applied to every boundary in the octave. Let σ follow how securely each interval is held and the formula becomes a sum over boundaries rather than n copies of one — the diatonic's error rises from 5.1 per cent to 11.8, and where the degrees are put matters again.

scales · Categorical-hearing
The best seven of the twelve is a scale nobody has ever used. All 462 ways of choosing seven of the twelve semitones with the tonic fixed, ranked by the identification error the harmonicity model gives them. The best is C C♯ F♯ G A♭ B♭ B at 8.8 per cent and the worst is 13.4; the diatonic major sits at rank 376, in the worse fifth of the ranking, at 11.8. The optimum is a cluster of semitones around the tonic and around the fifth, and the reason is visible in the criterion rather than in music: a boundary next to the unison or the fifth is a boundary with very little noise on it, so the cheapest way to satisfy this measure is to crowd the degrees where the model says the ear is sharpest. A criterion whose optimum is a scale nobody plays is a criterion that is not what scales are chosen for, and the useful reading of this drawing is that rather than its winner.

The best seven of the twelve

Once the noise is allowed to differ from boundary to boundary, a scale can be chosen to minimise identification error — and the choice is a search over four hundred and sixty-two sets rather than an argument. Run, it returns a cluster of semitones around the tonic and around the fifth, and puts the diatonic major at rank 376 of 462, in the worse fifth of the ranking. A criterion whose optimum is a scale nobody has ever played is a criterion that is not what scales are chosen for, and the reason it fails is legible in the model rather than in the music.

scales · Categorical-hearing
Each tradition's own steps against the equal division of the same size. Six scales, each drawn against the equal division into the same number of degrees, under both models of how securely an interval is held. On the harmonicity model the tempered diatonic is 11 per cent better than seven equal steps; on the profile model the same comparison is 0.9 per cent, which is nothing. The two models disagree about the one comparison anybody would want the measure for, and only one of them is free of circularity: the probe-tone profile was measured on listeners raised inside the diatonic tradition, so using it to explain why the diatonic is well chosen assumes the answer. The harmonicity model assumes only that a simple ratio is easier to hold than a complicated one.

The unequal scale that is easier to name

The whole of the earlier result was that a scale's step pattern cannot matter, and every tradition it drew agreed with the equal division of its own size to three decimal places. With one noise per boundary the comparison is live again, and the tempered diatonic beats seven equal steps by eleven per cent — while the pentatonics gain nothing and the maqam scales gain two tenths of one per cent. Only one of the two security models produces the effect, and it is the one that was not measured on listeners raised inside the tradition it is being used to explain.

scales · Categorical-hearing

Named alongside it

The objects these essays reach for when they reach for this one.

CentsDifference limenJust-noticeable differenceMicrotonalityScale degreeIdentificationMaqamCategory boundaryEqual divisionJust intonationBeatCategory width

All concepts