How many boxes an octave holds
Assumes: The ear sorts into boxes, and the boxes are the theory · How small a difference is audible
The ear sorts intervals into boxes and the boxes are the theory. That rung drew five boxes and measured the width of the transitions between them; the rung after it found that one acoustic value can belong to two boxes at once. Both figures were handed twelve categories to the octave, because twelve is what the keyboard has.
Nothing on this site has asked how many boxes there could be. It is a question with a number in it, the number can be computed from measurements this collection already carries, and the answer turns out to be two numbers rather than one.
The first ceiling, which is resolution
The obvious limit is how finely a difference can be heard at all. That is the difference limen, and this site has measured what it costs the tuning argument: about five cents in the middle of the range, rising at the bottom and the top.
Divide an octave by it and the answer is large.
Two or three hundred categories in an octave is not a musical proposal, and nobody has ever made it. But it is the right first answer, because it establishes that whatever limits a scale to seven notes or twelve, it is not the ear’s ability to hear the difference. There is room in an octave for forty times as many notes as anybody uses.
The second ceiling, which is naming
Telling two things apart is not the same faculty as naming which one has arrived. A scale degree is a name: hearing it means assigning a heard pitch to one of a fixed set of labels, and the assignment has to survive the pitch arriving on its own.
The identification model this site already draws gives that a number. Its logistic transition — 25 per cent to 75 per cent over about thirty cents for trained listeners — corresponds to an internal estimate carrying about eleven cents of noise. So the question becomes arithmetic: cut the octave into n equal categories, let a pitch fall anywhere in its own band, add eleven cents of noise, and ask how often the nearest category centre is the right one.
Six or seven is a very different number from three hundred, and it is the number scales actually have.
The accuracy has a closed form, and it depends only on the count
The model is integrated numerically above, and it need not be. Running it on a variety of category sets makes the reason visible.
A misnaming happens when eleven cents of noise carries the estimate across a boundary, and it can only happen near one — in the middle of a category, eleven cents moves nothing. So the total error is a fixed loss at each boundary, and the number of boundaries is the number of categories whatever their spacing. That gives
accuracy = 1 − n · σ√(2/π) / 1200,
which for σ = 11 cents is a loss of 8.78 cents of the octave per category. Seven categories lose 61.5 of 1200, which is 94.9 per cent — the figure the numerical integration returns, to the digit.
Which means an unequal mode scores exactly the same
That has a consequence the caveats below get wrong. A diatonic mode’s semitone steps are predicted there to score worse than its whole-tone steps, and they do — but the mode does not.
| degrees | accuracy | |
|---|---|---|
| seven equal | 7 | 94.9% |
| diatonic, 2 2 1 2 2 2 1 | 7 | 94.9% |
| harmonic minor | 7 | 94.9% |
| hijaz, 1 3 1 2 2 2 1 | 7 | 94.9% |
| pentatonic, 2 2 3 2 3 | 5 | 96.3% |
| slendro, near-equal | 5 | 96.3% |
Identical, to three figures, and the reason is in the per-band numbers: a 150-cent band is named right 94.1 per cent of the time and a 200-cent band 95.6, exactly as the caveat expects — and the narrow bands carry proportionally less of the octave, by exactly the amount that compensates. Unevenness moves the errors around and does not change how many there are.
Pushing it to an absurd partition confirms it. Six categories crammed into the first hundred cents with the seventh spanning the remaining eleven hundred still scores 95.0 per cent. The model is insensitive to spacing by construction, which is a strength for the argument — the count really is the whole of it — and a limitation worth naming, since a real listener surely does find a hundred-cent category harder than a two-hundred-cent one.
And the rule of thumb is off by a factor of four
The last caveat gives the ceiling as roughly 1200 divided by four times the noise. Solving the closed form at a 95 per cent criterion gives
n ≤ 0.05 × 1200 / (σ√(2/π)) = 75 / σ,
which is 1200 over sixteen times the noise, not four. At σ = 11 that is 6.8, so six — which is the answer the curve gives. The four-times version would predict twenty-seven categories at the same noise, which is four times too many and would put the ceiling well above every gamut in the table rather than below all of them.
The relation the caveat was reaching for is right and worth keeping: the ceiling is inversely proportional to the noise, and every tradition’s mode count sits at or just under it. The constant is 75, and it checks at every noise level tried — 12 categories at σ = 6, 9 at 8, 5 at 15, 3 at 20.
It is worth being precise about what has and has not been imported here. Nothing in this calculation is Miller’s “seven plus or minus two”, and nothing in it is a claim about channel capacity in general. The only quantity taken from outside is the width of one identification transition, measured on intervals; everything after that is arithmetic. That the answer lands where the famous number lands is a coincidence worth noticing and not an argument, and treating it as one would be borrowing a result from a different experiment on a different continuum.
The other thing the model does not do is decide where the categories go. It is told they are equal and asked how many fit; a real mode’s degrees are unequally spaced, and a scale is not a set of pitches but a set of positions a tradition treats as structural. The count is the part with a ceiling.
Two things called the size of the scale
The marks separate cleanly and the separation is the finding. Every system in this site’s own collection of traditions carries two numbers, and only one of them is about categories.
| Notes in the gamut | Degrees in one mode | Named correctly | |
|---|---|---|---|
| Javanese slendro | 5 | 5 | 96.3% |
| Western common practice | 12 | 7 | 94.9% |
| Hindustani theory | 22 | 5 | 96.3% |
| Arabic theory | 24 | 7 | 94.9% |
| Turkish theory | 53 | 7 | 94.9% |
The right-hand column is the accuracy for the mode, and every entry clears the criterion. The accuracy for the gamut clears it in exactly one case — slendro, where the gamut is the mode.
So the ceiling applies to the mode and not to the gamut, and the two have been called the same thing throughout the literature and throughout this site. A gamut is a supply of pitches, a notation, a set of positions a note can be tuned to. A mode is a set of categories a listener assigns, and the assignment is what has a capacity.
Turkish theory is the clearest case because the numbers are furthest apart. The Arel–Ezgi–Uzdilek system names fifty-three Holdrian commas to the octave, each 22.64 cents. A listener asked to name one of fifty-three would be right three times in five. A makam uses seven of them, and seven is comfortable. Fifty-three is a notation for describing where the seven are, not a set of things to be heard as things.
Where twelve sits, and why it is uncomfortable
Twelve scores 91.2 per cent, which is below the criterion and above the gamut sizes. That is the right result rather than an awkward one, and it matches something every musician knows: naming a chromatic degree heard in isolation is genuinely hard, and nobody does it that way.
What is done instead is to hear the degree inside a mode — seven categories, not twelve — with the other five arriving as inflections of those seven. The notation says so: a chromatic note is spelled as a sharpened or flattened degree, never as a category of its own, which is exactly what the vertical axis of the stave encodes. Seven positions for twelve pitches is not a historical accident of the page; it is the page agreeing with the capacity.
So the perceptual argument sets a ceiling somewhere near seven and the instruments carry twelve, nineteen, thirty-one and fifty-three. Those numbers were not chosen by anybody counting categories, and the reason they exist is arithmetic about the fifth rather than anything a listener can do.
Seven, from a third direction
Seven has now been arrived at three times on this site by three unrelated routes, and this is the third.
The combinatorial route ran a census over every subset of the twelve for four structural properties and found exactly one set with all four, at seven notes. The perceptual route is the one above: seven is the largest number of equal categories that clears a 95 per cent naming criterion at the measured category width. And the practical route is the table — every mode in every tradition entered here has five or seven degrees.
None of the three knows about the others. The census was arithmetic about a cyclic group; the capacity is arithmetic about a Gaussian; the table is what people play.
The suspicion to raise about that is the obvious one: three routes to the same integer are three routes only if they have no shared premise. These two very nearly do not. The census’s answer depends on the universe being twelve; the capacity’s does not depend on twelve at all — it would give the same six or seven in a universe of nineteen or of fifty-three, because it is about the width of a category in cents rather than about a division. What the two share is the octave, and only as the thing being divided.
Which makes the pairing a little sharper than a coincidence and a good deal weaker than a derivation. Something has to explain why seven notes chosen unevenly from twelve is the survivor of a structural census and sits at the naming ceiling, and neither argument explains the other.
One consequence of reading the ceiling this way is worth stating because it is a prediction rather than a description. If the capacity belongs to the mode and not to the gamut, then a tradition can enlarge its gamut indefinitely without asking anything more of a listener — and the traditions that have the largest gamuts are precisely the ones with the most elaborate written theory and the least elaborate modes. Turkish theory’s fifty-three commas exist to record where seven degrees are placed and to distinguish one makam’s placement from another’s; they are never all in play at once, and no piece asks a listener to tell the twenty-ninth comma from the thirtieth. A gamut is a coordinate system and a mode is a vocabulary, and only the second is spoken.
The prediction that follows is that a tradition which genuinely used its whole gamut as categories would have to have a small one. That is what slendro is, and it is the only row in the table where the two columns agree.
Which computation produced the numbers
The resolution ceiling is Wier, Jesteadt and Green’s fit for the frequency difference limen, converted to cents and divided into 1200. It is the same function every acuity figure on this site uses.
The naming model is a decision rule rather than a curve fit. A category is a band of width 1200/n cents; a stimulus falls uniformly within its own band; the listener’s estimate is that value plus Gaussian noise of standard deviation σ; the response is the nearest category centre. Integrating over the band gives the accuracy, and the integral is evaluated numerically over four hundred points, which is exact to more digits than the input deserves.
σ = 11 cents is the one number from outside, and it is not an independent import: it is the logistic scale this site’s own category figures already use, chosen there to reproduce the roughly thirty-cent transition the identification studies report for trained listeners. So the capacity result and the boundary-width result rest on the same measurement, and if that measurement is wrong both move together.
The degree counts are read off this site’s own entries for those traditions rather than asserted. The gamut sizes are the theoretical divisions each system names.
Whose listeners, and whose music
The thirty-cent transition is for trained listeners in laboratory conditions, identifying intervals played in isolation. Untrained listeners have wider transitions, so their capacity is lower; a tradition’s own practitioners, hearing degrees inside a familiar mode with a drone or a reference, do far better than the model allows, because they are not doing the task the model describes.
That last point bounds the whole argument. The capacity computed here is for absolute assignment of an isolated interval to a category, and almost no real listening is that. A raga is heard over a tanpura, a maqam over a tonic, a Western melody inside an established key — every one of them a reference against which the judgement is relative, which is a different and easier task.
Which is why the coincidence at seven should be treated carefully. The honest statement is narrower than “seven because of channel capacity”: it is that the number of degrees traditions use sits at the edge of what isolated identification can do, and the size of the gamuts they use does not — and that the two numbers being different is what needs explaining rather than either number alone.
What the picture cannot show
Equal categories are an idealisation, and it turns out to be a harmless one. A mode with unequal steps has categories of unequal width, and the wide ones really are easier to name than the narrow ones — but the weighted total is unchanged, because the error is a boundary effect and the number of boundaries does not depend on the spacing. The section above computes the diatonic, the harmonic minor and hijaz and gets 94.9 per cent for all three. What that exposes is a limitation of the model rather than a result: it cannot express the intuition that a hundred-cent category is harder than a two-hundred-cent one, because in it only the count matters.
Noise is taken as constant in cents. The difference limen is not: it is nearly twice as large at 125 Hz as at a kilohertz, so the capacity is genuinely register-dependent and the curve here is a middle-of-the-range figure.
Nothing here is about learning. Category boundaries are learned, half of consonance is learned, and a listener raised inside a twenty-two-śruti practice is not the listener this model describes. What the model would predict is that such a listener does better on the degrees of their own mode and no better on the gamut, which is testable and not tested here.
And the criterion is a choice. Ninety-five per cent gives six; ninety per cent gives thirteen; ninety-nine per cent gives nothing above one, because eleven cents of noise on a stimulus that may sit anywhere in its band cannot be that reliable at any division. The shape of the curve is the result and the single integer read off it is an artefact of where the line is drawn — which is why the table matters more than the ceiling.
The noise figure moves it just as hard. At six cents of internal noise the ceiling is twelve; at twenty it is three. So the result is not “seven” but a relation — and the closed form above gives it exactly: the ceiling is 75 over the noise in cents, which is 1200 divided by sixteen times the noise rather than four. Every tradition’s mode count sits at or just under it for the noise its own listeners have.
Where this ladder goes next
Three rungs of this ladder have taken the categories as given: their existence, their boundaries, and the case where one sound belongs to two of them. This one counts them and finds that the count is limited by naming rather than by hearing.
There is one more consequence worth stating, because it turns the result into a prediction rather than a description. If the ceiling is set by naming and not by hearing, then a tradition that wants more than seven degrees has to buy them with something other than resolution — a drone, a fixed reference, an instrument whose frets do the assignment. And that is what the traditions with large gamuts have: the tanpura under a raga, the tonic under a maqam, the fretting of a tanbur. Each of those converts an absolute identification into a relative one, which is the task the model says is easy and the model here does not describe.
What is not yet asked is whether the boundaries move. The model here has fixed centres and a fixed noise; real category boundaries shift with context, with the preceding interval and with the tuning system a listener has just been hearing. The measurement that would settle it is an identification run after an adapting sequence, and this site has the machinery for both halves of it and has never put them together.
Part 4 of 11
One essay in the series on Categorical-hearing. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 11.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Categorical perceptionDifference limenEnumerationEqual divisionJust-noticeable differenceMaqamMicrotonalityScale degree
- How much an anchor would have to be worth categorical perception, difference limen, just-noticeable difference, scale degree
- The best seven of the twelve categorical perception, difference limen, enumeration, scale degree
- A scale built downward from a fourth equal division, maqam, microtonality
- The part of the error a key cannot touch difference limen, just-noticeable difference, scale degree
- The quantity a rival account says is not there difference limen, just-noticeable difference, scale degree
- The unequal scale that is easier to name categorical perception, maqam, scale degree