Intervals and chords

How many boxes an octave holds

An identification model of pitch is usually handed twelve categories, and nothing ever asked how many an octave can hold. There are two answers and they are a factor of forty apart. Resolution allows between 151 and 356 — a listener can tell that many pitches apart in a direct comparison. Naming one of them without a comparison is a different faculty and it runs out at six or seven, which is where every mode in every tradition compared here sits. Turkish theory names fifty-three commas to the octave and a makam uses seven of them, and the gap between those two numbers is the whole of the argument.

Assumes: The ear sorts into boxes, and the boxes are the theory · How small a difference is audible

The ear sorts intervals into boxes and the boxes are the theory. That rung drew five boxes and measured the width of the transitions between them; the rung after it found that one acoustic value can belong to two boxes at once. Both figures were handed twelve categories to the octave, because twelve is what the keyboard has.

Nothing on this site has asked how many boxes there could be. It is a question with a number in it, the number can be computed from measurements this collection already carries, and the answer turns out to be two numbers rather than one.

The first ceiling, which is resolution

The obvious limit is how finely a difference can be heard at all. That is the difference limen, and this site has measured what it costs the tuning argument: about five cents in the middle of the range, rising at the bottom and the top.

Divide an octave by it and the answer is large.

The smallest audible difference, and what has to clear it. The difference limen for frequency, converted from Wier, Jesteadt and Green's 1977 fit into cents, against the intervals and commas the rest of these essays argue about. Anything drawn below the curve is a quantity nobody can hear as a change of pitch; anything well above it is a quantity a listener can be asked about. The limen is for pure tones, successive, with trained listeners — the most favourable case there is, and therefore the right one to test a claim against.
Fig. 1 The difference limen in cents across the range. At 250 Hz it is 5.2 cents, at a kilohertz 3.4, at 125 Hz 7.9. Twelve hundred cents divided by those gives 229 boxes at 250 Hz, 356 at a kilohertz, and 151 down at 125 — the number of distinguishable pitches an octave contains, if the only question is whether two of them can be told apart when played one after the other.

Two or three hundred categories in an octave is not a musical proposal, and nobody has ever made it. But it is the right first answer, because it establishes that whatever limits a scale to seven notes or twelve, it is not the ear’s ability to hear the difference. There is room in an octave for forty times as many notes as anybody uses.

The second ceiling, which is naming

Telling two things apart is not the same faculty as naming which one has arrived. A scale degree is a name: hearing it means assigning a heard pitch to one of a fixed set of labels, and the assignment has to survive the pitch arriving on its own.

The identification model this site already draws gives that a number. Its logistic transition — 25 per cent to 75 per cent over about thirty cents for trained listeners — corresponds to an internal estimate carrying about eleven cents of noise. So the question becomes arithmetic: cut the octave into n equal categories, let a pitch fall anywhere in its own band, add eleven cents of noise, and ask how often the nearest category centre is the right one.

How many notes an octave can hold, asked twice. The share of trials on which a category is named correctly, against how many equal categories the octave is cut into, for a listener whose internal estimate carries 11 cents of noise — the logistic scale this site's identification figures already use, which is the thirty-cent transition the studies report. At 95 per cent accuracy the ceiling is 6 categories, and seven scores 94.9 per cent — on the line. The other ceiling is resolution: 151 to 356 difference limens fit in an octave depending on register, which is a factor of forty larger. Every system marked below sits between the two, and the marks separate: the number of degrees a mode uses clears the criterion, and the size of the gamut it chooses them from does not.
Fig. 2 The answer, against n. Five categories are named correctly on 96.3 per cent of trials, seven on 94.9, twelve on 91.2, twenty-four on 82.4 and fifty-three on 61.9. At a criterion of 95 per cent the ceiling is six categories, and seven sits within a tenth of a per cent of it. The marks along the bottom are what real systems use, and they fall into two groups.

Six or seven is a very different number from three hundred, and it is the number scales actually have.

The accuracy has a closed form, and it depends only on the count

The model is integrated numerically above, and it need not be. Running it on a variety of category sets makes the reason visible.

A misnaming happens when eleven cents of noise carries the estimate across a boundary, and it can only happen near one — in the middle of a category, eleven cents moves nothing. So the total error is a fixed loss at each boundary, and the number of boundaries is the number of categories whatever their spacing. That gives

accuracy = 1 − n · σ√(2/π) / 1200,

which for σ = 11 cents is a loss of 8.78 cents of the octave per category. Seven categories lose 61.5 of 1200, which is 94.9 per cent — the figure the numerical integration returns, to the digit.

Which means an unequal mode scores exactly the same

That has a consequence the caveats below get wrong. A diatonic mode’s semitone steps are predicted there to score worse than its whole-tone steps, and they do — but the mode does not.

degrees accuracy
seven equal 7 94.9%
diatonic, 2 2 1 2 2 2 1 7 94.9%
harmonic minor 7 94.9%
hijaz, 1 3 1 2 2 2 1 7 94.9%
pentatonic, 2 2 3 2 3 5 96.3%
slendro, near-equal 5 96.3%

Identical, to three figures, and the reason is in the per-band numbers: a 150-cent band is named right 94.1 per cent of the time and a 200-cent band 95.6, exactly as the caveat expects — and the narrow bands carry proportionally less of the octave, by exactly the amount that compensates. Unevenness moves the errors around and does not change how many there are.

Pushing it to an absurd partition confirms it. Six categories crammed into the first hundred cents with the seventh spanning the remaining eleven hundred still scores 95.0 per cent. The model is insensitive to spacing by construction, which is a strength for the argument — the count really is the whole of it — and a limitation worth naming, since a real listener surely does find a hundred-cent category harder than a two-hundred-cent one.

And the rule of thumb is off by a factor of four

The last caveat gives the ceiling as roughly 1200 divided by four times the noise. Solving the closed form at a 95 per cent criterion gives

n ≤ 0.05 × 1200 / (σ√(2/π)) = 75 / σ,

which is 1200 over sixteen times the noise, not four. At σ = 11 that is 6.8, so six — which is the answer the curve gives. The four-times version would predict twenty-seven categories at the same noise, which is four times too many and would put the ceiling well above every gamut in the table rather than below all of them.

The relation the caveat was reaching for is right and worth keeping: the ceiling is inversely proportional to the noise, and every tradition’s mode count sits at or just under it. The constant is 75, and it checks at every noise level tried — 12 categories at σ = 6, 9 at 8, 5 at 15, 3 at 20.

It is worth being precise about what has and has not been imported here. Nothing in this calculation is Miller’s “seven plus or minus two”, and nothing in it is a claim about channel capacity in general. The only quantity taken from outside is the width of one identification transition, measured on intervals; everything after that is arithmetic. That the answer lands where the famous number lands is a coincidence worth noticing and not an argument, and treating it as one would be borrowing a result from a different experiment on a different continuum.

The other thing the model does not do is decide where the categories go. It is told they are equal and asked how many fit; a real mode’s degrees are unequally spaced, and a scale is not a set of pitches but a set of positions a tradition treats as structural. The count is the part with a ceiling.

Two things called the size of the scale

The marks separate cleanly and the separation is the finding. Every system in this site’s own collection of traditions carries two numbers, and only one of them is about categories.

Notes in the gamut Degrees in one mode Named correctly
Javanese slendro 5 5 96.3%
Western common practice 12 7 94.9%
Hindustani theory 22 5 96.3%
Arabic theory 24 7 94.9%
Turkish theory 53 7 94.9%

The right-hand column is the accuracy for the mode, and every entry clears the criterion. The accuracy for the gamut clears it in exactly one case — slendro, where the gamut is the mode.

So the ceiling applies to the mode and not to the gamut, and the two have been called the same thing throughout the literature and throughout this site. A gamut is a supply of pitches, a notation, a set of positions a note can be tuned to. A mode is a set of categories a listener assigns, and the assignment is what has a capacity.

Turkish theory is the clearest case because the numbers are furthest apart. The Arel–Ezgi–Uzdilek system names fifty-three Holdrian commas to the octave, each 22.64 cents. A listener asked to name one of fifty-three would be right three times in five. A makam uses seven of them, and seven is comfortable. Fifty-three is a notation for describing where the seven are, not a set of things to be heard as things.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 204 that separate adjacent categories — so the change of mind happens in 12% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 3 The mode rather than the gamut, drawn as the boxes a listener actually assigns. Seven degrees of maqam Rast in the Turkish naming, spaced as that theory places them, with the identification model at the width the studies report. Every crossing is comfortably narrower than the gap it sits in — which is what a 94.9 per cent naming score looks like as a picture. The same seven degrees are written in a fifty-three-comma gamut in one tradition and a twenty-four-quarter-tone gamut in another, and the two disagree about the third degree by a third of a semitone. Two gamuts, one mode, and the mode is the part with a capacity.

Where twelve sits, and why it is uncomfortable

Twelve scores 91.2 per cent, which is below the criterion and above the gamut sizes. That is the right result rather than an awkward one, and it matches something every musician knows: naming a chromatic degree heard in isolation is genuinely hard, and nobody does it that way.

What is done instead is to hear the degree inside a mode — seven categories, not twelve — with the other five arriving as inflections of those seven. The notation says so: a chromatic note is spelled as a sharpened or flattened degree, never as a category of its own, which is exactly what the vertical axis of the stave encodes. Seven positions for twelve pitches is not a historical accident of the page; it is the page agreeing with the capacity.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 4 The identification model at the resolution the studies report, on the five categories used at the outset. The transition from one name to the next takes about thirty cents, which is six difference limens: the boundary is not sharp, and its width is the whole of what the capacity calculation depends on.
Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 5 All twelve as categories, which is where the 91.2 per cent comes from and why it is uncomfortable. The crossings are the same width as in the seven-degree figure above and the gaps between centres are smaller, so a larger share of the axis is in doubt — naming a chromatic degree heard in isolation is genuinely hard, and nobody does it that way. What is done instead is to hear the degree inside a mode of seven, with the other five arriving as inflections of those seven; the notation says so, since a chromatic note is spelled as a sharpened or flattened degree and never as a category of its own. Seven positions on a stave for twelve pitches is the page agreeing with the capacity.

So the perceptual argument sets a ceiling somewhere near seven and the instruments carry twelve, nineteen, thirty-one and fifty-three. Those numbers were not chosen by anybody counting categories, and the reason they exist is arithmetic about the fifth rather than anything a listener can do.

Every equal division from 5 to 60, and how wrong it is. For each number of equal steps in the octave, how far its best fifth and its best major third fall from the pure ratios, in cents. The divisions people have actually used are the ones with small errors in both, and no other criterion was applied to pick them out.
Fig. 6 Why the large gamuts exist at all, which is not a perceptual reason. Each equal division is scored by how far its best fifth falls from a pure one, and the good ones — 12, 19, 31, 41, 53 — are the convergents of an irrational number. Fifty-three is on that list because of arithmetic about the fifth, not because anybody proposed hearing fifty-three things. The gamut is chosen to make the intervals right; the mode is chosen to be nameable; and the two criteria have nothing to do with each other.

Seven, from a third direction

Seven has now been arrived at three times on this site by three unrelated routes, and this is the third.

The combinatorial route ran a census over every subset of the twelve for four structural properties and found exactly one set with all four, at seven notes. The perceptual route is the one above: seven is the largest number of equal categories that clears a 95 per cent naming criterion at the measured category width. And the practical route is the table — every mode in every tradition entered here has five or seven degrees.

None of the three knows about the others. The census was arithmetic about a cyclic group; the capacity is arithmetic about a Gaussian; the table is what people play.

The suspicion to raise about that is the obvious one: three routes to the same integer are three routes only if they have no shared premise. These two very nearly do not. The census’s answer depends on the universe being twelve; the capacity’s does not depend on twelve at all — it would give the same six or seven in a universe of nineteen or of fifty-three, because it is about the width of a category in cents rather than about a division. What the two share is the octave, and only as the thing being divided.

Which makes the pairing a little sharper than a coincidence and a good deal weaker than a derivation. Something has to explain why seven notes chosen unevenly from twelve is the survivor of a structural census and sits at the naming ceiling, and neither argument explains the other.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 243 that separate adjacent categories — so the change of mind happens in 10% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 7 The other end of the range, on one measured Javanese slendro. Five degrees roughly 240 cents apart put the crossings a long way from the centres, and the naming score is 96.3 per cent — the one case in the table where the gamut and the mode are the same object and the gamut clears the criterion. That is the coincidence worth being suspicious of: a combinatorial census over the subsets of twelve arrives at seven and knows nothing about listeners; this curve arrives at six or seven from the width of a category and knows nothing about intervals. The two share the octave and nothing else, which makes the pairing sharper than an accident and much weaker than a derivation.
Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 8 The same fact seen from inside a single box. A just major third at 386 cents and an equal-tempered one at 400 are fourteen cents apart and both sit in the flat middle of the same category, where identification is at or near certainty. A category wide enough to hold a fourteen-cent error is a category there cannot be many of — which is the capacity curve stated the other way round, and it is why an earlier essay could say that a mistuned third stays a third.

One consequence of reading the ceiling this way is worth stating because it is a prediction rather than a description. If the capacity belongs to the mode and not to the gamut, then a tradition can enlarge its gamut indefinitely without asking anything more of a listener — and the traditions that have the largest gamuts are precisely the ones with the most elaborate written theory and the least elaborate modes. Turkish theory’s fifty-three commas exist to record where seven degrees are placed and to distinguish one makam’s placement from another’s; they are never all in play at once, and no piece asks a listener to tell the twenty-ninth comma from the thirtieth. A gamut is a coordinate system and a mode is a vocabulary, and only the second is spoken.

The prediction that follows is that a tradition which genuinely used its whole gamut as categories would have to have a small one. That is what slendro is, and it is the only row in the table where the two columns agree.

Which computation produced the numbers

The resolution ceiling is Wier, Jesteadt and Green’s fit for the frequency difference limen, converted to cents and divided into 1200. It is the same function every acuity figure on this site uses.

The naming model is a decision rule rather than a curve fit. A category is a band of width 1200/n cents; a stimulus falls uniformly within its own band; the listener’s estimate is that value plus Gaussian noise of standard deviation σ; the response is the nearest category centre. Integrating over the band gives the accuracy, and the integral is evaluated numerically over four hundred points, which is exact to more digits than the input deserves.

σ = 11 cents is the one number from outside, and it is not an independent import: it is the logistic scale this site’s own category figures already use, chosen there to reproduce the roughly thirty-cent transition the identification studies report for trained listeners. So the capacity result and the boundary-width result rest on the same measurement, and if that measurement is wrong both move together.

The degree counts are read off this site’s own entries for those traditions rather than asserted. The gamut sizes are the theoretical divisions each system names.

Whose listeners, and whose music

The thirty-cent transition is for trained listeners in laboratory conditions, identifying intervals played in isolation. Untrained listeners have wider transitions, so their capacity is lower; a tradition’s own practitioners, hearing degrees inside a familiar mode with a drone or a reference, do far better than the model allows, because they are not doing the task the model describes.

That last point bounds the whole argument. The capacity computed here is for absolute assignment of an isolated interval to a category, and almost no real listening is that. A raga is heard over a tanpura, a maqam over a tonic, a Western melody inside an established key — every one of them a reference against which the judgement is relative, which is a different and easier task.

Which is why the coincidence at seven should be treated carefully. The honest statement is narrower than “seven because of channel capacity”: it is that the number of degrees traditions use sits at the edge of what isolated identification can do, and the size of the gamuts they use does not — and that the two numbers being different is what needs explaining rather than either number alone.

What the picture cannot show

Equal categories are an idealisation, and it turns out to be a harmless one. A mode with unequal steps has categories of unequal width, and the wide ones really are easier to name than the narrow ones — but the weighted total is unchanged, because the error is a boundary effect and the number of boundaries does not depend on the spacing. The section above computes the diatonic, the harmonic minor and hijaz and gets 94.9 per cent for all three. What that exposes is a limitation of the model rather than a result: it cannot express the intuition that a hundred-cent category is harder than a two-hundred-cent one, because in it only the count matters.

Noise is taken as constant in cents. The difference limen is not: it is nearly twice as large at 125 Hz as at a kilohertz, so the capacity is genuinely register-dependent and the curve here is a middle-of-the-range figure.

Nothing here is about learning. Category boundaries are learned, half of consonance is learned, and a listener raised inside a twenty-two-śruti practice is not the listener this model describes. What the model would predict is that such a listener does better on the degrees of their own mode and no better on the gamut, which is testable and not tested here.

And the criterion is a choice. Ninety-five per cent gives six; ninety per cent gives thirteen; ninety-nine per cent gives nothing above one, because eleven cents of noise on a stimulus that may sit anywhere in its band cannot be that reliable at any division. The shape of the curve is the result and the single integer read off it is an artefact of where the line is drawn — which is why the table matters more than the ceiling.

The noise figure moves it just as hard. At six cents of internal noise the ceiling is twelve; at twenty it is three. So the result is not “seven” but a relation — and the closed form above gives it exactly: the ceiling is 75 over the noise in cents, which is 1200 divided by sixteen times the noise rather than four. Every tradition’s mode count sits at or just under it for the noise its own listeners have.

Where this ladder goes next

Three rungs of this ladder have taken the categories as given: their existence, their boundaries, and the case where one sound belongs to two of them. This one counts them and finds that the count is limited by naming rather than by hearing.

There is one more consequence worth stating, because it turns the result into a prediction rather than a description. If the ceiling is set by naming and not by hearing, then a tradition that wants more than seven degrees has to buy them with something other than resolution — a drone, a fixed reference, an instrument whose frets do the assignment. And that is what the traditions with large gamuts have: the tanpura under a raga, the tonic under a maqam, the fretting of a tanbur. Each of those converts an absolute identification into a relative one, which is the task the model says is easy and the model here does not describe.

What is not yet asked is whether the boundaries move. The model here has fixed centres and a fixed noise; real category boundaries shift with context, with the preceding interval and with the tuning system a listener has just been hearing. The measurement that would settle it is an identification run after an adapting sequence, and this site has the machinery for both halves of it and has never put them together.

Part 4 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 11.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionDifference limenEnumerationEqual divisionJust-noticeable differenceMaqamMicrotonalityScale degree