Scales and modes

The best seven of the twelve

Once the noise is allowed to differ from boundary to boundary, a scale can be chosen to minimise identification error — and the choice is a search over four hundred and sixty-two sets rather than an argument. Run, it returns a cluster of semitones around the tonic and around the fifth, and puts the diatonic major at rank 376 of 462, in the worse fifth of the ranking. A criterion whose optimum is a scale nobody has ever played is a criterion that is not what scales are chosen for, and the reason it fails is legible in the model rather than in the music.

Assumes: A boundary beside a fifth · How many boxes an octave holds

The previous rung broke the eighth rung’s null. With one internal noise applied to every boundary, an n-degree scale costs n·σ·√(2/π) whatever its steps are; with a noise that follows how securely each interval is held, the cost is a sum over boundaries and the arrangement decides which seven are summed.

That makes a scale a thing to optimise. There are exactly four hundred and sixty-two ways of taking seven of the twelve semitones with the tonic fixed, each has one error, and the ranking is a loop rather than an argument.

The best seven of the twelve is a scale nobody has ever used. All 462 ways of choosing seven of the twelve semitones with the tonic fixed, ranked by the identification error the harmonicity model gives them. The best is C C♯ F♯ G A♭ B♭ B at 8.8 per cent and the worst is 13.4; the diatonic major sits at rank 376, in the worse fifth of the ranking, at 11.8. The optimum is a cluster of semitones around the tonic and around the fifth, and the reason is visible in the criterion rather than in music: a boundary next to the unison or the fifth is a boundary with very little noise on it, so the cheapest way to satisfy this measure is to crowd the degrees where the model says the ear is sharpest. A criterion whose optimum is a scale nobody plays is a criterion that is not what scales are chosen for, and the useful reading of this drawing is that rather than its winner.
Fig. 1 All 462 seven-of-twelve sets with the tonic fixed, ranked by the identification error the harmonicity model gives them, with the diatonic major marked.

The answer is absurd, and the shape of the absurdity is the result.

The winner

The best set is C, C♯, F♯, G, A♭, B♭, B — two semitone clusters, one at the tonic and one at the tritone-to-fifth, with a four-semitone hole between them. Its error is 8.8 per cent against the diatonic’s 11.8 and the worst set’s 13.4.

Nothing plays that. It has no fifth from any degree but the tonic, it has three consecutive semitones, and its largest step is a major third. Whatever it is, it is not a scale in any sense a tradition would recognise.

The diatonic major sits at rank 376 of 462 — in the worse fifth of the ranking, closer to the worst set than to the best.

Why the search finds a cluster

The mechanism is visible and it is a property of the criterion rather than of hearing.

A boundary’s noise is set by how securely the interval at that boundary is held, and the security function has two large peaks: the unison, at weight one, and the fifth, at 0.39. Everything else is between 0.10 and 0.28.

So the cheapest possible boundary is one sitting on or beside the unison, and the way to put a boundary beside the unison is to put two degrees a semitone apart there — the boundary between them falls at fifty cents, which the model reads as the unison’s own security. The winner’s seven boundaries are at 50, 350, 650, 750, 900, 1050 and 1150 cents, and the two at fifty and eleven hundred and fifty are given a σ of exactly eleven cents while the rest run to twenty-nine. The same trick works at the fifth.

The optimiser is not choosing a scale. It is choosing where to put boundaries, and the degrees are only the means. A cluster is the cheapest arrangement because it puts the most boundaries in the cheapest places, and a scale spread across the octave is forced to put most of its boundaries in expensive ones.

What the search would have returned two rungs ago

It is worth running the same enumeration under the eighth rung’s assumption, because the answer is instructive and takes no computation at all.

6 step patterns, 2 category counts, and only the count matters. Each scale drawn as its own steps across the octave, with the share of trials a listener whose internal noise is 11 cents names correctly — and, beside it, the same figure for a scale of equally spaced degrees of the same size. The two agree to three decimal places on every row, though the narrowest step here is 100 cents on the diatonic major, tempered and the widest is 300. Naming is lost at boundaries, an n-degree scale has n of them wherever they are put, and each costs the same as long as no category is narrow enough for the noise to carry an estimate clean across it — 9.1 standard deviations, at the worst here.
Fig. 2 The earlier figure: every scale against the equal division of its own size, under one internal noise. The columns agree to three decimal places, which is what a criterion that cannot distinguish arrangements looks like.

Under one σ every one of the 462 sets has the same error — seven boundaries, each costing σ√(2/π) — so the ranking is a flat line and the search has no answer. That is the null in its strongest form: not that the diatonic is unremarkable, but that nothing can be remarkable, because the criterion assigns every arrangement the same number.

So the search only exists because of the ninth rung, and the absurd winner it returns is a direct consequence of the same relaxation that made the search possible. A criterion that could not rank arrangements has been replaced by one that ranks them wrongly, which is a specific kind of progress and is worth naming as one.

What that says about the criterion

A criterion whose optimum is a scale nobody has ever played is not a description of what scales are for.

That is the useful reading of this rung and it is worth stating without hedging. Identification error is not what a scale is chosen to minimise. If it were, music would use clustered scales, and the arithmetic here says clustered scales would be measurably easier to name.

There are three obvious things a scale is chosen for that this criterion has no term for, and each of them would penalise the cluster heavily.

A scale is chosen to supply intervals, and the diatonic set’s own census is about which intervals it contains and how often. The cluster contains almost none of the consonances and six copies of the semitone.

A scale is chosen to be transposable, and two sizes of every step is why the names work. The cluster has no such structure at all.

And a scale is chosen to be sung, which is a constraint on step sizes that nothing here contains — a melody is nearly all small steps and a scale with a four-semitone hole in it makes that impossible.

So the search has not found that the diatonic is badly designed. It has found that this measure is one of several and is not the binding one, which is a thing worth knowing about a measure this ladder has spent nine rungs building.

The spread, which is the size of what is at stake

Between the best set and the worst there is 4.6 percentage points — 8.8 against 13.4 — on a criterion whose absolute level is a normalisation.

Read as a ratio that is a factor of 1.5 between the easiest and hardest seven-note scale to identify, and the diatonic sits at 11.8, which is 65 per cent of the way from the best to the worst.

Not every interval of the octave is held equally securely. Two models of how well a listener holds each interval, drawn as a weight against the unison. The harmonicity model scores an interval by the simplest just ratio near it, one over its Tenney height — so the fifth is at 0.39, the major third at 0.23 and the tritone at 0.10, which is the least secure interval in the octave. The profile model scores it by the probe-tone stability of the pitch class, which is a measurement on listeners rather than a claim about ratios. Neither is a measurement of a listener's noise at that interval, and both are stated so that the earlier assumption — one noise everywhere — can be replaced by something rather than by nothing. The two agree that the unison is the securest and disagree about nearly everything else.
Fig. 3 The security function the whole ranking is generated from. Its two peaks are what the search exploits and its floor at the tritone is what every scale spread across the octave has to pay.

Whether a factor of 1.5 in identification error is musically consequential is a question this ladder can answer with its own arithmetic. The second rung found that a major third can be seventeen cents wrong and still be named a major third, which is what makes temperament possible; the spread here corresponds to a change in effective noise of about a factor of 1.5, which is the difference between an eleven-cent listener and a sixteen-cent one.

The training sweep already priced that gap: it is roughly the difference between a trained listener and a moderately trained one. So the whole range this search covers, across every seven-note scale on the twelve-semitone grid, is smaller than the difference between two listeners.

That is the strongest reason to think the criterion is not what scales are chosen for. A design pressure that a year of ear training would erase is not a design pressure a tradition would organise itself around.

The defect that is in the model rather than in the answer

There is also a specific error in how σ is assigned, and the cluster exposes it.

A boundary separates two categories. The noise on a listener’s decision between a major third and a fourth is presumably some combination of how well each of those is held — not a property of the 450-cent point that happens to lie between them. The model assigns σ by looking up the security of the boundary’s own interval from the tonic, which is a different quantity and is the thing the cluster exploits.

Fifty cents is not the unison. A boundary at fifty cents separates a unison from a semitone, and a listener deciding between those two is not enjoying the unison’s security — they are making the hardest fine discrimination in the octave, between two categories fifty cents apart. The model gives that boundary the least noise available and it should have the most.

Fixing it means scoring a boundary by the two categories it separates and by how far apart they are, which is a different model with a different optimum. That model is not built here, and naming the defect is what this rung can honestly do with the space it has.

The worst set, which is also not a scale

The bottom of the ranking is as informative as the top and is easier to read, because nothing exploits anything down there.

The worst set is C, D, E♭, E, G, A, B♭ at 13.4 per cent — a scale with three consecutive semitones in the middle and a minor third between the fourth and fifth degrees. Its boundaries land at 100, 250, 350, 550, 800, 950 and 1100 cents, which puts three of them in the expensive middle of the octave and none of them anywhere near the unison or the fifth.

So the ranking’s two extremes are both clusters and they differ in where the cluster is. A cluster at the tonic and the fifth is the best arrangement available; a cluster in the middle of the octave is the worst. The whole spread of 4.6 percentage points is a statement about where in the octave the degrees are crowded rather than about whether they are.

That reading is the most defensible thing the search produces, and it is one the model’s defect does not touch — because both extremes suffer the defect equally, and the difference between them is the security function’s shape rather than its lookup rule.

What survives the defect

One comparison does survive, and it is the one the next rung is about.

Each tradition's own steps against the equal division of the same size. Six scales, each drawn against the equal division into the same number of degrees, under both models of how securely an interval is held. On the harmonicity model the tempered diatonic is 11 per cent better than seven equal steps; on the profile model the same comparison is 0.9 per cent, which is nothing. The two models disagree about the one comparison anybody would want the measure for, and only one of them is free of circularity: the probe-tone profile was measured on listeners raised inside the diatonic tradition, so using it to explain why the diatonic is well chosen assumes the answer. The harmonicity model assumes only that a simple ratio is easier to hold than a complicated one.
Fig. 4 Six scales against the equal division into the same number of degrees. The comparison between a scale and its own equal division does not involve the clustering trick at all, because neither side of it clusters.

The cluster’s advantage is bought by putting two degrees fifty cents apart, and no scale in the comparison above does that. The narrowest step anywhere in those six is a hundred cents, so every boundary in every one of them is at least fifty cents from a degree and the pathological case never arises.

Within that restricted set the model behaves. It says the diatonic is better than seven equal steps by eleven per cent, the maqam scales are level with theirs, and the pentatonics gain nothing — an ordering that is defensible and is the subject of the rung after this one.

A model can be wrong at its optimum and right in its middle, and that is the position this one is in. What it cannot do is what a search asks of it, which is to be trusted at its extreme.

The one thing the winner and the diatonic have in common

There is a feature shared by the top of the ranking and by every scale anybody plays, and it is worth pulling out because it is the only structural agreement the search produces.

Every set near the top of the ranking has a fifth in it — a degree at seven semitones from the tonic — and so does every one of the six traditions this ladder measures. The reason on the search’s side is mechanical: a degree at the fifth puts two boundaries within reach of the security function’s second peak, which is worth more than any other single placement.

The reason on music’s side is different and much older. The fifth is the interval a chain of pure tunings is built from, it is the first thing a spectrum supplies after the octave, and no scale tradition of any size lacks one.

So two entirely unrelated arguments arrive at the same degree, and it is the degree every scale has. That is the strongest agreement between this criterion and the practice, and it is also the least informative, because a criterion that failed to produce a fifth would have been discarded before it was drawn.

What the agreement does bound is how badly the criterion is doing. It is not producing noise: it has a structure, that structure has the fifth in it, and where it departs from the practice it departs for a reason that can be named. A criterion that got the fifth wrong would be a criterion with nothing in it.

How much of this depends on the noise being eleven cents

The whole ranking is generated at one value of σ, and this ladder has been careful about that parameter since its sixth rung.

Every result so far, against the listener's own noiseThree findings drawn against the one parameter all of them assume: how finely the listener resolves a pitch. At 11 cents — a trained listener, and the value every earlier essay used — the octave holds 6 nameable categories, twelve equal ones are named right 91 per cent of the time, and a 20-to-one expectation moves a boundary by 3.6 cents. At 35 cents it is 2 categories, 72 per cent, and 37 cents. The capacity falls roughly as one over sigma and the shift rises as its square, so the three curves separate rather than moving together.9653322every earlier essay10152025303500.20.40.60.81the listener's internal noise, centseach quantity as a share of its largest valuenameable categories9 down to 2 per octavenaming twelve right94% down to 72%boundary moved byexpectation: 1.9 to 37 c
Fig. 5 The earlier sweep: every result produced so far, drawn against the listener’s own internal noise rather than at one value of it. σ is the parameter the whole model turns on and the one nobody had turned.

Here it turns nothing. Every σⱼ in the sum is σ divided by the square root of a fixed weight, so σ is a common factor: the whole error scales linearly with it and every ratio in the ranking is untouched. The diatonic’s rank of 376 is the same at eight cents and at thirty-five.

That is a genuinely robust feature and it is worth saying because it is unusual for this ladder. The sixth rung found that the collection’s headline capacity result inverts over the plausible range of σ; this one finds that the ranking does not move at all. The difference is that a ranking is a comparison and a capacity is a threshold, and a common factor cancels out of the first and decides the second.

What σ does decide is whether the differences matter. At eight cents the whole spread from best to worst set is 3.4 percentage points and at thirty-five it is fifteen, so the same ranking describes a negligible effect for a trained listener and a substantial one for an untrained one.

What the pictures cannot show

The search fixes the tonic and takes seven of twelve, so it searches a grid rather than the continuum. A scale is not obliged to put its degrees at multiples of a hundred cents — the maqam traditions do not — and a continuous search would find a lower optimum than 8.8 per cent by placing boundaries exactly on the security peaks rather than near them.

The count 462 is a binomial coefficient and not a sample, so the ranking is complete and the rank of 376 is exact. What is not exact is the ordering near the middle: sets within a few tenths of a per cent of each other are separated by numbers the model cannot support to that precision, and the diatonic has about forty sets within that band.

The whole search is run under one of the two security models. Under the profile model the diatonic ranks 396 rather than 376 and the winner is a different cluster — so the two models agree that the diatonic loses and disagree about everything else, which is exactly the agreement worth the least.

The security function is also evaluated at the boundary’s exact position, and the nearest-just-ratio lookup it uses is a step function: a boundary at 549 cents gets the fourth’s security and one at 551 gets the tritone’s, which is a factor of three across two cents. Nothing in a listener does that. A smooth function of distance from the nearest simple ratio would be the right shape and would blunt the cluster’s advantage without removing it, and it is not built here.

And nothing here is a claim about listeners. A search over a criterion is a search over a criterion, and the only empirical content in it is whichever security model turns out to describe real confusion rates, which the previous rung’s experiment would settle.

What a search is for when its answer is wrong

This ladder has now run two exhaustive searches and got two different kinds of answer, and the pair is worth setting together because it says what the method is good for.

The diatonic set’s own census enumerates every subset of the twelve and filters it by four properties, and the diatonic is the only survivor — a strong positive result from a search over the same space this one searches.

This search over the same 462 sets puts the diatonic at rank 376 and returns something nobody plays. Same enumeration, different criterion, opposite verdict.

Two searches over one space disagreeing is not a contradiction; it is a measurement of which criterion binds. The census’s four properties are satisfied by exactly one set and identification error is satisfied best by a set no tradition has, so whatever a scale is selected by is much closer to the first than to the second.

That is worth more than either search on its own, and it is only available because both were run over the same enumeration. A criterion that could not be applied to all 462 could not have been compared with the census at all.

Whose scales, and what the ranking is good for

The twelve-semitone grid is a Western construction and the search inherits it, so the whole exercise is about which seven of the twelve rather than about scales in general. Every tradition that does not divide the octave into twelve is outside it by construction.

What the ranking is good for is not choosing a scale. It is a null model: a statement of what a scale would look like if identification error were the only thing that mattered, so that the ways real scales differ from it are the ways other constraints bind.

Read that way the result is informative rather than embarrassing. Real scales avoid clusters, and the cluster is what this criterion wants — so whatever real scales are optimising, it is something that penalises adjacent degrees heavily. That is a constraint the ladder can now name and could not before, and the diatonic set’s own anchor has four properties that all do exactly that.

Where this ladder goes next

Ten rungs. The categories exist; they are wide enough for temperament; one value can belong to two; about seven are nameable; the boundaries barely move; all of it was at one value of one parameter; the parameter was swept; the last assumption was removed and it collapsed to one line; the line’s own assumption was removed and the pattern came back; and now the pattern has been optimised and the optimum is nonsense.

What remains is the comparison the optimisation cannot reach and the previous rung set up. A scale against the equal division of its own size is a comparison in which neither side clusters, so the model’s defect does not bite — and it is the comparison the eighth rung’s null was specifically about. Under one σ they are identical by arithmetic. Under many they are not, and which of them wins is a question about the ear rather than about a search.

Part 10 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionCentsDifference limenEnumerationJust intonationScale degree