Scales and modes

The number every claim here has been quoting

Sigma is the internal noise a listener's pitch judgements carry, it is measured in a laboratory on isolated intervals, and music never presents an isolated interval. Every resolution claim here rests on the laboratory value. Sweep it and the headline finding moves: at eleven cents the ear has about six nameable categories per octave, and at six it has twelve — which is the chromatic scale, and turns 'the ear has fewer boxes than the notation' into 'it has exactly as many'.

Assumes: The listener the model was never run for · How many boxes an octave holds

The sixth rung of this ladder established that every figure in it is a statement about a listener at one value of one parameter, and that the parameter is the whole model. Its last paragraph said what that means and could not do anything about:

Sigma is measured in a laboratory on isolated intervals and applied here to music, and music never presents an isolated interval — so the effective sigma inside a piece is smaller, by an amount that depends on the key, the drone and the immediately preceding notes. Every ladder on this site that quotes a resolution is quoting the laboratory number.

That is not a small dependency and it is not confined to one ladder. This rung does not measure it — measuring it needs listeners and this collection has none. It prices it: sweeps the parameter, and reports which of the collection’s own conclusions move and by how much.

Every resolution claim here, against the number it rests on. How many equal steps of the octave can be named at 95 per cent accuracy, against the internal noise the model gives a listener. The laboratory value this collection quotes everywhere is 11 cents, which gives 6 nameable categories — the "about seven per octave" every claim here has been repeating. The laboratory measures it on isolated intervals and music never presents one, so the effective value inside a piece is smaller by an amount nobody has measured: at 4 cents it is 18, which is the chromatic scale, and the conclusion changes from "the ear has fewer boxes than the notation" to "it has exactly as many". Nothing here measures it. This is what it is worth if it moves.
Fig. 1 How many equal divisions of the octave a listener can reliably name, against the internal noise they carry. The laboratory value quoted here is eleven cents and gives six — the “about seven per octave” every claim here repeats. At six cents it is twelve, which is the chromatic scale.

What sigma is, and where the eleven comes from

Sigma is the standard deviation of a listener’s internal estimate of a pitch interval — the spread of where they think an interval is, around where it actually is. It is not the same thing as a just-noticeable difference: a limen is about telling two things apart, and sigma is about placing one thing on a scale from memory.

The published values come from identification and adjustment experiments in which a listener hears an isolated interval, out of any context, and names it or matches it. Eleven cents is a representative figure and it is the one the fourth rung used to compute how many categories an octave holds.

That experiment is the right one for the question it asks and it removes the two things music always supplies: a key, which tells a listener what the interval is likely to be before they hear it, and a preceding context, which supplies a reference the interval can be measured against rather than remembered against. Both should reduce the effective spread, and neither has ever been quantified in a musical setting.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 2 The model itself, from an earlier essay: identification curves for adjacent interval categories at several values of the internal noise. A steeper crossing is a listener who is more certain; a shallower one is a listener for whom two categories overlap. Every quantity in this essay is a property of the steepness of these curves.

What moves, and how much

The sweep says the collection’s conclusions are not equally sensitive, and the pattern is worth stating carefully.

The category count is very sensitive. It goes as roughly one over sigma: 18 nameable divisions at four cents, 12 at six, 9 at eight, 6 at eleven, 5 at fourteen, 4 at eighteen and 3 at twenty-four. Halving the parameter doubles the answer.

The comparison with twelve-tone equal temperament is where that bites. At eleven cents a listener can name six divisions and the notation asks for twelve, which is the fourth rung’s headline: the notation has more boxes than the ear. At six cents the two are equal. So the collection’s most-quoted result about categorical hearing inverts at a factor of two in a parameter nobody has measured in a musical context.

The accuracy at a fixed division is much less sensitive. Identification accuracy in twelve equal divisions runs from 96.8 per cent at four cents to 80.9 at twenty-four — a range of sixteen points across a factor of six in the parameter. So any claim of the form “a listener gets most of these right” is safe at every value; the claim that moves is the ceiling.

And the comma’s size in sigma units moves proportionally, from 5.4 at four cents to 0.9 at twenty-four. That is the quantity every audibility claim in the comma ladder rests on, and it is the difference between a comma being unmistakable and being at the edge of noticeable.

Every equal division from 5 to 60, and how wrong it is. For each number of equal steps in the octave, how far its best fifth and its best major third fall from the pure ratios, in cents. The divisions people have actually used are the ones with small errors in both, and no other criterion was applied to pick them out.
Fig. 3 What the ceiling is a ceiling on: the equal divisions the essays on scales are about, and how finely each one divides the octave. At the laboratory sigma a listener cannot name the steps of any of them; at half it, they can name twelve; at a quarter, nineteen. Every argument here about whether a division is usable has this parameter under it.

What the sensitivity says about which claims to trust

The useful product of an audit is not a corrected number; it is a sorting of the collection’s claims by how much they depend on the thing that is uncertain. Doing that sorting gives three groups.

Robust. Anything that is a statement about ordering or about a difference between two conditions at the same sigma. That the boundaries between interval categories sit at the midpoints; that a category is wide enough for temperament to live inside; that one acoustic value can belong to two categories; that trained and untrained listeners differ — all of those are comparisons within a fixed value and survive any change to it.

At risk. Anything that compares a category count to twelve. That includes the fourth rung’s whole argument and, transitively, several claims elsewhere on this site about why twelve is a defensible number of divisions. It is at risk from two directions rather than one, which a later section works out: the noise is one asserted number and the accuracy criterion the count is read off is a second, they move the answer by comparable amounts, and neither has an argument behind it.

Untouched. Everything that is about the structure of scales rather than about a listener — the diatonic set’s census, the modes, the comma arithmetic itself. Those are statements about integers and they do not have a listener in them.

That sorting is the thing to carry away, and it is slightly reassuring: the parameter is load-bearing in one place rather than everywhere, and the one place is identified.

Every result so far, against the listener's own noiseThree findings drawn against the one parameter all of them assume: how finely the listener resolves a pitch. At 11 cents — a trained listener, and the value every earlier essay used — the octave holds 6 nameable categories, twelve equal ones are named right 91 per cent of the time, and a 20-to-one expectation moves a boundary by 3.6 cents. At 35 cents it is 2 categories, 72 per cent, and 37 cents. The capacity falls roughly as one over sigma and the shift rises as its square, so the three curves separate rather than moving together.9653322every earlier essay10152025303500.20.40.60.81the listener's internal noise, centseach quantity as a share of its largest valuenameable categories9 down to 2 per octavenaming twelve right94% down to 72%boundary moved byexpectation: 1.9 to 37 c
Fig. 4 The other axis opened earlier: what training does to the same parameter. A trained listener carries a smaller sigma than an untrained one, by roughly the factor this essay is sweeping over — so the difference between two listeners is the same size as the uncertainty in the number, which means a claim about “a listener” is already a claim about which listener.

The one place the sweep gives a genuinely new argument

Most of the audit is defensive — a list of which claims survive. There is one place where sweeping the parameter produces something the fixed value could not, and it is about the equal divisions this site has an anchor for.

Nineteen, thirty-one and fifty-three is an essay about divisions of the octave that close the chain of fifths better than twelve does, and the standing objection to all of them is that their steps are too fine to be heard as distinct categories. That objection is a claim about sigma and it has always been made at the laboratory value.

The sweep prices it. Nineteen divisions needs a sigma of about four cents to be nameable at ninety-five per cent, and thirty-one needs about two and a half — which is well below any published figure for any listener, trained or not. So the objection survives the whole sweep, comfortably, at every value the parameter could plausibly take.

That is a stronger result than the fixed value gave, and it runs the other way from the essay’s headline. The thing that inverts is the comparison with twelve; the comparison with nineteen and above does not, and no reasonable revision of sigma makes a thirty-one-division scale a set of nameable categories.

How many notes an octave can hold, asked twice. The share of trials on which a category is named correctly, against how many equal categories the octave is cut into, for a listener whose internal estimate carries 11 cents of noise — the logistic scale this site's identification figures already use, which is the thirty-cent transition the studies report. At 95 per cent accuracy the ceiling is 6 categories, and seven scores 94.9 per cent — on the line. The other ceiling is resolution: 151 to 356 difference limens fit in an octave depending on register, which is a factor of forty larger. Every system marked below sits between the two, and the marks separate: the number of degrees a mode uses clears the criterion, and the size of the gamut it chooses them from does not.
Fig. 5 The same ceiling drawn against the divisions themselves rather than against sigma. Twelve is reachable at six cents and below; nineteen at four; thirty-one nowhere on this axis. A division of the octave can be a good tuning without its steps being nameable categories — which is what the microtonal traditions have always said and is now priced.

The number is a property of a listener, and the practical question is how far a system can sit from that listener’s categories before the naming stops working.

How far a system can sit from a listener's categories before naming fails. A listener with twelve equal categories and a noise of 11 cents, hearing a system whose degrees are all shifted by the amount on the horizontal axis. The curve is flat and then a cliff: 100 per cent at 20 cents, 97 at 30, 83 at 40, and 54 at 50, which is the boundary itself and where the answer is a coin. What matters about a foreign tuning is therefore not its average distance from the twelve but whether any single degree is near the middle.
Fig. 6 A listener with twelve equal categories and eleven cents of noise, hearing a system whose degrees are all shifted by the amount on the horizontal axis. The curve is flat and then a cliff: 100 per cent correct at 20 cents of shift, 97 at 30, 83 at 40 and 54 at 50.

Twenty cents is free and fifty is a coin toss, which is a much sharper statement than the raw noise figure suggests and is the form the number takes when it is applied to anything. It is also why every historical temperament on this site is inaudible as a naming problem: none of them moves a degree far enough to reach the cliff.

Which computation produced the numbers

The identification model is the first rung’s: a stimulus is a category centre plus Gaussian noise of standard deviation sigma, a response is correct when the noise leaves the estimate inside half a step of the centre, and the accuracy is that probability integrated over stimuli uniform within a category.

The nameable count is the largest number of equal divisions at which that accuracy reaches ninety-five per cent, searched upward. The criterion is a convention, and a section below sweeps it rather than dismissing it — the counts double at ninety per cent rather than growing by a third, and the essay’s headline comparison goes with them.

The comma-in-sigma figure is the syntonic comma over sigma, which is the site’s own constant divided by the sweep variable and is arithmetic rather than a model.

Nothing here is a measurement of anything. It is one model evaluated at seven values of one parameter, and the value of the exercise is entirely in the shape of the dependence rather than in any point on it.

What the audit costs to run, which is nearly nothing

There is a methodological point worth stating because it applies well beyond this ladder.

Sweeping an asserted parameter and reporting which conclusions move is a cheap operation — it is the same computation run seven times — and it produces information that the single value cannot. This collection has recorded, for four consecutive phases, that several of its results rest on asserted parameters and each says so. Saying so is not the same as pricing it, and pricing it turns out to take an afternoon.

The output is a sorting, and a sorting is the useful shape. It does not fix anything; it says which of the collection’s claims a reader should hold loosely, and it identifies the single measurement that would tighten the largest number of them at once.

Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 19 at 63.2 cents, 31 at 38.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.
Fig. 7 The other resolutions this site quotes, which the same audit ought to be run on: the frequency limen, the duration limen, the interval limen and the identification ceiling, each measured in a laboratory and each applied here to music. This essay sweeps one of the four. Three have never been swept and every one of them has the same problem.

The criterion is a second parameter, and it was not swept

The audit above sweeps one number and holds another, and the one it holds turns out to matter as much. The nameable count is the largest division meeting a ninety-five per cent criterion, and the section above dismissed that choice on the grounds that ninety per cent would give counts “roughly a third larger at every sigma” with every conclusion unchanged.

Both halves of that are wrong.

sigma nameable at 95% nameable at 90% ratio
4 18 37 2.06
6 12 25 2.08
8 9 18 2.00
11 (the laboratory value) 6 13 2.17
14 5 10 2.00
18 4 8 2.00
24 3 6 2.00

The counts double rather than growing by a third, at every value, and the factor is 2.0 to 2.2 throughout. And the conclusion that changes is the headline one: at a ninety per cent criterion a listener carrying the laboratory sigma of eleven cents can name thirteen equal divisions, which is more than twelve. So the inversion this essay presents as needing a halving of sigma needs no change to sigma at all — it needs one convention moved five points, and the fourth rung had already printed the thirteen without drawing the consequence.

That does not weaken the audit; it strengthens what the audit was for. The collection’s most-quoted result about categorical hearing turns on two asserted numbers rather than one, they are independent, and moving either one across a plausible range flips it.

The other of the four resolutions the figure below names can be checked in the same breath, and it comes out the other way. Divide the octave by the frequency difference limen and a listener resolves 356 steps at a kilohertz against six they can name — a ratio of fifty-nine. Sweep both: let sigma run from two to thirty-two cents and let the published limen be wrong by anything from a half to four times, and the narrowest the gap ever gets is 2.4, at the most favourable corner of the grid. The two-ceilings result is the robust one in this collection, and it is robust by a wide margin in a way the category count is not.

Where the model stops

A single Gaussian is a strong claim about a listener. Real identification data is not always symmetric about a category centre and categories are not always equal in width; the fifth rung found that boundaries move much less under a shifted prior than a simple signal-detection account would predict, which is already evidence that the model is missing something.

Sigma is treated as one number and it is at least three. There is the noise in hearing the interval, the noise in remembering it, and the noise in reporting it, and a laboratory identification task measures the sum. Music removes some of the second and none of the first, so the “effective sigma in music is smaller” claim is really a claim about one component of three, and how much of the eleven cents that component is has never been apportioned.

Two asserted numbers is not twice the uncertainty of one. They interact: a criterion of ninety per cent at eleven cents and a criterion of ninety-five at six cents give thirteen and twelve, which are nearly the same answer from opposite corners. So the collection’s headline is not a single quantity awaiting a single measurement; it is a surface over two conventions, and the experiment the last section calls for would fix only one of its axes.

Ninety-five per cent is doing more work than it looks. The nameable count is the largest division at which a criterion is met, and a criterion is a cliff: a division that misses it by a percentage point is reported as unusable and one that meets it by a percentage point as usable. The counts quoted are therefore step functions of a smoothly varying quantity, and two sigmas that give the same count can be meaningfully different listeners.

And the direction is asserted rather than derived. The sixth rung said the effective sigma in music should be smaller, on the grounds that context supplies a reference. There is a case for the reverse: music is a divided-attention task with several things happening at once, and a listener attending to a melody, a harmony and a rhythm has less capacity for interval identification than one attending to nothing else. This essay sweeps in both directions for that reason, and the sweep is symmetric about the laboratory value rather than one-sided.

Whose ears, and when

The eleven-cent figure is a modern laboratory value from Western listeners, most of them musically trained, tested on intervals from a twelve-division system they have heard all their lives.

Every part of that is a restriction. The third rung is about listeners naming the same acoustic distance differently under two systems, and the sixth is about the listener the model was never run for. A tradition with a finer division and a lifetime of exposure to it — Turkish makam, Indian classical music, a gamelan tradition — is a case where the effective sigma might be smaller for reasons that have nothing to do with context and everything to do with what the categories are.

Which makes the sweep a claim about a range of possible listeners rather than about one uncertain number, and the range is the more useful object. A collection that quotes “about seven per octave” is quoting a Western trained listener in a booth, and the honest statement is that the number is between three and eighteen depending on who is listening and to what.

What the picture cannot show

Whether context helps at all. The whole essay is built on a conjecture — that a key and a preceding context reduce the effective spread — and the conjecture is untested. It is plausible and it is the sort of thing that has come out the other way before.

It cannot show a listener who is wrong in a structured way. The model’s noise is independent from trial to trial, and a real listener’s errors are correlated — they mishear a particular interval in a particular direction because of what they expect, which is the whole content of the fifth rung’s finding about priors. A structured error of the same size as an unstructured one has quite different consequences for how many categories are usable.

And it cannot show what a category is for. Counting nameable categories assumes the useful thing is naming, and a listener does a great many things with an interval other than name it: hearing it as in or out of tune, hearing it as a tendency, hearing it as a member of a chord. The boundary that barely moves is evidence that identification and discrimination come apart, and the count in the hero figure is a count of one of them.

Where this ladder goes next

Seven rungs. The categories exist; they are wide enough for temperament; one value can belong to two; there are about seven per octave nameable; the boundaries are harder to move than expected; all five are statements at one value of one parameter; and now the parameter is swept and the collection’s own headline is found to turn on it.

The rung after it is the one the audit makes obvious and cannot perform. Every value on this essay’s horizontal axis is a hypothesis about a listener, and the experiment that distinguishes them is one of the simplest in this whole collection: an identification task run twice on the same listeners, once on isolated intervals and once on the same intervals inside an established key. The difference between the two sigmas is the number every ladder here has been assuming and none has measured, it is a single quantity, it needs no equipment this field does not already own, and it would resolve an inversion in this site’s most-quoted result about hearing.

Part 7 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionEqual temperamentIdentificationInternal noiseJust-noticeable differenceLimenMicrotonalityPitch resolution