Scales and modes

The listener the model was never run for

Five earlier essays rest on one number — how finely a listener resolves a pitch — and every one of them used a trained listener's eleven cents. The model's dependence on it is not gentle: the capacity goes as its reciprocal and the expectation shift as its square, so an untrained listener at thirty-five cents has two nameable categories per octave rather than six, and a foreign tuning system is not mis-transcribed by them but absorbed.

Assumes: The boundary that barely moves · How many boxes an octave holds

The boundary that barely moves closed by naming the variable this ladder has held fixed as firmly as any. Every number in it is for a trained listener: the eleven-cent internal noise, the capacity ceiling, the expectation shift, the mis-transcription cliff. The parameter is in every function, and nobody had turned it.

Turning it does not produce a slightly worse listener.

Every result so far, against the listener's own noiseThree findings drawn against the one parameter all of them assume: how finely the listener resolves a pitch. At 11 cents — a trained listener, and the value every earlier essay used — the octave holds 6 nameable categories, twelve equal ones are named right 91 per cent of the time, and a 20-to-one expectation moves a boundary by 3.6 cents. At 35 cents it is 2 categories, 72 per cent, and 37 cents. The capacity falls roughly as one over sigma and the shift rises as its square, so the three curves separate rather than moving together.9653322every earlier essay10152025303500.20.40.60.81the listener's internal noise, centseach quantity as a share of its largest valuenameable categories9 down to 2 per octavenaming twelve right94% down to 72%boundary moved byexpectation: 1.9 to 37 c
Fig. 1 Three earlier findings against the one parameter all of them assume. At eleven cents — the trained listener every earlier essay used — the octave holds six nameable categories, twelve equal ones are named right 91 per cent of the time, and a twenty-to-one expectation moves a boundary by 3.6 cents. At thirty-five cents it is two categories, 72 per cent, and 37 cents. The three curves separate rather than moving together, because the capacity falls as one over sigma and the shift rises as its square.

Six categories per octave becomes two. That is the finding, and everything else in this rung is a consequence of it.

Why the dependence is not gentle

The three quantities depend on the same number in three different ways, and that is what makes an untrained listener a different model rather than a degraded one.

Capacity goes as one over sigma. How many boxes an octave holds asks how many equal categories can be named to a stated criterion, and the answer is the octave divided by a fixed multiple of the noise. The multiple is worth getting right, because a rule of thumb is only useful if it returns the answer: solving for the ceiling at each value of sigma and reading the spacing back gives sixteen to twenty times sigma, averaging about eighteen — not four.

σ ceiling at 95% spacing at the ceiling
8 9 16.7 σ
11 6 18.2 σ
15 5 16.0 σ
20 3 20.0 σ
35 2 17.1 σ

Eighteen sigma is a very wide spacing and it is worth understanding why the criterion demands one. A boundary between two categories needs the estimate to stay on its own side, which at 95 per cent takes about four sigma; a system of categories named right 95 per cent of the time takes far more, because the criterion is over every position inside every category rather than over the worst pair. So the four-sigma figure is the right one for a boundary and the wrong one by four and a half for a capacity, and the two are easy to confuse because both are cast as an octave divided by a multiple of the noise.

Halve the resolution and halve the count survives, and its resolution does not. The ceilings are integers and they go 9, 6, 5, 3, 3, 2, 2 across the sweep, so the model cannot tell a listener at twenty cents from one at twenty-five, or one at thirty from one at thirty-five: both pairs return the same count. The quantity is coarse at exactly the end of the range this rung is about, which means the six becomes two headline is safe and any finer reading of the untrained end is not.

Naming accuracy falls slowly and then quickly. Twelve equal categories are named correctly 94 per cent of the time at eight cents, 91 at eleven, 80 at twenty-five, 72 at thirty-five. That looks like a gentle decline until it is read the right way round: a listener naming twelve intervals with 72 per cent accuracy is wrong about one note in four, which is not a listener who has twelve categories at all.

The expectation shift goes as the square. A boundary sits where two categories are equally likely, and shifting the prior odds moves it by an amount proportional to the variance rather than to the spread. So a twenty-to-one context moves a trained listener’s boundary by 3.6 cents and an untrained listener’s by 37 — a factor of ten from a factor of three in the underlying noise.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 67 cents, against the 100 that separate adjacent categories — so the change of mind happens in 67% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 2 The identification functions of a listener at twenty-five cents rather than eleven: the same five categories, the same centres, and transitions two and a half times as wide. A note at 350 cents is named “minor third” and “major third” nearly equally often, and a note at 250 is not reliably either. The categories have not moved; they have stopped being separate.

The three together are why this is a rung rather than a caveat. If all three quantities scaled the same way, an untrained listener would be a trained one seen through a blur and every earlier result would survive with its numbers multiplied. They do not scale the same way: the capacity falls linearly, the accuracy falls slowly and then off a cliff, and the shift rises quadratically. A listener at thirty-five cents has fewer categories, is unreliable inside the ones they have, and is far more suggestible about where their edges are — and the third of those is what turns a blurred version of the model into a different one.

What happens to a foreign tuning

The consequence that matters for this collection is about the cross-system rung, and it reverses its reading.

The same distance under two names found that a listener whose categories are twelve equal ones, hearing a system whose degrees are not, mis-transcribes: each foreign degree is assigned to the nearest native box, and the assignment is confident and wrong. A maqam’s neutral third at 350 cents is halfway between two of the listener’s boxes, and the naming is a coin toss.

That result is a trained listener’s, and it needs the listener to have boxes narrow enough for 350 to fall between two of them.

Maqam Rast, Arabic theory, named by a listener who has twelve categories. Each degree of this tradition's scale, and how often a listener whose categories are the twelve equal divisions would name it as the same category twice running. The model is the identification one this figure already draws, with the centres unpinned from the stimulus. The mean is 83 per cent and the worst degree is at 350 cents, named correctly 50 per cent of the time — because it sits almost exactly on a boundary, where no amount of quiet in the listener helps.
Fig. 3 The degrees of one maqam named by a listener whose categories are twelve equal ones, at an untrained resolution. With two or three effective categories per octave the neutral third is not between boxes at all — it is comfortably inside one, along with the minor third and the major third, and the listener has no experience of anything being wrong.

The capacity table gives that a sharper form. At the untrained end the ceiling is two or three categories to the octave, and a maqam’s neutral third, minor third and major third span 70 cents — a third of one category at that resolution. So the three are not merely hard to tell apart; they are interior points of one box, no nearer its edge than each other, and there is no sense in which the listener has assigned any of them wrongly.

An untrained listener does not mis-transcribe a foreign tuning. They absorb it. The mis-transcription result requires a listener with enough resolution to be confidently wrong, and a listener with two categories per octave has nothing to be confidently wrong about — the neutral third, the minor third and the major third are one thing.

That is a genuinely different failure and it points the other way. The trained listener hears a foreign system as out of tune; the untrained listener hears it as music. The complaint about unfamiliar intonation is the complaint of somebody with fine categories, which is very nearly the opposite of how it is usually framed.

Maqam Rast, Turkish theory, named by a listener who has twelve categories. Each degree of this tradition's scale, and how often a listener whose categories are the twelve equal divisions would name it as the same category twice running. The model is the identification one this figure already draws, with the centres unpinned from the stimulus. The mean is 100 per cent and the worst degree is at 385 cents, named correctly 100 per cent of the time — which is comfortably inside its own category.
Fig. 4 A second tradition’s degrees named by a trained twelve-category listener, for comparison. Here the mis-transcription is real: the degrees that fall between boxes are named unreliably and the ones that do not are named confidently and in the wrong system’s vocabulary. The two figures are the same computation at two values of one number, and they describe two entirely different listening experiences.

There is a corollary about this site’s own habit of quoting cents. A tuning is not a table of cents argued that a table cannot say what a paired tuning sounds like; this rung says a table cannot say what a scale sounds like either, and for a reason that is about the reader rather than the tuning. A five-cent distinction between two published versions of a tradition’s third is a distinction a trained listener could just about hear and an untrained listener could not hear at all, and neither the table nor the essay quoting it says which reader it is addressed to.

The mismatch cliff, and where it moves to

The fourth rung’s other result was a cliff: a listener whose category centres are a fixed offset from the ones being played names them well up to a point and then not at all.

How many notes an octave can hold, asked twice. The share of trials on which a category is named correctly, against how many equal categories the octave is cut into, for a listener whose internal estimate carries 11 cents of noise — the logistic scale this site's identification figures already use, which is the thirty-cent transition the studies report. At 95 per cent accuracy the ceiling is 6 categories, and seven scores 94.9 per cent — on the line. The other ceiling is resolution: 151 to 356 difference limens fit in an octave depending on register, which is a factor of forty larger. Every system marked below sits between the two, and the marks separate: the number of degrees a mode uses clears the criterion, and the size of the gamut it chooses them from does not.
Fig. 5 The capacity result at the trained value, which is where the six comes from: naming accuracy against how many equal categories the octave is cut into, with the 95-per-cent criterion and the traditions marked. Cutting the octave into more boxes than the resolution supports is not a system with more notes in it; it is a system whose notes cannot be told apart.

Running the mismatch sweep at the untrained value moves the cliff a long way. A trained listener holds above 90 per cent accuracy up to a 37-cent offset; at twenty-five cents of noise that falls to 19, and at thirty-five it starts below 90 per cent with no offset at all.

That last case is worth stating as a limit on the model rather than as a result about a listener. A criterion that is failed at zero offset cannot measure an offset: the sweep has nothing to sweep. So for a listener at thirty-five cents the mismatch question is not answered badly, it is unaskable, and the same is true of every question in this ladder that is posed as how far can a system be displaced before naming fails. Below about thirty cents of resolution the displacement questions stop having answers, which is a different kind of boundary from the ones the rung has been drawing and is the point at which turning the parameter further stops producing information.

An untrained listener is already in the regime the mismatch result describes, with no mismatch. The failure the fourth rung constructed by moving one system against another arrives for free, and it arrives from the listener rather than from the systems.

Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 14.4 cents at ten to one and 18.7 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it.
Fig. 6 The three mechanisms that can move a boundary, drawn at the untrained value of sigma. The expectation curve — the one that goes as the square — is now the largest of the three by a wide margin, where at the trained value it was the smallest. So the earlier headline finding, that boundaries are much harder to move than expected, is a trained listener’s finding and reverses for an untrained one.

That reversal is the one result in this rung that changes an earlier conclusion rather than qualifying it. The fifth rung’s title is a claim about how hard boundaries are to move; at twenty-five cents of noise a twenty-to-one context moves one by nearly nineteen cents, which is a fifth of a semitone and is not a boundary that barely moves.

Which computation produced the numbers

Everything runs on the same one-parameter model the ladder has used since its first rung, with the parameter moved.

A category is a Gaussian around its centre with standard deviation sigma. Identification is the probability that the drawn value’s nearest centre is the true one, evaluated by integrating the density over the interval that belongs to each category. categoryAccuracy is that integral for n equal categories in an octave, and categoryCeilings is the largest n for which it clears a stated criterion — 95 per cent throughout, which is the fourth rung’s own choice and is stated there.

The boundary shift is Bayes with two equal-variance Gaussians: the boundary sits where the likelihood ratio equals the prior odds, and solving gives a displacement of σ² ln(odds) divided by the distance between the centres. That is where the square comes from, and it is arithmetic rather than a fit.

The cross-system naming is the fourth rung’s function unchanged: each degree of a tradition is assigned to the nearest of the listener’s centres and the accuracy is the integral over the assigned box.

The one thing that is not computed is the value of sigma itself. Eleven cents is the logistic scale that reproduces the ~30-cent identification transitions trained listeners are reported to show; twenty-five to thirty-five is used here for an untrained listener because it reproduces the much wider transitions untrained listeners are reported to show. Both are read off published transition widths rather than measured here, and the range is wide because the published range is wide.

The one place the direction is already known

There is a body of evidence this rung does not need to speculate about, and it is worth naming because it is the reason the sigma range used here is not invented.

Identification transitions are measured, and they are measured on both kinds of listener. A trained listener’s identification function swings from a quarter to three quarters over about thirty cents; an untrained listener’s swings over something between two and four times that. The eleven cents this ladder has been using is the logistic scale that reproduces the first, and the twenty-five to thirty-five used here is what reproduces the second. Neither is fitted to anything in this essay.

What is not measured, and what would settle the whole rung, is the capacity at the untrained value: how many equal intervals in an octave an untrained listener can in fact name to a criterion. The model says two or three. Nobody has run it, because an identification study with three categories is not an interesting experiment to a laboratory interested in twelve.

Where the model stops

Training is not one axis. Sigma is a single number standing for everything a musical education does, and a real listener can have fine discrimination and no labels, or confident labels and poor discrimination. The model conflates the two, and the second is much the more common: most listeners can hear that two intervals differ long before they can say which is which.

And a category system does not have to be equal. Every capacity result here cuts the octave into n equal boxes, which is the arrangement that maximises the count for a given noise. A real system’s degrees are unequal — two sizes of every step is the whole reason the diatonic names work — so a listener with three effective categories does not have three equally spaced ones, and where they fall is a question this model cannot ask.

And the listening is not modelled at all. Every number here is for an isolated interval judged out of context. Music supplies context relentlessly — a key, a drone, a preceding phrase — and the shared reference that context provides reduces the effective noise for exactly the reason the pitch-acuity ladder is currently arguing about. So the untrained listener’s sigma in a piece is smaller than their sigma in a laboratory, by an amount nobody has measured.

Whose music, and what the numbers do not license

The arithmetic is about a listener. The word “untrained” is doing a lot of work in it, and it is worth being careful, because there is an old and bad argument in this area.

What the model says is that a listener with coarse pitch resolution has few nameable categories. What it does not say is that such a listener hears less music, enjoys it less, or is missing anything — consonance is half learned, rhythm is not in this model at all, and the entire apparatus of timbre, dynamics, phrasing and form is untouched by the parameter being turned. A listener with two pitch categories per octave has full access to everything this collection’s other eight fields are about.

What it also does not say is that any tradition’s listeners are untrained. Every musical culture trains its listeners, in its own categories, from infancy; the maqam listener whose neutral third is a category has fine resolution there and the twelve-tone listener does not. The model is symmetric and the word “trained” is not, which is why the figures here are labelled in cents rather than in adjectives.

Where the reading does bite is on this collection’s own habit. Nearly every perceptual number quoted on this site comes from an experiment run on conservatoire students, because they are who volunteers for interval-identification studies. That is a sampling fact with a consequence: it means the site’s numbers describe a listener at the fine end of the range and the essays have been writing “a listener” throughout.

What the picture cannot show

Whether the boxes are the same boxes. The model assumes an untrained listener has fewer of the same categories. They might instead have different ones — a “step” and a “leap”, or a “consonant” and a “dissonant” — which is a different partition of the same octave and would be invisible to every figure here.

Whether sigma is even one number for one listener. Resolution varies with register — the limen is finest around a kilohertz and coarser at both ends — so a single sigma is already an average over the octave it is being used to divide. The capacity figures cut the octave into equal boxes and evaluate them against one noise; a proper version would cut it into boxes whose widths follow the limen, and would give a different count.

And nothing here measures anybody. The whole rung is one published model evaluated at a second published parameter range, which is the cheapest kind of result and is worth having only because the parameter had never been moved.

Where this ladder goes next

Six rungs. The categories exist; they are wide enough for temperament to live inside; one acoustic value can belong to two of them; there are about seven per octave that can be named; the boundaries are much harder to move than the fourth rung suggested; and now, all five of those are statements about a listener at one value of one parameter, and the parameter is the whole model.

What the ladder still owes is the thing the context limitation names. Sigma is measured in a laboratory on isolated intervals and applied here to music, and music never presents an isolated interval — so the effective sigma inside a piece is smaller, by an amount that depends on the key, the drone and the immediately preceding notes. Every ladder on this site that quotes a resolution is quoting the laboratory number. What the difference is worth is a listening experiment, and it is the second time in two rungs that this collection has arrived at the same missing measurement from a different direction.

Part 6 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionCategory boundaryCategory widthCross-culturalDifference limenEnculturationIdentificationMicrotonality