The listener the model was never run for
Assumes: The boundary that barely moves · How many boxes an octave holds
The boundary that barely moves closed by naming the variable this ladder has held fixed as firmly as any. Every number in it is for a trained listener: the eleven-cent internal noise, the capacity ceiling, the expectation shift, the mis-transcription cliff. The parameter is in every function, and nobody had turned it.
Turning it does not produce a slightly worse listener.
Six categories per octave becomes two. That is the finding, and everything else in this rung is a consequence of it.
Why the dependence is not gentle
The three quantities depend on the same number in three different ways, and that is what makes an untrained listener a different model rather than a degraded one.
Capacity goes as one over sigma. How many boxes an octave holds asks how many equal categories can be named to a stated criterion, and the answer is the octave divided by a fixed multiple of the noise. The multiple is worth getting right, because a rule of thumb is only useful if it returns the answer: solving for the ceiling at each value of sigma and reading the spacing back gives sixteen to twenty times sigma, averaging about eighteen — not four.
| σ | ceiling at 95% | spacing at the ceiling |
|---|---|---|
| 8 | 9 | 16.7 σ |
| 11 | 6 | 18.2 σ |
| 15 | 5 | 16.0 σ |
| 20 | 3 | 20.0 σ |
| 35 | 2 | 17.1 σ |
Eighteen sigma is a very wide spacing and it is worth understanding why the criterion demands one. A boundary between two categories needs the estimate to stay on its own side, which at 95 per cent takes about four sigma; a system of categories named right 95 per cent of the time takes far more, because the criterion is over every position inside every category rather than over the worst pair. So the four-sigma figure is the right one for a boundary and the wrong one by four and a half for a capacity, and the two are easy to confuse because both are cast as an octave divided by a multiple of the noise.
Halve the resolution and halve the count survives, and its resolution does not. The ceilings are integers and they go 9, 6, 5, 3, 3, 2, 2 across the sweep, so the model cannot tell a listener at twenty cents from one at twenty-five, or one at thirty from one at thirty-five: both pairs return the same count. The quantity is coarse at exactly the end of the range this rung is about, which means the six becomes two headline is safe and any finer reading of the untrained end is not.
Naming accuracy falls slowly and then quickly. Twelve equal categories are named correctly 94 per cent of the time at eight cents, 91 at eleven, 80 at twenty-five, 72 at thirty-five. That looks like a gentle decline until it is read the right way round: a listener naming twelve intervals with 72 per cent accuracy is wrong about one note in four, which is not a listener who has twelve categories at all.
The expectation shift goes as the square. A boundary sits where two categories are equally likely, and shifting the prior odds moves it by an amount proportional to the variance rather than to the spread. So a twenty-to-one context moves a trained listener’s boundary by 3.6 cents and an untrained listener’s by 37 — a factor of ten from a factor of three in the underlying noise.
The three together are why this is a rung rather than a caveat. If all three quantities scaled the same way, an untrained listener would be a trained one seen through a blur and every earlier result would survive with its numbers multiplied. They do not scale the same way: the capacity falls linearly, the accuracy falls slowly and then off a cliff, and the shift rises quadratically. A listener at thirty-five cents has fewer categories, is unreliable inside the ones they have, and is far more suggestible about where their edges are — and the third of those is what turns a blurred version of the model into a different one.
What happens to a foreign tuning
The consequence that matters for this collection is about the cross-system rung, and it reverses its reading.
The same distance under two names found that a listener whose categories are twelve equal ones, hearing a system whose degrees are not, mis-transcribes: each foreign degree is assigned to the nearest native box, and the assignment is confident and wrong. A maqam’s neutral third at 350 cents is halfway between two of the listener’s boxes, and the naming is a coin toss.
That result is a trained listener’s, and it needs the listener to have boxes narrow enough for 350 to fall between two of them.
The capacity table gives that a sharper form. At the untrained end the ceiling is two or three categories to the octave, and a maqam’s neutral third, minor third and major third span 70 cents — a third of one category at that resolution. So the three are not merely hard to tell apart; they are interior points of one box, no nearer its edge than each other, and there is no sense in which the listener has assigned any of them wrongly.
An untrained listener does not mis-transcribe a foreign tuning. They absorb it. The mis-transcription result requires a listener with enough resolution to be confidently wrong, and a listener with two categories per octave has nothing to be confidently wrong about — the neutral third, the minor third and the major third are one thing.
That is a genuinely different failure and it points the other way. The trained listener hears a foreign system as out of tune; the untrained listener hears it as music. The complaint about unfamiliar intonation is the complaint of somebody with fine categories, which is very nearly the opposite of how it is usually framed.
There is a corollary about this site’s own habit of quoting cents. A tuning is not a table of cents argued that a table cannot say what a paired tuning sounds like; this rung says a table cannot say what a scale sounds like either, and for a reason that is about the reader rather than the tuning. A five-cent distinction between two published versions of a tradition’s third is a distinction a trained listener could just about hear and an untrained listener could not hear at all, and neither the table nor the essay quoting it says which reader it is addressed to.
The mismatch cliff, and where it moves to
The fourth rung’s other result was a cliff: a listener whose category centres are a fixed offset from the ones being played names them well up to a point and then not at all.
Running the mismatch sweep at the untrained value moves the cliff a long way. A trained listener holds above 90 per cent accuracy up to a 37-cent offset; at twenty-five cents of noise that falls to 19, and at thirty-five it starts below 90 per cent with no offset at all.
That last case is worth stating as a limit on the model rather than as a result about a listener. A criterion that is failed at zero offset cannot measure an offset: the sweep has nothing to sweep. So for a listener at thirty-five cents the mismatch question is not answered badly, it is unaskable, and the same is true of every question in this ladder that is posed as how far can a system be displaced before naming fails. Below about thirty cents of resolution the displacement questions stop having answers, which is a different kind of boundary from the ones the rung has been drawing and is the point at which turning the parameter further stops producing information.
An untrained listener is already in the regime the mismatch result describes, with no mismatch. The failure the fourth rung constructed by moving one system against another arrives for free, and it arrives from the listener rather than from the systems.
That reversal is the one result in this rung that changes an earlier conclusion rather than qualifying it. The fifth rung’s title is a claim about how hard boundaries are to move; at twenty-five cents of noise a twenty-to-one context moves one by nearly nineteen cents, which is a fifth of a semitone and is not a boundary that barely moves.
Which computation produced the numbers
Everything runs on the same one-parameter model the ladder has used since its first rung, with the parameter moved.
A category is a Gaussian around its centre with standard deviation sigma. Identification is the probability that the drawn value’s nearest centre is the true one, evaluated by integrating the density over the interval that belongs to each category. categoryAccuracy is that integral for n equal categories in an octave, and categoryCeilings is the largest n for which it clears a stated criterion — 95 per cent throughout, which is the fourth rung’s own choice and is stated there.
The boundary shift is Bayes with two equal-variance Gaussians: the boundary sits where the likelihood ratio equals the prior odds, and solving gives a displacement of σ² ln(odds) divided by the distance between the centres. That is where the square comes from, and it is arithmetic rather than a fit.
The cross-system naming is the fourth rung’s function unchanged: each degree of a tradition is assigned to the nearest of the listener’s centres and the accuracy is the integral over the assigned box.
The one thing that is not computed is the value of sigma itself. Eleven cents is the logistic scale that reproduces the ~30-cent identification transitions trained listeners are reported to show; twenty-five to thirty-five is used here for an untrained listener because it reproduces the much wider transitions untrained listeners are reported to show. Both are read off published transition widths rather than measured here, and the range is wide because the published range is wide.
The one place the direction is already known
There is a body of evidence this rung does not need to speculate about, and it is worth naming because it is the reason the sigma range used here is not invented.
Identification transitions are measured, and they are measured on both kinds of listener. A trained listener’s identification function swings from a quarter to three quarters over about thirty cents; an untrained listener’s swings over something between two and four times that. The eleven cents this ladder has been using is the logistic scale that reproduces the first, and the twenty-five to thirty-five used here is what reproduces the second. Neither is fitted to anything in this essay.
What is not measured, and what would settle the whole rung, is the capacity at the untrained value: how many equal intervals in an octave an untrained listener can in fact name to a criterion. The model says two or three. Nobody has run it, because an identification study with three categories is not an interesting experiment to a laboratory interested in twelve.
Where the model stops
Training is not one axis. Sigma is a single number standing for everything a musical education does, and a real listener can have fine discrimination and no labels, or confident labels and poor discrimination. The model conflates the two, and the second is much the more common: most listeners can hear that two intervals differ long before they can say which is which.
And a category system does not have to be equal. Every capacity result here cuts the octave into n equal boxes, which is the arrangement that maximises the count for a given noise. A real system’s degrees are unequal — two sizes of every step is the whole reason the diatonic names work — so a listener with three effective categories does not have three equally spaced ones, and where they fall is a question this model cannot ask.
And the listening is not modelled at all. Every number here is for an isolated interval judged out of context. Music supplies context relentlessly — a key, a drone, a preceding phrase — and the shared reference that context provides reduces the effective noise for exactly the reason the pitch-acuity ladder is currently arguing about. So the untrained listener’s sigma in a piece is smaller than their sigma in a laboratory, by an amount nobody has measured.
Whose music, and what the numbers do not license
The arithmetic is about a listener. The word “untrained” is doing a lot of work in it, and it is worth being careful, because there is an old and bad argument in this area.
What the model says is that a listener with coarse pitch resolution has few nameable categories. What it does not say is that such a listener hears less music, enjoys it less, or is missing anything — consonance is half learned, rhythm is not in this model at all, and the entire apparatus of timbre, dynamics, phrasing and form is untouched by the parameter being turned. A listener with two pitch categories per octave has full access to everything this collection’s other eight fields are about.
What it also does not say is that any tradition’s listeners are untrained. Every musical culture trains its listeners, in its own categories, from infancy; the maqam listener whose neutral third is a category has fine resolution there and the twelve-tone listener does not. The model is symmetric and the word “trained” is not, which is why the figures here are labelled in cents rather than in adjectives.
Where the reading does bite is on this collection’s own habit. Nearly every perceptual number quoted on this site comes from an experiment run on conservatoire students, because they are who volunteers for interval-identification studies. That is a sampling fact with a consequence: it means the site’s numbers describe a listener at the fine end of the range and the essays have been writing “a listener” throughout.
What the picture cannot show
Whether the boxes are the same boxes. The model assumes an untrained listener has fewer of the same categories. They might instead have different ones — a “step” and a “leap”, or a “consonant” and a “dissonant” — which is a different partition of the same octave and would be invisible to every figure here.
Whether sigma is even one number for one listener. Resolution varies with register — the limen is finest around a kilohertz and coarser at both ends — so a single sigma is already an average over the octave it is being used to divide. The capacity figures cut the octave into equal boxes and evaluate them against one noise; a proper version would cut it into boxes whose widths follow the limen, and would give a different count.
And nothing here measures anybody. The whole rung is one published model evaluated at a second published parameter range, which is the cheapest kind of result and is worth having only because the parameter had never been moved.
Where this ladder goes next
Six rungs. The categories exist; they are wide enough for temperament to live inside; one acoustic value can belong to two of them; there are about seven per octave that can be named; the boundaries are much harder to move than the fourth rung suggested; and now, all five of those are statements about a listener at one value of one parameter, and the parameter is the whole model.
What the ladder still owes is the thing the context limitation names. Sigma is measured in a laboratory on isolated intervals and applied here to music, and music never presents an isolated interval — so the effective sigma inside a piece is smaller, by an amount that depends on the key, the drone and the immediately preceding notes. Every ladder on this site that quotes a resolution is quoting the laboratory number. What the difference is worth is a listening experiment, and it is the second time in two rungs that this collection has arrived at the same missing measurement from a different direction.
Part 6 of 11
One essay in the series on Categorical-hearing. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Categorical perceptionCategory boundaryCategory widthCross-culturalDifference limenEnculturationIdentificationMicrotonality
- Three answers to how finely a pitch can be heard categorical perception, difference limen, microtonality
- A boundary beside a fifth categorical perception, difference limen
- How much an anchor would have to be worth categorical perception, difference limen
- The chord is still major, and that is why temperament works categorical perception, category width