Perception and the listener

The boundary that barely moves

Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.

Assumes: How many boxes an octave holds · The ear sorts into boxes, and the boxes are the theory

The fourth rung of this ladder counted the categories an octave can hold and ended by conceding what its own model does not contain: the centres are fixed and the noise is fixed, and real category boundaries are reported to shift with context, with the preceding interval, and with the tuning system a listener has just been hearing. It said the site had the machinery for both halves of that and had never put them together.

Putting them together produces a smaller result than the concession implied, and the smallness is the finding.

Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 2.8 cents at ten to one and 3.6 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it.
Fig. 1 Three mechanisms that could move a category boundary, each drawn against how strong the context is. They differ by a factor of four in size and, more usefully, two of them move the boundary toward the context and one moves it away. Nothing in the picture needs a calibration to be told apart from the others.

Expectation is worth less than a comma

The first mechanism is the one most people mean by “context shifts the boundary”: the listener expects one category more than the other, and a stimulus near the line gets pushed toward the expected answer.

That is a decision problem with a known answer. Given an estimate corrupted by Gaussian noise of standard deviation σ, and two categories whose centres are Δ apart, the boundary that minimises errors sits not at the midpoint but a distance σ² ln(odds) / Δ from it, where the odds are the prior ratio between the two categories.

Every quantity in that expression is already on this site. σ is 11 cents — the logistic scale the identification figures use, chosen so the drawn transition matches the thirty cents that studies of trained listeners report. Δ is 100 cents for adjacent semitone categories.

At two-to-one odds the boundary moves 0.84 cents. At ten to one, 2.8. At twenty to one, 3.6.

A comma is 21.5 cents. The difference between a just and an equal-tempered major third is 13.7, and a third that far out is still a major third. An expectation strong enough to be worth calling a context moves the boundary by a fraction of the smallest interval this collection ever argues about.

That is a bound rather than a measurement, and it is a bound of a particular kind: it is what an optimal listener would do. A real listener could do more, but doing more would be worse, and the whole reason categorical perception is an efficient strategy is that the categories track the world rather than the listener’s hopes.

It is also a bound that depends on σ squared, which makes the smallness much more fragile than a single number suggests. σ is 11 cents for the trained listeners the identification studies use; sweeping it:

σ 2:1 10:1 20:1
8 0.4 1.5 1.9
11 0.8 2.8 3.6
15 1.6 5.2 6.7
20 2.8 9.2 12.0
25 4.3 14.4 18.7

At σ = 25 a twenty-to-one expectation is worth eighteen and a half cents, which is most of a comma, and the shift passes a comma outright at σ = 26.8. So the finding is not expectation is worth less than a comma; it is expectation is worth less than a comma for a listener whose categories are sharp.

That qualification runs the wrong way for the claim, because the listeners with the widest transitions are exactly the ones with the weakest categories — Burns and Ward found the categorical effect strongly in trained musicians and weakly or not at all in the untrained, which is to say the untrained have the larger σ. The mechanism this rung dismisses as negligible is largest for the listeners whose boundaries are least fixed to begin with, and that is a coherent picture rather than a contradiction: a listener with vague categories is a listener whose categories are more nearly what they expect.

The other quantity in the expression bites the same way. The shift goes as one over Δ, so a listener with twenty-four categories rather than twelve has boundaries twice as movable: at σ = 11 and twenty-to-one, 7.2 cents rather than 3.6. A quarter-tone tradition’s listener therefore has both finer categories and more mobile boundaries, from one formula, and it is the second half that the fixed-centre figures cannot draw.

Re-learning the centres is a different thing entirely

The second mechanism is not a shift of the boundary at all. It is a shift of the centres, with the boundary following.

A listener who has spent twenty minutes inside a tuning system whose major third is 386 cents rather than 400 may come to treat 386 as the centre of the category. If the neighbouring centres move too, the boundary between them moves by the mean of their movements. If only one moves, the boundary moves by half of it.

That mechanism is worth up to half the size of the re-tuning, which for meantone against equal temperament is about 7 cents and for a system whose degrees sit 50 cents away is 25. It is an order of magnitude larger than the expectation effect at any plausible strength.

And it is not a fact about boundaries. It is learning, and learning has a time course, a decay and a dependence on how much was heard, none of which the model above contains.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 2 The picture the earliest essay is built on and the one both mechanisms are perturbations of. Each category’s probability against interval size, with the transitions swinging from three-quarters to one-quarter over 24 cents. A shift of three cents is an eighth of one of those transitions; a shift of twenty-five moves a boundary a quarter of the way to the next centre.

The one that goes the other way

The third mechanism is the classic selective-adaptation effect, and it is the reason the experiment is worth specifying: it moves the boundary away from the adapting value.

Repeated exposure to a stimulus fatigues whatever is tuned to it, so the response to a subsequent stimulus is biased away — the aftereffect that makes a still image drift after watching a waterfall, and that has been reported in speech-sound identification for half a century. Its size here is unknown, and the figure draws it as a shape with a free constant rather than as a number, because a curve whose height can be set to anything is not evidence.

What the free constant does not touch is the sign.

Two of the three mechanisms move the boundary toward the context and one moves it away. So the experiment that separates them needs no absolute calibration at all: play an adapting sequence built around a flat major third, then run an identification series, and ask only which way the boundary went. If it went flat, the listener is either expecting or re-learning, and the size distinguishes those two. If it went sharp, none of the collection’s arithmetic applies and the mechanism is adaptation.

The sign test is worth more than the size test for a second reason the σ sweep supplies. The size test distinguishes expectation from re-learning by a factor of about seven at σ = 11 — three and a half cents against twenty-five — and by a factor of only 1.3 at σ = 25, where expectation alone reaches eighteen. So a size measurement separates the two mechanisms for a trained listener and not for an untrained one, while the sign separates adaptation from both at any σ whatever. An experiment that measures only the direction is therefore the one that works on everybody, and it is also the cheaper of the two to run.

That is a cheap experiment, it is not one this site can run, and the reason it is worth stating precisely is that the three predictions were an order of magnitude apart before anybody listened to anything.

The degree nobody can name

There is one place where the fixed-centre model, unperturbed, says something sharp about real music, and it comes out of running the identification model across systems rather than within one.

A listener whose categories are the twelve equal divisions hears a scale whose degrees are somewhere else. Each degree gets assigned to whichever of the listener’s twelve centres it is nearest, and the noise decides how reliably.

Maqam Rast, Arabic theory, named by a listener who has twelve categories. Each degree of this tradition's scale, and how often a listener whose categories are the twelve equal divisions would name it as the same category twice running. The model is the identification one this figure already draws, with the centres unpinned from the stimulus. The mean is 86 per cent and the worst degree is at 350 cents, named correctly 50 per cent of the time — because it sits almost exactly on a boundary, where no amount of quiet in the listener helps.
Fig. 3 Maqam Rast in the Arabic theory fixed at Cairo in 1932, heard by a listener with twelve categories. Five of its seven degrees are named perfectly. Two are not named at all: the neutral third at 350 cents and the neutral seventh at 1050 sit exactly halfway between two of the listener’s centres, so the answer is a coin toss and no amount of quiet in the listener improves it.

Fifty per cent, exactly, and it is not noise. A degree at the midpoint of two categories is unnameable in principle by a listener who has only those two: the model’s error is not that it is uncertain but that the question has no answer in its vocabulary. Every trial is a fresh coin.

That is also the one case where the shift arithmetic above is not negligible, and it is negligible for the wrong reason. A degree sitting exactly on a boundary is the most movable stimulus there is: a two-to-one expectation is enough to break the tie completely, because breaking a tie needs only a shift in a direction rather than a shift of a size. So a listener who has heard enough of a maqam to expect a third rather than a fourth will name the neutral third consistently — not because they have learned the category, but because a coin with a thumb on it stops being a coin. The fifty per cent is a figure for a listener with no expectations at all, and it is the one place in this rung where three and a half cents is the whole of what is needed.

The same run over the other traditions this collection carries gives a very different picture.

Maqam Rast, Turkish theory, named by a listener who has twelve categories. Each degree of this tradition's scale, and how often a listener whose categories are the twelve equal divisions would name it as the same category twice running. The model is the identification one this figure already draws, with the centres unpinned from the stimulus. The mean is 100 per cent and the worst degree is at 385 cents, named correctly 100 per cent of the time — which is comfortably inside its own category.
Fig. 4 Maqam Rast in the Turkish theory, whose degrees are named in Holdrian commas and which puts the third at 385 cents rather than 350. Every degree is named at ninety-nine per cent or better. The two theories of what is nominally the same maqam differ by 35 cents on one degree, and that 35 cents is the difference between a scale a twelve-category listener can transcribe and one it cannot.

That is a result about two theories rather than about two musics, and it is worth being careful with. Two national theories of one maqam disagree by a third of a semitone — the scales rung established that. A tuning is not a table of cents either, so both theories are idealisations of what is played. What is added here is a consequence: the Arabic theory’s third is at the one place in the octave where a twelve-category listener has nothing to say, and the Turkish theory’s is not. Whichever of them is closer to any given performance, the transcription problem is not symmetric.

Slendro, one Javanese gamelan, named by a listener who has twelve categories. Each degree of this tradition's scale, and how often a listener whose categories are the twelve equal divisions would name it as the same category twice running. The model is the identification one this figure already draws, with the centres unpinned from the stimulus. The mean is 92 per cent and the worst degree is at 955 cents, named correctly 68 per cent of the time — because it sits almost exactly on a boundary, where no amount of quiet in the listener helps.
Fig. 5 Slendro, measured on one Javanese gamelan, run through the same model. Mean 92 per cent, with the fifth degree at 955 cents the weakest at 68. The pattern is different again: slendro’s degrees are not near the twelve and not on the boundaries either, so a twelve-category listener names them consistently and names them wrong.

That is also the case the same distance under two names made from inside the twelve: one acoustic value can belong to two categories, and which one it is comes from context rather than from the sound. Here the value belongs to neither, and no context inside a twelve-category system can supply an answer.

Consistently and wrong is a different failure from inconsistently. A listener who calls 955 cents a major sixth every time has a stable, communicable, incorrect transcription; a listener who calls 350 cents a minor third half the time and a major third the other half has no transcription at all. The first produces the recorded history of Western notations of gamelan; the second produces the recorded history of Western notations of maqam, which is a history of special symbols.

How much mismatch a listener can absorb

The general form of that calculation is one variable: shift the whole heard system away from the listener’s grid and watch the naming fall apart.

How far a system can sit from a listener's categories before naming fails. A listener with twelve equal categories and a noise of 11 cents, hearing a system whose degrees are all shifted by the amount on the horizontal axis. The curve is flat and then a cliff: 100 per cent at 20 cents, 97 at 30, 83 at 40, and 54 at 50, which is the boundary itself and where the answer is a coin. What matters about a foreign tuning is therefore not its average distance from the twelve but whether any single degree is near the middle.
Fig. 6 A listener with twelve categories hearing a system shifted bodily away from them. The accuracy holds essentially perfect out to twenty cents, is still 97 per cent at thirty, falls to 83 at forty and to 54 at the boundary itself. The shape is the reason a foreign tuning system is either transcribable or not, with very little in between.
Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents. Every one of them is coarser than discrimination and finer than melodic judgement.
Fig. 7 The three resolutions the arithmetic sits between, with the twelve and the twenty-four drawn as step sizes. The relevant one is the middle band: a category is about a hundred cents wide and the listener’s noise is eleven, so there is a great deal of room before a mismatch matters. The melodic band is wider still, which is a separate argument arriving at the same conclusion — the categories are generous.

The accuracy holds essentially perfect out to a shift of about 20 cents, is still 97 at 30, falls to 83 at 40 and to 54 at 50. The curve is flat and then a cliff, which is what a boundary is, and it means that the interesting question about a foreign tuning system is never the average distance from the twelve. It is whether any single degree is near 50.

There is a second reading of the same cliff, and it is about notation rather than about listeners. The stave counts letters rather than measuring pitch, and an accidental is the repair for a scale with five values missing; what the curve above says is that the repair works perfectly until a degree lands mid-category and then fails completely. A notation with twelve names is a listener with twelve categories, and it has the same cliff in the same place.

That is a usable test and it is a short one. Of the four traditions entered on this site, exactly one has a degree within 5 cents of a boundary, and it is the one with the special notation.

Which computation produced the numbers

The identification model is the one this ladder has used since its first rung, read as a decision: a stimulus arrives, the listener’s internal estimate is that value plus Gaussian noise of standard deviation σ, and the response is whichever category centre the estimate is nearest. σ is 11 cents and comes from outside — it is the logistic scale that reproduces the 25-to-75 per cent transition width the identification literature reports for trained listeners, and it is the only number here not computed.

The cross-system run changes one thing and one thing only: the centres are the listener’s and the stimulus is the heard system’s, so a degree need not be at any centre. Correct means naming it as the same category it is objectively nearest to, twice running — which is why a degree exactly on a boundary scores one half rather than zero.

The prior-shift expression is the standard Bayesian decision boundary for two equal-variance Gaussians, with no free parameters once σ and Δ are given.

The prototype model is arithmetic: the boundary between two categories is the midpoint of their centres, so it moves by the mean of their shifts.

The adaptation model is a shape with a gain in it and is drawn as one. Its constant is not known to this collection and is not guessed at; what is used is the sign, which is not in doubt.

Whose music, and when

The transcription claim is about a specific and datable practice: European and American ethnomusicologists notating Arabic and Javanese music from the late nineteenth century onward, using a staff and twelve categories. The special symbols that appeared for the neutral intervals — the half-flat and half-sharp of the Arabic theoretical literature, standardised alongside the 1932 Cairo convention itself — arrived exactly where the arithmetic above says notation had to break.

The gamelan case went the other way and the arithmetic says why it could. Nineteenth-century Western transcriptions of slendro assigned its degrees to Western note names and were internally consistent, which is what made them look adequate and what made the tuning’s actual character invisible for decades.

Neither of those is a claim about which system is correct. Both theories of rast are theories, held by living traditions, and the degrees a gamelan actually plays are the ones its bars are filed to. What the model describes is a listener with twelve categories, which is a listener rather than a truth.

What the picture cannot show

The noise is one number and listeners are not one listener. σ = 11 cents is for trained listeners in a laboratory. An untrained listener’s is larger, and a larger σ makes the prior shift bigger — it goes as σ² — so at 25 cents of noise the twenty-to-one expectation is worth 18 cents rather than 3.6. The claim that expectation is negligible is a claim about trained ears.

The categories are assumed to be equally spaced and equally variable. Real interval categories are not: the fifth and the octave are narrower categories than the thirds, which is what the beat-counting limen predicts and what the model here does not contain.

Nothing here is a boundary that has been measured to move. The three curves are three predictions, and the experiment that would choose between them has not been run by anybody this collection can cite for these stimuli. What has been established is that the predictions are far enough apart to be told apart.

And the cross-system figures assume the listener assigns to the nearest centre and nothing else. A real listener hearing a maqam has a melodic context, a drone, and a great deal of information the model discards — which is exactly the machinery a tradition with a large gamut supplies, and it is why the neutral third is not, in practice, a coin toss for anybody who knows the music.

Where this ladder goes next

Five rungs. The categories exist; they are wide enough for temperament to live inside; one acoustic value can belong to two of them; there are about seven of them per octave that can be named; and now, the boundaries between them are much harder to move than the concession at the end of the fourth rung suggested.

What is left is the variable this ladder has held fixed as firmly as any: the listener’s training. Every number above is for a trained listener, σ is a trained listener’s, and the capacity result of the previous rung was a trained listener’s too. The model’s dependence on σ is not gentle — the prior shift goes as its square and the naming accuracy falls off a cliff at a mismatch that scales with it — so an untrained listener is not a slightly worse version of this model but a qualitatively different one, in which expectation matters and a foreign tuning system is absorbed rather than mis-transcribed. That is a computation, the parameter is already in every function here, and nobody has turned it.

Part 5 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AdaptationCategorical perceptionCategory boundaryCentsIdentificationJust-noticeable differenceMaqamNeutral third