How much an anchor would have to be worth
Assumes: An interval is two errors · The ear sorts into boxes, and the boxes are the theory
An interval is two errors treated a melodic interval as the difference of two independent pitch estimates and added their limens in quadrature. It also recorded, in its own words, where it could be attacked: if a listener holds a reference — a tonic, a drone, the memory of the first note as an absolute — then the two errors share a component and partly cancel.
That rung declined to draw the correlated model, on the grounds that a curve with a dial on it that can be turned to reach any answer is not evidence. The objection is correct and it is not the end of the matter, because the dial is not free.
A factor of two requires a correlation of three quarters. That is a specific, testable, and rather large number, and stating it is the whole of what this rung does that the previous one could not.
Why it is a decomposition and not a fit
The useful move is to stop writing the model as a correlation at all.
Write each note’s estimate as the truth plus two errors: a common one the two notes of an interval share, because both are referred to the same key, the same drone, the same recently-heard reference; and an independent one they do not. Then a single note’s variance is σc² + σi², and the difference of two notes has variance 2σi² — because the common part is in both estimates with the same sign and subtracts out completely.
ρ = σc²/(σc² + σi²), which is not a parameter to be fitted but a ratio of two quantities that can each be measured on the same listener in the same afternoon. Measure a single-note limen. Measure an interval limen. The second divided by √2 is σi, the first with σi taken out of it in quadrature is σc, and there is nothing left over.
This is the same manoeuvre the loudness ladder made with its two models and the same one the missing-fundamental ladder made when it closed: a parameter that looks free stops being free the moment two independent measurements constrain it, and finding the second measurement is usually the whole of the work.
That is the difference between a dial and a rung. The previous essay was right that a free ρ proves nothing; what it did not have was the observation that ρ is over-determined by two measurements that both already exist as standard psychophysical tasks.
That last observation is a real constraint on how large ρ can be. The Fourier bound cannot be shared between two notes: it is set by each note’s own duration and it is in every estimate independently. So at short note lengths, where the bound dominates, ρ has a ceiling that has nothing to do with how good the listener’s anchor is.
“Well below” is not a number, and the ceiling is computable exactly: σi is at least the bound, so ρ is at most 1 − (bound/limen)². At A440:
| note length | limen | of which the bound | ρ ceiling | best improvement |
|---|---|---|---|---|
| 0.25 s | 7.85 | 7.85 | 0.00 | none |
| 0.4 s | 4.91 | 4.91 | 0.00 | none |
| 0.6 s | 4.04 | 3.28 | 0.34 | 1.23× |
| 1 s | 4.04 | 1.97 | 0.76 | 2.06× |
| 2 s | 4.04 | 0.98 | 0.94 | 4.11× |
At a quarter of a second the ceiling is not “well below 0.75” — it is exactly zero. Below the crossover at 0.486 seconds the bound is the limen, there is no listener-limited component for an anchor to share, and no reference of any kind can improve an interval judgement at all. The hero figure at the top of this essay is drawn at a quarter of a second, which is a note length at which its own answer is unreachable.
Putting that on the tempo axis is where it bites. A crotchet at 60 gives a ceiling of 0.76 and could just support the factor of two; at 90 it is 0.47 and at 120 it is 0.06 — a best-possible improvement of three per cent. At A440 the mechanism this rung proposes to measure can only operate on notes at or below about sixty to the minute, and the threshold moves with pitch: at 1760 hertz the bound falls fast enough that a crotchet at 217 would still do.
So the prediction sharpens from the anchor should help long notes and not short ones into something with an edge on it: it is high, slow notes where an anchor can be worth anything, and the experiment’s two conditions have to be run there or they will find nothing whatever the listener’s reference is doing.
That edge is sharper than the physics warrants and the reason is a modelling convention. The single-note limen is the larger of the two bounds rather than their quadrature sum, so the ceiling snaps to zero the moment the bound overtakes the ear. Combining them in quadrature instead softens it without changing the conclusion: 0.21 at a quarter second, 0.51 at a half, 0.81 at one — a best improvement of 1.12×, 1.43× and 2.29×. Either way the factor of two needs about a second.
The experiment, and what it would have to find
The run has two conditions and one variable.
The ceiling also decides where the run has to be done, which the two conditions below do not say. A stimulus of quarter-second notes at A440 is the natural laboratory choice — it is short enough to be a melodic interval and long enough to have a pitch — and it is the one duration at which the model guarantees a null result. A null at a quarter of a second would be evidence about the Fourier bound and not about the listener, and it would be indistinguishable from the substantive null the section below describes. The two conditions have to be run at a second or more, on notes high enough that a second is musically ordinary, or the experiment cannot tell its own two failure modes apart.
Condition one: no key. Two successive tones, an interval apart, in isolation — no preceding context, no drone, ideally at pitches that do not sit in any familiar scale. Measure the smallest change in the interval a listener can detect.
Condition two: an established key. The same two tones preceded by a cadence or a scale that puts them at recognisable scale degrees, so that a categorical anchor is available. Measure the same thing.
The ratio of the two limens is √(1 − ρ) and gives ρ directly. And the size the effect has to reach to matter is the hero figure’s marks: under about 1.2, it is a curiosity; at 1.5 it changes which melodic intervals are protected; at 2 it means every number in the previous rung is out by a factor of two.
The simultaneity is the control. If a listener’s melodic interval limen improves in a key and their harmonic one does not, the improvement is in the pitch-estimation stage and this model is the right one. If both improve, something more general is going on — attention, arousal, familiarity — and the correlation reading is wrong.
It is worth saying what a null result would mean, because it is not nothing. If the two conditions give the same limen, then a listener inside a key is not referring successive notes to a shared standard when judging how far apart they are — which is a substantive claim about a mechanism that a great deal of tonal-perception writing takes for granted. The independent model would then be right rather than merely convenient, and every number in the previous rung would stand as computed rather than as an upper bound.
What a categorical anchor would actually do
There is a complication, and it is the one that makes the prediction interesting rather than merely quantitative.
The ear sorts into boxes, and the boxes are the theory. A scale degree is not a reference point; it is a category with a width, and referring a note to a category is a quantisation rather than a measurement. That has two consequences pulling opposite ways.
For an interval inside the category system — a fifth, a third, a step of the scale in use — a category is a strong shared reference and ρ should be large. For an interval outside it — a neutral third, a quarter tone, anything the listener’s own system has no box for — the same machinery is a liability: the note is assigned to the nearest box and the assignment discards the very deviation the task is asking about.
So the prediction is not that a key improves interval discrimination. It is that a key improves it for in-category intervals and degrades it for out-of-category ones, and that the improvement should be largest for notes near category centres and smallest for notes near boundaries — which the categorical ladder has independently measured as barely movable.
That is a shape rather than a number, and a shape is much harder to get by accident.
Which computation produced the numbers
The single-note limen is durationLimen unchanged from the third rung: the larger of the published steady-tone discrimination limen at that frequency and the Fourier bound of 1/2T, converted to cents. At A440 and a quarter of a second it is 7.85 cents, and the bound rather than the ear is what is limiting there.
The interval limen is √(a² + b² − 2ρab) with a and b the two notes’ limens. For an interval small enough that a ≈ b this reduces to a√(2(1 − ρ)), which is why the whole curve is one shape scaled: the ratio to the independent case is √(1 − ρ) with no dependence on frequency, note length or interval size.
The correlation required for a factor k is therefore 1 − 1/k², which is 0.31 for 1.2, 0.56 for 1.5 and 0.75 for 2. Those are the marks on the hero figure and they are exact.
The split into shared and independent components is the same algebra read the other way, with the single-note limen as the total and ρ as the share of its variance that is common.
Where the model stops
The common error is assumed to be common in the same amount to both notes. Two notes of an interval are at different pitches, and a reference held as a scale degree is a different distance from each of them; there is no reason the shared component should be identical rather than merely correlated. Allowing them to differ replaces one parameter with two and the experiment above measures only the combination.
A correlation of one is nonsense and the model does not know it. At ρ = 1 the interval limen is zero, which is the model announcing that it has been pushed past where it means anything. The decomposition version does not have that defect — σi cannot be zero because it contains the Fourier bound — and that is a second reason to prefer it.
And nothing here contains time. A shared reference has to be held, and pitch memory decays on a timescale of seconds. Two notes a semiquaver apart share a reference almost perfectly; two notes a phrase apart share a degraded one. So ρ is a function of the gap between the notes, and every figure here is drawn at a single gap of zero.
Whose music, and what turns on it
The claim is about listeners, so it applies wherever there is a scale. What turns on it is a claim this collection has made repeatedly about repertoires.
The published band for melodic interval intonation — the amount by which a sung or played interval can be off before a trained listener objects — is twenty-five to fifty cents, and the fourth rung derived a number inside it from two bounds computed for other reasons. If ρ is 0.75, the derived number is half of what was drawn and no longer lands in the band, which would mean either that the band is measuring something else or that the anchor is not there.
That matters most for the traditions this site has argued about at the edges. A degree is where it goes next is about a system in which a scale degree is defined by its behaviour rather than by its pitch, and a listener inside such a system has a very strong categorical reference indeed; the third the model has no opinion about is about intervals for which a twelve-category listener has no box. The prediction above says those two listeners should differ in opposite directions on the same stimulus, which is the cleanest cross-cultural test this collection has been able to state.
What the picture cannot show
Whether the anchor is a category or a pitch. A drone gives an absolute reference; a key gives a categorical one; the memory of the first note gives a third thing. All three produce a correlated error and this model cannot tell them apart, and the three make different predictions about what happens when the key is established and then removed.
Nor whether the two conditions differ only in the anchor. Establishing a key means playing more notes, and more notes mean more time, more attention and a possible expectation effect on the target itself. A control with a non-tonal context of the same length is the obvious remedy and it is not obvious that a non-tonal context of the same length exists for a listener who has one.
The ceiling is as much a modelling artefact as a finding. It rests on the single-note limen being the larger of two bounds and on the Fourier bound being wholly independent between two notes. The first is a convention this ladder adopted for a different reason; the second is an argument rather than a measurement, and a listener who integrated a note’s frequency over a window longer than the note would violate it.
And nothing here is measured. Every number above is a model evaluated at a parameter the model does not determine. What has been added to the previous rung is not a measurement but a specification: the size the effect must have, the two limens that would measure it, and the control that would say whether it is the right explanation.
The size that is already known
One piece of the answer is not missing, and it is worth putting beside the rest.
A drone is the strongest possible shared reference: an actual sounding pitch, continuously available, requiring no memory at all. And the intonation this collection has measured against a drone is very fine indeed — a tuner working against a held note resolves a fraction of a cent, because the judgement is not two estimates at all but a beat count.
That is the ceiling case and it says the mechanism is real: a shared reference does collapse a two-estimate problem into a one-estimate problem, when the reference is physically present. The question this rung asks is how much of that survives when the reference is remembered rather than sounding, and the answer is somewhere between the drone’s fraction of a cent and the independent model’s nine.
Where this ladder goes next
Five rungs. How finely two pitches can be told apart; that the octave is not where it should be; that the whole family is a function of note length; that an interval is two errors and is heard less finely than either note; and now, that the independence in that model is a measurable ratio rather than an assumption, with a stated size it would have to reach.
The rung after it is the one the decay limitation names. Every figure in this ladder is about two notes with nothing between them, and a melody is notes with other notes between them — so the shared reference is not held across a gap but across material, and the material is itself pitched. Whether an intervening note strengthens the anchor by confirming the key or weakens it by overwriting the memory is a question with two published literatures and one arithmetic, and this collection has the arithmetic.
Part 5 of 12
One essay in the series on Pitch-acuity. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Categorical perceptionCentsDifference limenIntonationJust-noticeable differenceNull modelPitch memoryScale degree
- A boundary beside a fifth categorical perception, cents, difference limen, scale degree
- An interval is two posteriors subtracted cents, difference limen, pitch memory, scale degree
- How many boxes an octave holds categorical perception, difference limen, just-noticeable difference, scale degree
- The best seven of the twelve categorical perception, cents, difference limen, scale degree
- A note that is never at its pitch cents, intonation, just-noticeable difference
- How small a difference is audible cents, difference limen, just-noticeable difference