Intervals and chords

How much an anchor would have to be worth

Two pitch errors added in quadrature assume an independence nobody measured — a listener inside a key hears a note as a scale degree, and a shared reference is exactly a correlated error. Turning the dial is not evidence. What is evidence is that the dial is not free: a shared error cancels out of a difference completely, so a listener's single-note limen and their interval limen give the two components with nothing left over, and halving the interval limen needs a correlation of exactly 0.75.

Assumes: An interval is two errors · The ear sorts into boxes, and the boxes are the theory

An interval is two errors treated a melodic interval as the difference of two independent pitch estimates and added their limens in quadrature. It also recorded, in its own words, where it could be attacked: if a listener holds a reference — a tonic, a drone, the memory of the first note as an absolute — then the two errors share a component and partly cancel.

That rung declined to draw the correlated model, on the grounds that a curve with a dial on it that can be turned to reach any answer is not evidence. The objection is correct and it is not the end of the matter, because the dial is not free.

How much correlation it would take to matterThe limen of a 7-semitone interval at a note length of 0.25 seconds, against the correlation between the two notes' errors. The independent model at the left gives 9.44 cents. Halving that needs a correlation of 0.75; a fifth off it needs 0.31. The curve is √(1 − ρ) and nothing else, so the correlation required for a stated improvement is arithmetic — which turns the question from “does a key help?” into “by how much, and here is the number it must reach”.×1.1×1.2×1.5×2×300.20.40.60.80246810correlation between the two notes' errorsthe interval's limen, centsone note: 7.85 ctwo, independent:9.44 cthe marks are theimprovements a keywould have to buy
Fig. 1 The limen of a fifth at a quarter-second note length, against the correlation between the two notes’ errors. The whole curve is √(1 − ρ) times the independent value, so the correlation needed to buy a stated improvement is arithmetic rather than opinion: a fifth off the independent limen needs 0.31, and halving it needs 0.75 exactly.

A factor of two requires a correlation of three quarters. That is a specific, testable, and rather large number, and stating it is the whole of what this rung does that the previous one could not.

Why it is a decomposition and not a fit

The useful move is to stop writing the model as a correlation at all.

Write each note’s estimate as the truth plus two errors: a common one the two notes of an interval share, because both are referred to the same key, the same drone, the same recently-heard reference; and an independent one they do not. Then a single note’s variance is σc² + σi², and the difference of two notes has variance 2σi² — because the common part is in both estimates with the same sign and subtracts out completely.

A shared error cancels out of a difference completely. A single-note limen of 7.85 cents split into a part the two notes of an interval share and a part they do not, at five values of the correlation. The shared part cancels out of a difference entirely, so the interval limen is √2 times the independent part alone: at a correlation of 0.75 it is 5.55 cents against the 11.10 the independent model gives. The point of the split is that the correlation is not a free parameter — it is what the ratio of two measurable limens says it is.
Fig. 2 One note’s limen of 7.85 cents split into its shared and independent halves at five correlations, with the resulting interval limen marked. The shared half does nothing to an interval, however large it is; only the independent half survives the subtraction. At a correlation of 0.75 the shared part is 6.8 cents and the independent part 3.9, and the interval limen is √2 times the second alone.

ρ = σc²/(σc² + σi²), which is not a parameter to be fitted but a ratio of two quantities that can each be measured on the same listener in the same afternoon. Measure a single-note limen. Measure an interval limen. The second divided by √2 is σi, the first with σi taken out of it in quadrature is σc, and there is nothing left over.

This is the same manoeuvre the loudness ladder made with its two models and the same one the missing-fundamental ladder made when it closed: a parameter that looks free stops being free the moment two independent measurements constrain it, and finding the second measurement is usually the whole of the work.

That is the difference between a dial and a rung. The previous essay was right that a free ρ proves nothing; what it did not have was the observation that ρ is over-determined by two measurements that both already exist as standard psychophysical tasks.

How long a note has to be before its pitch is worth arguing about. The smallest audible frequency difference at 440 Hz, against how long the note lasts. The flat line is the steady-tone difference limen of 4.0 cents that every tuning argument on this site rests on. The falling line is the bound a finite duration imposes on its own frequency, 1/2T in cents, which no listener can beat. They cross at 486 milliseconds: below that the note is the limit and above it the listener is. A tenth of a second gives 19.6 cents and a quarter gives 7.9, against the commas drawn across the figure.
Fig. 3 The single-note limen the split is taken of: how finely one note’s pitch can be specified against how long it lasts, with the note lengths of ordinary tempi on it. Below about a fifth of a second the limiting term is not the ear at all but the Fourier bound — a tone of duration T cannot have a frequency finer than about 1/2T — and that term is in the independent half of the split by construction, because it is a property of the signal rather than of the listener’s reference.

That last observation is a real constraint on how large ρ can be. The Fourier bound cannot be shared between two notes: it is set by each note’s own duration and it is in every estimate independently. So at short note lengths, where the bound dominates, ρ has a ceiling that has nothing to do with how good the listener’s anchor is.

“Well below” is not a number, and the ceiling is computable exactly: σi is at least the bound, so ρ is at most 1 − (bound/limen)². At A440:

note length limen of which the bound ρ ceiling best improvement
0.25 s 7.85 7.85 0.00 none
0.4 s 4.91 4.91 0.00 none
0.6 s 4.04 3.28 0.34 1.23×
1 s 4.04 1.97 0.76 2.06×
2 s 4.04 0.98 0.94 4.11×

At a quarter of a second the ceiling is not “well below 0.75” — it is exactly zero. Below the crossover at 0.486 seconds the bound is the limen, there is no listener-limited component for an anchor to share, and no reference of any kind can improve an interval judgement at all. The hero figure at the top of this essay is drawn at a quarter of a second, which is a note length at which its own answer is unreachable.

Putting that on the tempo axis is where it bites. A crotchet at 60 gives a ceiling of 0.76 and could just support the factor of two; at 90 it is 0.47 and at 120 it is 0.06 — a best-possible improvement of three per cent. At A440 the mechanism this rung proposes to measure can only operate on notes at or below about sixty to the minute, and the threshold moves with pitch: at 1760 hertz the bound falls fast enough that a crotchet at 217 would still do.

So the prediction sharpens from the anchor should help long notes and not short ones into something with an edge on it: it is high, slow notes where an anchor can be worth anything, and the experiment’s two conditions have to be run there or they will find nothing whatever the listener’s reference is doing.

That edge is sharper than the physics warrants and the reason is a modelling convention. The single-note limen is the larger of the two bounds rather than their quadrature sum, so the ceiling snaps to zero the moment the bound overtakes the ear. Combining them in quadrature instead softens it without changing the conclusion: 0.21 at a quarter second, 0.51 at a half, 0.81 at one — a best improvement of 1.12×, 1.43× and 2.29×. Either way the factor of two needs about a second.

The experiment, and what it would have to find

The run has two conditions and one variable.

The ceiling also decides where the run has to be done, which the two conditions below do not say. A stimulus of quarter-second notes at A440 is the natural laboratory choice — it is short enough to be a melodic interval and long enough to have a pitch — and it is the one duration at which the model guarantees a null result. A null at a quarter of a second would be evidence about the Fourier bound and not about the listener, and it would be indistinguishable from the substantive null the section below describes. The two conditions have to be run at a second or more, on notes high enough that a second is musically ordinary, or the experiment cannot tell its own two failure modes apart.

Condition one: no key. Two successive tones, an interval apart, in isolation — no preceding context, no drone, ideally at pitches that do not sit in any familiar scale. Measure the smallest change in the interval a listener can detect.

Condition two: an established key. The same two tones preceded by a cadence or a scale that puts them at recognisable scale degrees, so that a categorical anchor is available. Measure the same thing.

The ratio of the two limens is √(1 − ρ) and gives ρ directly. And the size the effect has to reach to matter is the hero figure’s marks: under about 1.2, it is a curiosity; at 1.5 it changes which melodic intervals are protected; at 2 it means every number in the previous rung is out by a factor of two.

How finely a fifth can be heard, against how fast it goes by. The smallest audible mistuning of a melodic fifth above 440 hertz, against how long each of its two notes lasts. The lower solid line is one note's own limen; the upper one is the interval's, larger because two independent errors add in quadrature. A tenth of a second gives 23.5 cents against the note's own 19.6, the floor for long notes is 5.4, and the syntonic comma is not cleared until each note lasts 111 milliseconds. The dotted line at 1.31 cents is the same interval heard as a simultaneity, where partials 3 and 2 coincide at 1320 hertz and one beat every 2 seconds can be counted there.
Fig. 4 The earlier figure, which is what the experiment would be testing: the interval limen against note length, independent model, with the same interval judged as a simultaneity for comparison. Every point on the melodic curve moves down by √(1 − ρ) if the anchor is real. The simultaneity curve does not move at all, because a beat between coinciding partials is counted rather than estimated and has no pitch memory in it.

The simultaneity is the control. If a listener’s melodic interval limen improves in a key and their harmonic one does not, the improvement is in the pitch-estimation stage and this model is the right one. If both improve, something more general is going on — attention, arousal, familiarity — and the correlation reading is wrong.

It is worth saying what a null result would mean, because it is not nothing. If the two conditions give the same limen, then a listener inside a key is not referring successive notes to a shared standard when judging how far apart they are — which is a substantive claim about a mechanism that a great deal of tonal-perception writing takes for granted. The independent model would then be right rather than merely convenient, and every number in the previous rung would stand as computed rather than as an upper bound.

What a categorical anchor would actually do

There is a complication, and it is the one that makes the prediction interesting rather than merely quantitative.

The ear sorts into boxes, and the boxes are the theory. A scale degree is not a reference point; it is a category with a width, and referring a note to a category is a quantisation rather than a measurement. That has two consequences pulling opposite ways.

For an interval inside the category system — a fifth, a third, a step of the scale in use — a category is a strong shared reference and ρ should be large. For an interval outside it — a neutral third, a quarter tone, anything the listener’s own system has no box for — the same machinery is a liability: the note is assigned to the nearest box and the assignment discards the very deviation the task is asking about.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 5 The categories the anchor would be made of, drawn as the identification functions measured earlier. A note near a centre is named confidently and a note near a boundary is not — so the strength of the anchor is not one number but a function of where in its category the note falls, which is a dependence no version of the correlated model above contains.

So the prediction is not that a key improves interval discrimination. It is that a key improves it for in-category intervals and degrades it for out-of-category ones, and that the improvement should be largest for notes near category centres and smallest for notes near boundaries — which the categorical ladder has independently measured as barely movable.

That is a shape rather than a number, and a shape is much harder to get by accident.

How much correlation it would take to matterThe limen of a 4-semitone interval at a note length of 0.1 seconds, against the correlation between the two notes' errors. The independent model at the left gives 24.99 cents. Halving that needs a correlation of 0.75; a fifth off it needs 0.31. The curve is √(1 − ρ) and nothing else, so the correlation required for a stated improvement is arithmetic — which turns the question from “does a key help?” into “by how much, and here is the number it must reach”.×1.1×1.2×1.5×2×300.20.40.60.80510152025correlation between the two notes' errorsthe interval's limen, centsone note: 19.56 ctwo, independent:24.99 cthe marks are theimprovements a keywould have to buy
Fig. 6 The same curve for a major third at a tenth of a second, which is a note at the fast end of a melody. The independent limen is much larger — the Fourier bound is doing all of the work — and the correlation needed for any given improvement is unchanged, because the shape does not depend on the note. What changes is that at a tenth of a second every part of the error is of the kind an anchor cannot touch, so the curve is drawn over a range of ρ the note itself cannot reach. It is the right picture of an arithmetic relation and the wrong picture of this note.

Which computation produced the numbers

The single-note limen is durationLimen unchanged from the third rung: the larger of the published steady-tone discrimination limen at that frequency and the Fourier bound of 1/2T, converted to cents. At A440 and a quarter of a second it is 7.85 cents, and the bound rather than the ear is what is limiting there.

The interval limen is √(a² + b² − 2ρab) with a and b the two notes’ limens. For an interval small enough that a ≈ b this reduces to a√(2(1 − ρ)), which is why the whole curve is one shape scaled: the ratio to the independent case is √(1 − ρ) with no dependence on frequency, note length or interval size.

The correlation required for a factor k is therefore 1 − 1/k², which is 0.31 for 1.2, 0.56 for 1.5 and 0.75 for 2. Those are the marks on the hero figure and they are exact.

The split into shared and independent components is the same algebra read the other way, with the single-note limen as the total and ρ as the share of its variance that is common.

Where the model stops

The common error is assumed to be common in the same amount to both notes. Two notes of an interval are at different pitches, and a reference held as a scale degree is a different distance from each of them; there is no reason the shared component should be identical rather than merely correlated. Allowing them to differ replaces one parameter with two and the experiment above measures only the combination.

A correlation of one is nonsense and the model does not know it. At ρ = 1 the interval limen is zero, which is the model announcing that it has been pushed past where it means anything. The decomposition version does not have that defect — σi cannot be zero because it contains the Fourier bound — and that is a second reason to prefer it.

And nothing here contains time. A shared reference has to be held, and pitch memory decays on a timescale of seconds. Two notes a semiquaver apart share a reference almost perfectly; two notes a phrase apart share a degraded one. So ρ is a function of the gap between the notes, and every figure here is drawn at a single gap of zero.

Whose music, and what turns on it

The claim is about listeners, so it applies wherever there is a scale. What turns on it is a claim this collection has made repeatedly about repertoires.

The published band for melodic interval intonation — the amount by which a sung or played interval can be off before a trained listener objects — is twenty-five to fifty cents, and the fourth rung derived a number inside it from two bounds computed for other reasons. If ρ is 0.75, the derived number is half of what was drawn and no longer lands in the band, which would mean either that the band is measuring something else or that the anchor is not there.

That matters most for the traditions this site has argued about at the edges. A degree is where it goes next is about a system in which a scale degree is defined by its behaviour rather than by its pitch, and a listener inside such a system has a very strong categorical reference indeed; the third the model has no opinion about is about intervals for which a twelve-category listener has no box. The prediction above says those two listeners should differ in opposite directions on the same stimulus, which is the cleanest cross-cultural test this collection has been able to state.

What the picture cannot show

Whether the anchor is a category or a pitch. A drone gives an absolute reference; a key gives a categorical one; the memory of the first note gives a third thing. All three produce a correlated error and this model cannot tell them apart, and the three make different predictions about what happens when the key is established and then removed.

Nor whether the two conditions differ only in the anchor. Establishing a key means playing more notes, and more notes mean more time, more attention and a possible expectation effect on the target itself. A control with a non-tonal context of the same length is the obvious remedy and it is not obvious that a non-tonal context of the same length exists for a listener who has one.

The ceiling is as much a modelling artefact as a finding. It rests on the single-note limen being the larger of two bounds and on the Fourier bound being wholly independent between two notes. The first is a convention this ladder adopted for a different reason; the second is an argument rather than a measurement, and a listener who integrated a note’s frequency over a window longer than the note would violate it.

And nothing here is measured. Every number above is a model evaluated at a parameter the model does not determine. What has been added to the previous rung is not a measurement but a specification: the size the effect must have, the two limens that would measure it, and the control that would say whether it is the right explanation.

The size that is already known

One piece of the answer is not missing, and it is worth putting beside the rest.

A drone is the strongest possible shared reference: an actual sounding pitch, continuously available, requiring no memory at all. And the intonation this collection has measured against a drone is very fine indeed — a tuner working against a held note resolves a fraction of a cent, because the judgement is not two estimates at all but a beat count.

That is the ceiling case and it says the mechanism is real: a shared reference does collapse a two-estimate problem into a one-estimate problem, when the reference is physically present. The question this rung asks is how much of that survives when the reference is remembered rather than sounding, and the answer is somewhere between the drone’s fraction of a cent and the independent model’s nine.

Where this ladder goes next

Five rungs. How finely two pitches can be told apart; that the octave is not where it should be; that the whole family is a function of note length; that an interval is two errors and is heard less finely than either note; and now, that the independence in that model is a measurable ratio rather than an assumption, with a stated size it would have to reach.

The rung after it is the one the decay limitation names. Every figure in this ladder is about two notes with nothing between them, and a melody is notes with other notes between them — so the shared reference is not held across a gap but across material, and the material is itself pitched. Whether an intervening note strengthens the anchor by confirming the key or weakens it by overwriting the memory is a question with two published literatures and one arithmetic, and this collection has the arithmetic.

Part 5 of 12

One essay in the series on Pitch-acuity. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionCentsDifference limenIntonationJust-noticeable differenceNull modelPitch memoryScale degree