Intervals and chords

The quantity a rival account says is not there

Two earlier essays measure how much two notes' errors are correlated through a shared anchor, and price what that correlation would be worth. There is a rival account in which a listener refers each note to a key and never forms the distance at all — under which the correlation is not small, it is a description of something that is not happening. The two accounts agree on almost everything and disagree on one manipulation, and the manipulation costs an afternoon.

Assumes: The notes in between · How much an anchor would have to be worth

An interval is two errors is the fourth rung of this ladder and it establishes the object the two rungs after it are about. Judging an interval means judging two pitches, each with its own error, so the error in the interval is the combination of two — and how they combine depends on whether they are independent.

How much an anchor would have to be worth prices the correlation: if the two notes are judged against a shared reference, their errors are correlated, and a correlation of one half halves the interval’s variance. The notes in between asks what happens to that shared reference when there is music between the two notes.

Both of them are about ρ, the correlation between two notes’ errors. And there is a published account of pitch judgement under which ρ is not a small number, it is a description of a thing that does not happen: a listener does not form the distance at all. Each note is referred to an established key, the answer is a pair of scale degrees, and there is no interval anywhere in the process to have an error.

The same interval, started on each of the twelve. An interval of 7 semitones started on each pitch class of a major key, against how strongly the key specifies its two notes — the mean of the probe-tone profile at each. The interval account says the listener encodes a distance, so the key cannot enter and the prediction is a horizontal line at 5.4 cents. The degree account says the listener refers each note to the key, so its precision on a note falls as the key's specification of that note weakens; scaled to agree at the most stable start, it rises from 5.4 cents on C to 8.5 on E♭. Every earlier figure measures a quantity the second account says is not being formed at all.
Fig. 1 The same interval — a fifth — started on each of the twelve pitch classes of a major key, against how strongly the key specifies its two notes. The interval account has no term for the key, so its prediction is a horizontal line. The degree account’s precision depends on how strongly each note is specified, so its prediction is a slope. Both are calibrated to agree at the most stable start, which is what makes the divergence a prediction rather than a scaling.

Two accounts, stated carefully

The interval account. A listener encodes the first note, then encodes the second relative to it. What is stored and compared is a distance. Its error is the error of one relative judgement, which is why an interval can be judged more finely than either of its notes can be identified — a fact this ladder has relied on since its fourth rung.

The degree account. A listener has an established tonal frame and refers each note to it. What is stored is a pair of positions in a scale, and the precision on each depends on how strongly the key specifies that position. A tonic is specified sharply; a chromatic note is not. Nothing forms a distance, so nothing has an interval error.

Both accounts have a great deal going for them and both have direct evidence. The interval account is what musicians describe when asked, and it explains transposition invariance — the fact that a melody survives being moved — with nothing added. The degree account is what counting produced the hierarchy is about, and it explains why a note’s stability is such a strong predictor of almost everything.

The reason the ladder never had to choose is that they make the same prediction almost everywhere. Both say a wider interval is judged less accurately; both say a longer note is judged better; both say two notes with material in between are worse than two notes with silence.

The one place they part

They part on a manipulation that neither is usually asked about: hold the interval and move it around the key.

Under the interval account nothing has changed. The distance is the same distance, the relative judgement is the same judgement, and the key is not in the model. The prediction is a flat line.

Under the degree account everything has changed. A fifth from C to G in C major is two of the three most strongly specified positions in the key; a fifth from E♭ to B♭ in the same key is two positions the key barely specifies at all. The precision on each note is worse, so the interval judgement built from them is worse.

Calibrate the two accounts to agree on the strongest case — which is the honest thing to do, since both have a free parameter and neither has an absolute scale — and the degree account predicts thresholds fifty per cent worse at the weak positions while the interval account predicts no difference at all.

The spread one account predicts and the other cannot. Both accounts are calibrated to agree on the most stable starting note, so the discriminating quantity is what happens at the least stable one. The interval account has no term for the key at all, so it predicts a spread of exactly one at every interval — the flat line. The degree account predicts 1.21 for 1 semitones, 1.38 for 2 semitones, 1.41 for 3 semitones, 1.46 for 4 semitones, 1.57 for 5 semitones, 1.57 for 7 semitones, 1.41 for 9 semitones, 1.69 for 12 semitones, rising with the size of the interval because a wider interval is more likely to have one of its two notes at an unstable degree. A difference of forty to seventy per cent in an interval-discrimination threshold is well inside what an ordinary identification experiment resolves, which is why this is a measurement rather than a thought experiment.
Fig. 2 The spread each account predicts across the twelve starting notes, for eight interval sizes. The interval account’s line is at one everywhere, by construction: it has no key term to produce a spread with. The degree account’s rises with interval size, because a wider interval has more chance of putting one of its two notes somewhere the key does not specify.

Why the spread grows with the interval

The rising line has a reason and it is worth stating, because it is the shape of the prediction rather than its size that would survive a change of parameters.

A narrow interval’s two notes are close together, so they tend to have similar stabilities: two adjacent scale degrees are usually both in the key or both out of it. A wide interval’s two notes are far apart, so their stabilities are nearly independent — and a judgement built from two independent draws is dominated by the worse of them.

So the degree account predicts that the key’s influence on interval discrimination grows with the interval, which is a second, independent signature. An experiment that found a spread but found it flat across interval sizes would be evidence against both accounts as stated, which is the kind of thing a prediction is for.

What the experiment is

It is one of the simplest in this collection and it needs no equipment the field does not own.

Establish a key with a short cadential context. Present an interval of a stated size, started on one of the twelve pitch classes, with the second note mistuned by a small and varying amount. Ask whether it was in tune. Repeat over the twelve starting notes and over three or four interval sizes.

the interval account   thresholds equal at all twelve starts
the degree account     thresholds 40 to 70 per cent worse at
                       the unstable starts, by more for wider
                       intervals

Two things make it a good experiment rather than a merely possible one. The predicted effect is large if the degree account is right outright — a factor of one and a half is not a subtle statistical claim — and the two accounts are calibrated to agree on the reference condition, so the measurement is of a pattern rather than of a level and does not depend on either free parameter. What it is not is a good experiment for the hybrid positions, and a later section works out why: the spread falls away much faster than the weight does.

The probe-tone profile, major against minor. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.
Fig. 3 Where the degree account’s numbers come from: the probe-tone profile, which is the strength with which a key specifies each of the twelve. It is a measurement rather than a model — counting produced the hierarchy is the essay about how — and it is the only ingredient in the prediction that did not have to be stipulated.

What the fifth rung was measuring, under each account

It is worth going back to the rung this one is about and asking what its number means under each reading, because the two answers are not variations on one thing.

The fifth rung sweeps ρ — the correlation between the two notes’ errors — and reports what each value buys: at ρ = 0.5 the interval’s standard deviation is 0.707 of the independent case, and to halve it ρ has to reach 0.75. That is a clean piece of arithmetic and it is agnostic about mechanism.

Under the interval account ρ is the correlation induced by encoding the second note relative to the first: the second note’s error contains the first’s, so they are correlated by construction and the correlation is a property of the encoding.

Under the degree account there is no second note relative to a first. The two errors are correlated because both are referred to the same tonal frame, so ρ is a property of the key — how sharply it is established, how long it survives, and how much of it a listener is holding.

The two produce the same number and the number means different things, and the sixth rung’s manipulation is the one that starts to separate them. Putting intervening notes between the two is neutral under the interval account except as interference, and under the degree account it is a restatement of the key — which is why that rung found two published accounts predicting opposite signs, and why it could not choose between them either.

How much correlation it would take to matterThe limen of a 7-semitone interval at a note length of 0.25 seconds, against the correlation between the two notes' errors. The independent model at the left gives 9.44 cents. Halving that needs a correlation of 0.75; a fifth off it needs 0.31. The curve is √(1 − ρ) and nothing else, so the correlation required for a stated improvement is arithmetic — which turns the question from “does a key help?” into “by how much, and here is the number it must reach”.×1.1×1.2×1.5×2×300.20.40.60.80246810correlation between the two notes' errorsthe interval's limen, centsone note: 7.85 ctwo, independent:9.44 cthe marks are theimprovements a keywould have to buy
Fig. 4 The earlier figure: what a shared anchor is worth, as the correlation between the two notes’ errors is swept. Every value on this axis has two readings, and the difference between them is what this essay is about — the same curve describes an encoding under one account and a key under the other.

Which computation produced the numbers

The interval account’s prediction is a constant: the relative-judgement standard deviation, taken as the ladder’s own 5.4 cents for a quarter-second note, which is the fifth rung’s single-note figure.

The degree account’s is built from two absolute judgements. Each note’s standard deviation is a base value divided by the square root of the key’s specification of that note, taken from PROBE_TONE normalised to its own maximum; the two share the key’s own anchor, with a correlation of one half. The interval’s standard deviation is then the usual combination of two correlated errors.

Three of those are asserted: the base absolute sigma, the exponent of one half on the stability, and the shared-anchor correlation. Every one is a decision and none of them is measured here, which is the standing weakness this collection keeps recording against itself and which the categorical-hearing ladder has begun to price by sweeping.

What the asserted numbers do not affect is the thing the figure is for. The two accounts are renormalised to agree at the most stable start, so a different base sigma moves both curves together and a different exponent changes the size of the spread and not its existence. The prediction that survives every choice is the sign: one account predicts a spread and the other predicts none.

A second manipulation, and why it is weaker

There is an obvious alternative experiment and it is worth saying why it is the second choice rather than the first.

Take the interval out of a key altogether — present it in isolation, with no context — and the degree account has nothing to refer the notes to, so it should collapse toward the interval account’s prediction. That is a real difference and it is the manipulation most of the published work has actually done.

The trouble is what it compares. An isolated interval and an in-context interval differ in a great many ways at once: attention, memory load, the presence of other pitches, the passage of time. A difference between the two is evidence about something, and pinning it on the mechanism requires ruling the rest out.

The manipulation in the hero figure has none of that. Every condition has the same context, the same number of notes, the same duration and the same interval — the only thing that changes is which twelve-tone position the pair occupies, which is a variable the interval account is committed to being blind to. That is what makes it a clean test rather than a suggestive one, and it is the sort of design the categorical-hearing ladder’s own audit argued for on a different question.

How long a note has to be before its pitch is worth arguing about. The smallest audible frequency difference at 440 Hz, against how long the note lasts. The flat line is the steady-tone difference limen of 4.0 cents that every tuning argument on this site rests on. The falling line is the bound a finite duration imposes on its own frequency, 1/2T in cents, which no listener can beat. They cross at 486 milliseconds: below that the note is the limit and above it the listener is. A tenth of a second gives 19.6 cents and a quarter gives 7.9, against the commas drawn across the figure.
Fig. 5 The earlier curve, which sets the scale of everything above: the effective limen against how long a note lasts. Every threshold in this essay is quoted for a quarter-second note, which is where this curve has flattened but not settled, and both accounts inherit the same duration term — so it is one more thing the manipulation holds fixed.

What the experiment could actually resolve

Two of the caveats below are sweeps waiting to be run, and running them changes what the experiment is for.

The hybrid. Nobody holds the degree account in the pure form the figures draw; the usual position is that a listener forms both a distance and a tonal position and weights them. Combine the two estimates by their reliabilities, with a weight w running from a pure interval account at zero to a pure degree account at one, and the predicted spread at an octave is:

w spread threshold precision needed to see it
0.10 1.013 0.4%
0.25 1.037 1.3%
0.50 1.099 3.5%
0.75 1.228 8.1%
1.00 1.687 24.3%

The effect is not linear in the weight and it is not large except at the very end. A listener who uses the tonal cue with half the weight of the interval cue produces a spread of ten per cent, not sixty-nine, because combining two estimates is dominated by the better of them and the interval estimate is the flat one. Solving for what an ordinary experiment could see — thresholds measured to ten per cent, which is a good day’s work — gives a detectable weight of about 0.81.

So the section above overstates what the manipulation buys. It does not measure a mixing weight across its range; it tests whether the weight is near one. A null result would rule out the pure degree account and leave every hybrid weaker than four fifths standing, which is most of the space the published positions occupy.

The exponent. The other asserted number is the power on the stability, and it moves the size of everything:

exponent spread at an octave
0.25 1.30
0.50 (as drawn) 1.69
1.00 2.85
1.50 4.81

That is a factor of nearly four across a range nobody can rule out, so the size quoted anywhere in this essay is worth nothing on its own. What survives both sweeps is the thing the figures were drawn for and it survives untouched: at w = 0 the spread is exactly one at every exponent, and at any w above zero it is above one at every exponent. The sign is a theorem about the model and the magnitude is a guess, which is the right division for a prediction that has to be run rather than believed.

Where the model stops

The degree account is stated in its strongest form and nobody holds it in that form. Most published positions are hybrids, and the section above prices what a hybrid predicts: a spread smaller than the pure account’s and larger than zero, but smaller by much more than the weight is — a half-weighted listener gives a tenth of the effect rather than half of it. So the experiment discriminates near the pure end and not across the range, which is a less useful outcome than first supposed and is worth knowing before running it.

The stability exponent is a guess about a mechanism. Why should precision improve as the square root of a probe-tone rating? There is no derivation here, and sweeping it moves the predicted spread from 1.30 to 4.81 across a range nobody can exclude — which is why the figure reports a pattern and why no number in this essay should be quoted as a prediction of a size.

And the profile is not a within-listener quantity. PROBE_TONE is a mean over listeners and over contexts, and using it as though it described the tonal frame of the listener in the chair is exactly the substitution the listener the model was never run for is about.

What the picture cannot show

It cannot show which account is right, and that is the point of it. What it shows is that the question is decidable and by what, which is a thing five rungs of this ladder have been able to avoid because their measurements do not distinguish.

Nor can it show what a musician would say. Musicians report hearing intervals, unambiguously and without hesitation. That is evidence about a report rather than about a mechanism, and the entire interest of the degree account is that a listener who is doing something else would report the same thing.

And it cannot rescue the fifth and sixth rungs if the answer goes the wrong way. If the degree account is substantially right, then the correlation those rungs measure is not a correlation between two errors in one comparison — it is the shared anchor of the key, which is a different object with a different half-life, and the notes in between would be measuring how long a key lasts rather than how long a note does.

Whose listeners, and when

The probe-tone profile is a measurement on Western listeners with Western tonal exposure, and the whole degree account presupposes an established key — so both the prediction and the experiment are claims about a listener who has one. A listener from a tradition with a different frame would have a different profile and a different predicted spread, and one with no such frame would fall back on the interval account by default, which is itself a prediction.

The historical shape is worth noticing. The interval account is the older one and it is the one music theory is written in: an interval is a named object, taught as a distance, and the whole apparatus of species and inversions is arithmetic on distances. The degree account arrives with the empirical work of the 1970s and 1980s and is expressed in a vocabulary the theory does not use. So a collection like this one, which computes from theory, is structurally biased toward the account that theory happens to be written in — which is the honest reason this rung had to be written.

Where this ladder goes next

Seven rungs. How finely two pitches can be told apart; the octave that is not two to one; the family as a function of note length; an interval as two errors; how much a shared anchor would be worth; what intervening material does to it; and now what the whole of that is worth if a listener is not forming intervals at all.

What the ladder owes after this is not another rung. It is the experiment, and this rung has stated it precisely enough to be run: a context, twelve starting notes, four interval sizes, and a threshold at each. Everything above it in the ladder is conditional on the answer, and the answer is one afternoon in a laboratory that this collection does not have.

Part 7 of 12

One essay in the series on Pitch-acuity. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Difference limenIntervalJust-noticeable differenceKey-findingPitch memoryProbe-toneScale degreeTonal hierarchy