Intervals and chords

The same distance, under two names

Four hundred cents is a major third or a diminished fourth, and on a keyboard nothing in the sound distinguishes them. An earlier essay was about the boundary between two categories; this is about two categories at one acoustic value, and the surprise is where the ambiguity comes from. In quarter-comma meantone a major third is 386 cents and a diminished fourth is 427 — two names, two pitches, forty-one cents apart. Equal temperament collapsed them, and what a listener now supplies from context used to be in the sound.

Assumes: The ear sorts into boxes, and the boxes are the theory

The first rung of this ladder is about what happens between two categories: slide one note continuously against another and the report jumps rather than sliding, with the change compressed into twenty cents or so.

That is a boundary problem — one continuum, two boxes, and a question about where the wall is. There is a second and stranger case in the same subject, and it has the opposite shape: two boxes at one point on the continuum, with nothing in the sound to say which.

The case

Play a C and the E above it. That is four hundred cents on a piano, and it is a major third.

Play a C and the F flat above it. That is also four hundred cents on a piano, on the same two keys, and it is a diminished fourth.

One sound, two intervals. A phrase written on a stave. Notation records what a player should do rather than what the air does, so it shows the note names exactly and the pitches only by convention — which is the reason so much of the evidence here is drawn some other way.
Fig. 1 Two intervals in notation. On a keyboard the second pair is the first pair — the same two keys, the same four hundred cents, the same waveform. In notation they are different intervals with different names, different sizes as scale steps, and different obligations about what comes next. The button plays them, and there is nothing to hear.

These are not two ways of describing one thing. A major third is a third: its two notes are two letter-names apart, it belongs to C major, and it behaves as a stable consonance. A diminished fourth is a fourth: its notes are three letter-names apart, it belongs to keys with many flats, and in practice it resolves inward. They are as different functionally as any two intervals in the system, and identical acoustically. That combination is rare enough to be worth dwelling on: nearly every distinction music theory makes is either audible or admitted to be a convention, and this one is audible in some tunings, inaudible in others, and treated as real in all of them.

Where the two names come from

The names are not arbitrary and they are not merely orthographic. They are positions on the chain of fifths.

The chain of fifths in quarter-comma meantone. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and one of them — G♯ to D♯, the 12th link, where the chain is forced to close — is the wolf, at 35.7 cents.
Fig. 2 The chain of pure-ish fifths from which every note name is generated. An interval’s name is the number of steps along this chain between its two notes: a major third is four fifths up, a fourth is one fifth down, and a diminished fourth is eight fifths down. The chain is where spelling lives, and it is one-dimensional and infinite.

Reading it off: a major third is +4 on the chain, a diminished fourth is −8. Those differ by twelve, and twelve steps along the chain is the interval the whole tuning ladder is about — twelve fifths against seven octaves, which is not zero.

So the two spellings differ by exactly one turn of the chain, and whether they are the same pitch depends entirely on whether the chain closes.

That gives the whole essay’s quantity a closed form, and it is worth having because it is simpler than any of the tables below. A major third is four fifths up and a diminished fourth is eight fifths down, so their difference is twelve fifths less seven octaves — which is to say

the enharmonic gap is exactly twelve times the amount by which the fifth misses 700 cents.

Nothing else enters. Not the third, not the temperament’s name, not which comma it was designed around. Every cent a tuning narrows its fifth opens the gap by twelve, and a tuning whose fifth is 700 cents has no gap because twelve times nothing is nothing.

fifth major 3rd diminished 4th gap
Pythagorean 701.96 407.8 384.4 −23.5
53 equal 701.89 407.5 384.9 −22.6
12 equal 700.00 400.0 400.0 0
1/6-comma meantone 698.37 393.5 413.0 19.6
1/5-comma meantone 697.65 390.6 418.8 28.2
31 equal 696.77 387.1 425.8 38.7
quarter-comma meantone 696.58 386.3 427.4 41.1
19 equal 694.74 378.9 442.1 63.2

Two things fall out of that table that the tuning tables elsewhere on this site do not show.

Equal temperament is not merely where the gap closes. It is where the two names cross over. Above 700 cents the diminished fourth is the narrower of the pair — 384 against 408 in Pythagorean — and below it the diminished fourth is the wider, 427 against 386 in quarter-comma. So a musician trained in Pythagorean tuning and one trained in meantone would both hear the distinction clearly and would disagree about which way round it goes. Twelve equal sits exactly at the crossing, which is a stronger statement than “the distinction is lost”: it is lost at the one point where losing it requires no rounding.

The chain of fifths in Pythagorean. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and one of them — F♯ to C♯, the 10th link, where the chain is forced to close — is the wolf, at -23.5 cents.
Fig. 3 And the chain on the other side of the crossing. Pythagorean tuning keeps eleven fifths pure and leaves the twelfth 23.5 cents narrow, which puts its fifth 1.955 cents above 700 rather than below — so its enharmonic gap comes out negative, and its diminished fourth is 23 cents narrower than its major third rather than 41 cents wider. Both tunings keep the distinction audibly; they disagree about which of the two names is the bigger interval. Twelve equal sits at the crossing, which is why the collapse there is exact rather than approximate.

And the distinction is very cheap to keep. Twenty cents is comfortably audible on a sustained chord, and twelve times a fifth’s deviation reaches twenty cents when the fifth misses 700 by 1.67 cents. Quarter-comma meantone misses it by 3.42, which is twice what is needed; even sixth-comma, the gentlest meantone in common use, misses it by 1.63 and lands within a whisker of the threshold. Essentially any tuning that is not equal keeps the two names apart audibly, which is why the split keyboards were worth building and why they stopped being worth building the moment they were not.

How many names, over how long a chain

The claim made below that every one of the twelve keys has at least two names in ordinary use, and five of them three, is a claim about how far along the chain “ordinary use” runs, and it is worth making that explicit because the number is not obvious.

Positions n and n + 12 on the chain give the same key, so twelve classes each get a second name only once the chain runs to 24 consecutive fifths, and exactly five get a third at 29. That is the range the claim assumes, and it reaches well into double accidentals — the span from F double flat to B double sharp is 35 positions and gives all twelve a second name and eleven of them a third.

Inside the range a key signature can actually reach, seven flats to seven sharps, there are only three enharmonic pairs: D♭ and C♯, G♭ and F♯, C♭ and B. Those are precisely the three pairs of enharmonically equivalent key signatures every musician learns, and the coincidence is not one — the pairs of keys are the pairs of tonics.

In meantone they are two different pitches

Equal temperament closes the chain by definition, so +4 and −8 land on the same key. Nothing else does.

Tuning Major third Diminished fourth Apart by
Quarter-comma meantone 386.3 ¢ 427.4 ¢ 41.1
Pythagorean 407.8 384.4 −23.5
Equal 400.0 400.0 0

Forty-one cents is the lesser diesis, 128:125, and it is a large interval — a fifth of a semitone, ten times the smallest audible pitch difference, and obvious to anybody on a sustained chord.

The chain of fifths in quarter-comma meantone. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and no one link carries the closure: the departure is spread across all 6 of them at -5.4 cents each.
Fig. 4 The same chain stopped after six fifths, which is what a chain looks like before anybody insists it close. Every link is 5.4 cents narrow, every one of them is usable, and there is no wolf anywhere — because a wolf is not a property of a tuning but of a decision to bend twelve of these links into a circle. Stop at six and the two spellings this essay is about are simply two different positions, forty-one cents apart, with nothing forcing them together.

So the question this essay opened with — how does a listener know which interval it is, when nothing in the sound says — is a question that did not exist before about 1700. On the instruments the notation was designed for, the sound said.

Equal temperament did not merely retune the notes. It destroyed a distinction that the notation still carries.

How many such pairs there are

The major third and diminished fourth are one instance and it is worth knowing how general the phenomenon is, because the answer is: completely.

Every position on the chain of fifths is a distinct note name, and the chain is infinite in both directions. Twelve equal temperament maps that infinite line onto twelve keys, so every key carries infinitely many names — C is also B sharp and D double flat and A triple sharp, and each of those is a different note in any tuning with an open chain.

The same 41 cents arrives by a second route that has nothing to do with fifths. Stack three pure major thirds — C to E, E to G♯, G♯ to B♯ — and the top note should be the octave. It is not: three 5:4 thirds make 1159.1 cents, and the shortfall is the same lesser diesis, 128:125. Any keyboard player in meantone discovers it immediately, because the tuning that makes each of those thirds pure is exactly the tuning that leaves them 41 cents short of closing.

The chain of fifths in just intonation. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and the departures are not alike: 9 at 0.0, 2 at -21.5, 1 at +19.6 cents. The error sits on a few links rather than on one or on all of them.
Fig. 5 A twelve-note just tuning read as a chain, which shows what “choosing twelve names” costs. Nine of its links are pure fifths and three are not — two 21.5 cents narrow, one 19.6 wide — because a twelve-note just scale is built from thirds as well as fifths and the chain has to absorb the difference wherever the two disagree. The tuning has already made the enharmonic decision before a note is played: it supplies twelve names out of infinitely many, and the notes it cannot play are the ones whose names it did not choose.

In practice only a few of these pairs are used, because notation only reaches about seven flats to seven sharps. But the count of usable enharmonic equivalences is not small: at the twelve-note limit every one of the twelve keys has at least two names in ordinary use, and five of them have three.

What a keyboard with more keys did about it

The historical evidence that this mattered is that people built instruments to avoid it.

Split-key keyboards, with separate levers for D sharp and E flat, are documented from the sixteenth century onward and were not rare; the archicembalo of Vicentino had thirty-one notes to the octave; several Italian and Spanish organs of the period have divided accidentals. All of them exist to keep two enharmonically equivalent notes apart, because in the tuning they were built for those notes are 41 cents apart and using one for the other is audibly wrong.

Nineteen and thirty-one are the divisions that give good thirds, and both of them keep this distinction: in thirty-one, a major third is ten steps and a diminished fourth is eleven. The property is not a coincidence of those numbers. A division keeps the pair apart exactly when its best fifth is not 700 cents, and both of those divisions have a fifth well below it.

The essay on why the chain was never closed makes the same point from the other end. A tradition that never assumes closure never has this problem, and several outside Europe do not.

Why the collapse happened at twelve and not elsewhere

There is a reason this particular ambiguity is the one everybody has and it is arithmetic rather than historical.

Twelve is the number of fifths that very nearly closes an octave — twelve of them overshoot seven octaves by 23.5 cents, which is small enough to distribute and pretend away. Nineteen and thirty-one and fifty-three also close, more accurately, and each of them closes further along the chain: in thirty-one, the chain has to run thirty-one steps before it returns, so the first enharmonic collapse is between notes thirty-one fifths apart rather than twelve.

The chain of fifths in equal temperament. The fifths laid end to end as the chain they are. The bar under each shows how far that fifth departs from a pure three-to-two, and no one link carries the closure: the departure is spread across all 12 of them at -2.0 cents each.
Fig. 6 Equal temperament’s chain: twelve identical links, each 1.955 cents narrow of pure, and no wolf because the error is divided rather than dumped. That uniformity is the whole of the collapse. The enharmonic gap is twelve times the fifth’s departure from 700 cents, the departure here is zero by construction, and so the gap is zero — not approximately, not by rounding, but as an identity. Every distinction this essay is about is destroyed by that one decision, and nothing about hearing selected it.

That is the general statement: an equal division is a decision about which pairs of names become one note, and twelve’s decision includes the pair this essay is about. Nothing about hearing selected it.

So what is a listener doing?

Given the sound is now genuinely ambiguous, something has to supply the missing information, and the answer is: everything except the interval.

What supplies the missing information is key-fitting from everything else present. A four-hundred-cent interval in a passage whose other notes fit C major is a major third; the same interval where the surrounding notes fit A♭ minor is a diminished fourth. The disambiguating evidence is never in the interval itself. It is entirely in what is around it, which is why a listener’s answer can be changed by a note that arrives afterwards.

Which makes the enharmonic case a categorical phenomenon of an unusual kind. The categories of the first rung are anchored in the sound: a minor third and a major third are 300 and 400 cents, and the category boundary is a place on the continuum. These categories have the same anchor, so the boundary is nowhere and the assignment has to come from the context.

The listener is not doing perception here at all in the sense the first rung’s experiments measured. They are doing inference, on evidence that arrives before and after the interval in question.

Where one interval stops being itself. Identification as a function of interval size: the probability that a listener names each category, modelled as a logistic with the boundary positions and sharpness a study reports. One category's share falls from three-quarters to one-quarter over 24 cents, against the 100 that separate adjacent categories — so the change of mind happens in 24% of the gap and the rest of it is not in doubt at all. That is what makes a mistuned third a third that is out, rather than a different interval.
Fig. 7 The identification curves established earlier, over the region in question. The boundaries between adjacent categories are places where identification swings over about thirty cents. There is no analogous curve for the major third against the diminished fourth, and there cannot be one: the two would be a single vertical line at 400 cents, which is not a boundary but a coincidence.

The composer’s use of it

The collapse created something as well as destroying something, and the thing it created is worth naming because it became a major device.

If a major third and a diminished fourth are the same sound, then a passage may enter as one and leave as the other. That is the enharmonic modulation, and it is not available in a tuning where the two differ — in meantone the modulation is audibly a mistuning, which is why it is essentially absent from music written before equal temperament became normal and pervasive after.

So a nineteenth-century composer moving from one key to a remote one by reinterpreting a diminished seventh is exploiting an ambiguity created by a tuning decision. The device and the temperament arrived together, and neither is intelligible without the other.

The same shape, one field along

There is a structural parallel in this collection worth drawing out, because it is the same defect wearing different clothes.

A necklace cannot record where a rhythm starts: rotation-invariance identifies a timeline with all of its own rotations, and which rotation is played is the thing the timeline is for. A pitch-class set cannot record which note is home, for the same reason one dimension along.

Here the collapse is imposed by an instrument rather than by a mathematical convenience, and it goes the other way: the notation keeps the distinction and the sound has lost it. So a performer reading from a score has information the listener does not, which is an unusual state of affairs and is the reason enharmonic spelling is so often dismissed as pedantry by people who have only ever heard equal temperament.

It is also why no rule set founded on sound can separate the pair. Rank every interval by roughness and a consonance–dissonance division cuts that ranking somewhere; where two spellings share a roughness value exactly — which is precisely what an enharmonic pair is in equal temperament — the division has nothing to cut on. Every historical rule set separates them anyway, and every historical rule set was written for instruments on which they were two different sounds.

Which computation produced the numbers

An interval’s spelling is its signed position on the chain of fifths, and its size in a tuning is that number of the tuning’s fifths reduced into an octave. A major third is four fifths up less two octaves; a diminished fourth is eight fifths down plus five octaves. In quarter-comma meantone, where the fifth is the fourth root of five, those are 386.3 and 427.4 cents; in Pythagorean, where the fifth is 3:2, they are 407.8 and 384.4.

The gap between them is twelve fifths less seven octaves, with the sign depending on whether the fifths are wide or narrow of 700 cents. It is 41.1 cents in quarter-comma meantone — the lesser diesis exactly — and −23.5 in Pythagorean, which is the Pythagorean comma. Both are computed from the fifth rather than looked up, and the agreement with the named commas is the check.

Because the expression is twelve fifths less a constant, it is linear in the fifth, and that is what makes the table of temperaments a straight line rather than a set of separate calculations. Every row of it is the same multiplication. The equal divisions are computed the same way, from a fifth of 1200 times the nearest whole number of steps over the division, which is why nineteen and thirty-one appear on the meantone side of the crossing and fifty-three on the Pythagorean side — a fact about where each division’s best fifth falls relative to 700 cents, and not a separate property of the division.

The name census walks the chain and reads off a letter and an accidental count from the position, which is the same one-dimensional map the spelling argument rests on; the pitch class is seven times the position, modulo twelve. Nothing in it knows about a keyboard.

Whose notation, and when

The spelling system is European staff notation and its logic is the chain of fifths — which is why the system needs double sharps and double flats, and why it has never been replaced by something simpler: any notation with one symbol per key throws away the distinction this essay is about, and composers have consistently refused notations that do.

Split keyboards and enharmonic instruments are documented from about 1480 to about 1700 in Italy, Spain, Germany and England. Their disappearance tracks the spread of well temperament and then equal temperament, which is what makes the story tidy: the instruments existed to preserve a distinction, and vanished when the distinction did. That is a cleaner historical argument than most in this area, because it does not depend on reading anybody’s intentions: an instrument with split keys costs more to build and more to play, and nobody pays for a feature that does nothing.

Traditions without a fixed twelve-note keyboard never made the collapse and do not have the ambiguity. A performer with continuous pitch — a singer, a string player — can and does place a diminished fourth differently from a major third, and the melodic-intonation argument says which direction: toward wherever it is going.

What the picture cannot show

It cannot show what performers actually do. Whether string quartets and choirs really distinguish enharmonic spellings in performance is an empirical question with a thin and inconsistent literature. The tuning arithmetic says they could; it does not say they do.

It cannot show the boundary case. Between a passage that is clearly in one key and one clearly in another there is a great deal of music that is neither, and in that music the spelling on the page is a decision by the composer or the editor rather than a fact to be recovered.

It cannot show how long the inference takes. Establishing a key from scratch takes a measurable number of events, and a spelling decision that depends on the key cannot be faster than that. An interval heard cold, with no context yet, has no spelling at all — which is a prediction and is not tested here.

And the key-fitting model is coarse. It counts pitch classes and weights them, with no term for order, register or which notes are metrically strong — which is exactly the gap the segmentation essay is about. It is being used here to show that context can carry the information, not that this is how the information is read.

The ladder from here

Three rungs: a boundary that is sharper than the sound, a category that survives mistuning, and now a pair of categories with no acoustic separation at all.

The obvious fourth is the one this essay keeps touching and does not do: measure it. Take listeners who read music, give them four-hundred-cent intervals in contexts that imply a major third and a diminished fourth, and ask them to notate what they heard. If the spelling is genuinely supplied by context the answers will track the context; if it is supplied by the notation they have been trained on, they will track the key signature. Nobody appears to have run it.

Part 3 of 11

One essay in the series on Categorical-hearing. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Categorical perceptionChain of fifthsEnharmonicIntervalLesser diesisMeantoneNotation