The same distance, under two names
Assumes: The ear sorts into boxes, and the boxes are the theory
The first rung of this ladder is about what happens between two categories: slide one note continuously against another and the report jumps rather than sliding, with the change compressed into twenty cents or so.
That is a boundary problem — one continuum, two boxes, and a question about where the wall is. There is a second and stranger case in the same subject, and it has the opposite shape: two boxes at one point on the continuum, with nothing in the sound to say which.
The case
Play a C and the E above it. That is four hundred cents on a piano, and it is a major third.
Play a C and the F flat above it. That is also four hundred cents on a piano, on the same two keys, and it is a diminished fourth.
These are not two ways of describing one thing. A major third is a third: its two notes are two letter-names apart, it belongs to C major, and it behaves as a stable consonance. A diminished fourth is a fourth: its notes are three letter-names apart, it belongs to keys with many flats, and in practice it resolves inward. They are as different functionally as any two intervals in the system, and identical acoustically. That combination is rare enough to be worth dwelling on: nearly every distinction music theory makes is either audible or admitted to be a convention, and this one is audible in some tunings, inaudible in others, and treated as real in all of them.
Where the two names come from
The names are not arbitrary and they are not merely orthographic. They are positions on the chain of fifths.
Reading it off: a major third is +4 on the chain, a diminished fourth is −8. Those differ by twelve, and twelve steps along the chain is the interval the whole tuning ladder is about — twelve fifths against seven octaves, which is not zero.
So the two spellings differ by exactly one turn of the chain, and whether they are the same pitch depends entirely on whether the chain closes.
That gives the whole essay’s quantity a closed form, and it is worth having because it is simpler than any of the tables below. A major third is four fifths up and a diminished fourth is eight fifths down, so their difference is twelve fifths less seven octaves — which is to say
the enharmonic gap is exactly twelve times the amount by which the fifth misses 700 cents.
Nothing else enters. Not the third, not the temperament’s name, not which comma it was designed around. Every cent a tuning narrows its fifth opens the gap by twelve, and a tuning whose fifth is 700 cents has no gap because twelve times nothing is nothing.
| fifth | major 3rd | diminished 4th | gap | |
|---|---|---|---|---|
| Pythagorean | 701.96 | 407.8 | 384.4 | −23.5 |
| 53 equal | 701.89 | 407.5 | 384.9 | −22.6 |
| 12 equal | 700.00 | 400.0 | 400.0 | 0 |
| 1/6-comma meantone | 698.37 | 393.5 | 413.0 | 19.6 |
| 1/5-comma meantone | 697.65 | 390.6 | 418.8 | 28.2 |
| 31 equal | 696.77 | 387.1 | 425.8 | 38.7 |
| quarter-comma meantone | 696.58 | 386.3 | 427.4 | 41.1 |
| 19 equal | 694.74 | 378.9 | 442.1 | 63.2 |
Two things fall out of that table that the tuning tables elsewhere on this site do not show.
Equal temperament is not merely where the gap closes. It is where the two names cross over. Above 700 cents the diminished fourth is the narrower of the pair — 384 against 408 in Pythagorean — and below it the diminished fourth is the wider, 427 against 386 in quarter-comma. So a musician trained in Pythagorean tuning and one trained in meantone would both hear the distinction clearly and would disagree about which way round it goes. Twelve equal sits exactly at the crossing, which is a stronger statement than “the distinction is lost”: it is lost at the one point where losing it requires no rounding.
And the distinction is very cheap to keep. Twenty cents is comfortably audible on a sustained chord, and twelve times a fifth’s deviation reaches twenty cents when the fifth misses 700 by 1.67 cents. Quarter-comma meantone misses it by 3.42, which is twice what is needed; even sixth-comma, the gentlest meantone in common use, misses it by 1.63 and lands within a whisker of the threshold. Essentially any tuning that is not equal keeps the two names apart audibly, which is why the split keyboards were worth building and why they stopped being worth building the moment they were not.
How many names, over how long a chain
The claim made below that every one of the twelve keys has at least two names in ordinary use, and five of them three, is a claim about how far along the chain “ordinary use” runs, and it is worth making that explicit because the number is not obvious.
Positions n and n + 12 on the chain give the same key, so twelve classes each get a second name only once the chain runs to 24 consecutive fifths, and exactly five get a third at 29. That is the range the claim assumes, and it reaches well into double accidentals — the span from F double flat to B double sharp is 35 positions and gives all twelve a second name and eleven of them a third.
Inside the range a key signature can actually reach, seven flats to seven sharps, there are only three enharmonic pairs: D♭ and C♯, G♭ and F♯, C♭ and B. Those are precisely the three pairs of enharmonically equivalent key signatures every musician learns, and the coincidence is not one — the pairs of keys are the pairs of tonics.
In meantone they are two different pitches
Equal temperament closes the chain by definition, so +4 and −8 land on the same key. Nothing else does.
| Tuning | Major third | Diminished fourth | Apart by |
|---|---|---|---|
| Quarter-comma meantone | 386.3 ¢ | 427.4 ¢ | 41.1 |
| Pythagorean | 407.8 | 384.4 | −23.5 |
| Equal | 400.0 | 400.0 | 0 |
Forty-one cents is the lesser diesis, 128:125, and it is a large interval — a fifth of a semitone, ten times the smallest audible pitch difference, and obvious to anybody on a sustained chord.
So the question this essay opened with — how does a listener know which interval it is, when nothing in the sound says — is a question that did not exist before about 1700. On the instruments the notation was designed for, the sound said.
Equal temperament did not merely retune the notes. It destroyed a distinction that the notation still carries.
How many such pairs there are
The major third and diminished fourth are one instance and it is worth knowing how general the phenomenon is, because the answer is: completely.
Every position on the chain of fifths is a distinct note name, and the chain is infinite in both directions. Twelve equal temperament maps that infinite line onto twelve keys, so every key carries infinitely many names — C is also B sharp and D double flat and A triple sharp, and each of those is a different note in any tuning with an open chain.
The same 41 cents arrives by a second route that has nothing to do with fifths. Stack three pure major thirds — C to E, E to G♯, G♯ to B♯ — and the top note should be the octave. It is not: three 5:4 thirds make 1159.1 cents, and the shortfall is the same lesser diesis, 128:125. Any keyboard player in meantone discovers it immediately, because the tuning that makes each of those thirds pure is exactly the tuning that leaves them 41 cents short of closing.
In practice only a few of these pairs are used, because notation only reaches about seven flats to seven sharps. But the count of usable enharmonic equivalences is not small: at the twelve-note limit every one of the twelve keys has at least two names in ordinary use, and five of them have three.
What a keyboard with more keys did about it
The historical evidence that this mattered is that people built instruments to avoid it.
Split-key keyboards, with separate levers for D sharp and E flat, are documented from the sixteenth century onward and were not rare; the archicembalo of Vicentino had thirty-one notes to the octave; several Italian and Spanish organs of the period have divided accidentals. All of them exist to keep two enharmonically equivalent notes apart, because in the tuning they were built for those notes are 41 cents apart and using one for the other is audibly wrong.
Nineteen and thirty-one are the divisions that give good thirds, and both of them keep this distinction: in thirty-one, a major third is ten steps and a diminished fourth is eleven. The property is not a coincidence of those numbers. A division keeps the pair apart exactly when its best fifth is not 700 cents, and both of those divisions have a fifth well below it.
The essay on why the chain was never closed makes the same point from the other end. A tradition that never assumes closure never has this problem, and several outside Europe do not.
Why the collapse happened at twelve and not elsewhere
There is a reason this particular ambiguity is the one everybody has and it is arithmetic rather than historical.
Twelve is the number of fifths that very nearly closes an octave — twelve of them overshoot seven octaves by 23.5 cents, which is small enough to distribute and pretend away. Nineteen and thirty-one and fifty-three also close, more accurately, and each of them closes further along the chain: in thirty-one, the chain has to run thirty-one steps before it returns, so the first enharmonic collapse is between notes thirty-one fifths apart rather than twelve.
That is the general statement: an equal division is a decision about which pairs of names become one note, and twelve’s decision includes the pair this essay is about. Nothing about hearing selected it.
So what is a listener doing?
Given the sound is now genuinely ambiguous, something has to supply the missing information, and the answer is: everything except the interval.
What supplies the missing information is key-fitting from everything else present. A four-hundred-cent interval in a passage whose other notes fit C major is a major third; the same interval where the surrounding notes fit A♭ minor is a diminished fourth. The disambiguating evidence is never in the interval itself. It is entirely in what is around it, which is why a listener’s answer can be changed by a note that arrives afterwards.
Which makes the enharmonic case a categorical phenomenon of an unusual kind. The categories of the first rung are anchored in the sound: a minor third and a major third are 300 and 400 cents, and the category boundary is a place on the continuum. These categories have the same anchor, so the boundary is nowhere and the assignment has to come from the context.
The listener is not doing perception here at all in the sense the first rung’s experiments measured. They are doing inference, on evidence that arrives before and after the interval in question.
The composer’s use of it
The collapse created something as well as destroying something, and the thing it created is worth naming because it became a major device.
If a major third and a diminished fourth are the same sound, then a passage may enter as one and leave as the other. That is the enharmonic modulation, and it is not available in a tuning where the two differ — in meantone the modulation is audibly a mistuning, which is why it is essentially absent from music written before equal temperament became normal and pervasive after.
So a nineteenth-century composer moving from one key to a remote one by reinterpreting a diminished seventh is exploiting an ambiguity created by a tuning decision. The device and the temperament arrived together, and neither is intelligible without the other.
The same shape, one field along
There is a structural parallel in this collection worth drawing out, because it is the same defect wearing different clothes.
A necklace cannot record where a rhythm starts: rotation-invariance identifies a timeline with all of its own rotations, and which rotation is played is the thing the timeline is for. A pitch-class set cannot record which note is home, for the same reason one dimension along.
Here the collapse is imposed by an instrument rather than by a mathematical convenience, and it goes the other way: the notation keeps the distinction and the sound has lost it. So a performer reading from a score has information the listener does not, which is an unusual state of affairs and is the reason enharmonic spelling is so often dismissed as pedantry by people who have only ever heard equal temperament.
It is also why no rule set founded on sound can separate the pair. Rank every interval by roughness and a consonance–dissonance division cuts that ranking somewhere; where two spellings share a roughness value exactly — which is precisely what an enharmonic pair is in equal temperament — the division has nothing to cut on. Every historical rule set separates them anyway, and every historical rule set was written for instruments on which they were two different sounds.
Which computation produced the numbers
An interval’s spelling is its signed position on the chain of fifths, and its size in a tuning is that number of the tuning’s fifths reduced into an octave. A major third is four fifths up less two octaves; a diminished fourth is eight fifths down plus five octaves. In quarter-comma meantone, where the fifth is the fourth root of five, those are 386.3 and 427.4 cents; in Pythagorean, where the fifth is 3:2, they are 407.8 and 384.4.
The gap between them is twelve fifths less seven octaves, with the sign depending on whether the fifths are wide or narrow of 700 cents. It is 41.1 cents in quarter-comma meantone — the lesser diesis exactly — and −23.5 in Pythagorean, which is the Pythagorean comma. Both are computed from the fifth rather than looked up, and the agreement with the named commas is the check.
Because the expression is twelve fifths less a constant, it is linear in the fifth, and that is what makes the table of temperaments a straight line rather than a set of separate calculations. Every row of it is the same multiplication. The equal divisions are computed the same way, from a fifth of 1200 times the nearest whole number of steps over the division, which is why nineteen and thirty-one appear on the meantone side of the crossing and fifty-three on the Pythagorean side — a fact about where each division’s best fifth falls relative to 700 cents, and not a separate property of the division.
The name census walks the chain and reads off a letter and an accidental count from the position, which is the same one-dimensional map the spelling argument rests on; the pitch class is seven times the position, modulo twelve. Nothing in it knows about a keyboard.
Whose notation, and when
The spelling system is European staff notation and its logic is the chain of fifths — which is why the system needs double sharps and double flats, and why it has never been replaced by something simpler: any notation with one symbol per key throws away the distinction this essay is about, and composers have consistently refused notations that do.
Split keyboards and enharmonic instruments are documented from about 1480 to about 1700 in Italy, Spain, Germany and England. Their disappearance tracks the spread of well temperament and then equal temperament, which is what makes the story tidy: the instruments existed to preserve a distinction, and vanished when the distinction did. That is a cleaner historical argument than most in this area, because it does not depend on reading anybody’s intentions: an instrument with split keys costs more to build and more to play, and nobody pays for a feature that does nothing.
Traditions without a fixed twelve-note keyboard never made the collapse and do not have the ambiguity. A performer with continuous pitch — a singer, a string player — can and does place a diminished fourth differently from a major third, and the melodic-intonation argument says which direction: toward wherever it is going.
What the picture cannot show
It cannot show what performers actually do. Whether string quartets and choirs really distinguish enharmonic spellings in performance is an empirical question with a thin and inconsistent literature. The tuning arithmetic says they could; it does not say they do.
It cannot show the boundary case. Between a passage that is clearly in one key and one clearly in another there is a great deal of music that is neither, and in that music the spelling on the page is a decision by the composer or the editor rather than a fact to be recovered.
It cannot show how long the inference takes. Establishing a key from scratch takes a measurable number of events, and a spelling decision that depends on the key cannot be faster than that. An interval heard cold, with no context yet, has no spelling at all — which is a prediction and is not tested here.
And the key-fitting model is coarse. It counts pitch classes and weights them, with no term for order, register or which notes are metrically strong — which is exactly the gap the segmentation essay is about. It is being used here to show that context can carry the information, not that this is how the information is read.
The ladder from here
Three rungs: a boundary that is sharper than the sound, a category that survives mistuning, and now a pair of categories with no acoustic separation at all.
The obvious fourth is the one this essay keeps touching and does not do: measure it. Take listeners who read music, give them four-hundred-cent intervals in contexts that imply a major third and a diminished fourth, and ask them to notate what they heard. If the spelling is genuinely supplied by context the answers will track the context; if it is supplied by the notation they have been trained on, they will track the key signature. Nobody appears to have run it.
Part 3 of 11
One essay in the series on Categorical-hearing. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
Categorical perceptionChain of fifthsEnharmonicIntervalLesser diesisMeantoneNotation
- Two names for one key chain of fifths, enharmonic, lesser diesis, meantone, notation
- The notations invented for the overflow chain of fifths, enharmonic, notation