What a temperament cannot do
Assumes: Twelve fifths and seven octaves, which are not the same thing
Equal temperament is usually described as the most accurate of the keyboard tunings, or as the fairest, or as the one that gave up the most. All three descriptions assume that the systems differ in how much error they contain, and that the choice between them trades total error against something else.
They do not differ in how much error they contain. They cannot.
The invariant, and its one-line proof
Take any tuning of twelve pitch classes in which the chain of fifths closes — that is, any tuning a keyboard can actually be in, since a keyboard has to be playable in every key even if some of them are horrible. Ask for the major third above each of the twelve notes, and add the twelve answers.
The answer is 4,800 cents, and the reason is that the major thirds partition the twelve notes into four cycles of three. C to E to G sharp to C is one such cycle; it closes on the octave by construction, so its three thirds sum to 1,200 cents whatever they individually are. There are four such cycles, covering all twelve notes exactly once each. Four octaves is 4,800 cents.
Nothing in that argument mentions how the fifths were distributed. It does not mention fifths at all. The total is fixed by the requirement that the tuning closes, and the requirement that it closes is what makes it a temperament rather than a chain that overshoots.
So the mean major third in every temperament is 400 cents, exactly. A pure 5:4 is 386.31 cents. The mean departure from pure is 13.686 cents, in every temperament that has ever been proposed and every one that ever will be.
The same argument on the fifths gives 8,400 cents — seven octaves, one cycle of twelve — so the mean fifth is 700 cents everywhere too, and the mean fifth error is 1.955 cents everywhere.
It is not about thirds and fifths. It is about every interval.
The cycle argument has to be run once per interval, and running it eleven times turns up something the two special cases hide. Take the interval of k semitones above each of the twelve notes and add the twelve answers:
| total | total | ||
|---|---|---|---|
| minor 2nd | 1,200 | tritone | 7,200 |
| major 2nd | 2,400 | fifth | 8,400 |
| minor 3rd | 3,600 | minor 6th | 9,600 |
| major 3rd | 4,800 | major 6th | 10,800 |
| fourth | 6,000 | minor 7th | 12,000 |
| major 7th | 13,200 |
The twelve instances of the interval of k semitones total exactly 1,200k cents, for every k from one to eleven, in every closing temperament. Verified against all six systems the site carries and against a temperament generated from twelve random numbers scaled to close: eleven totals, seven systems, every one exact.
And there is a proof of the general statement that is shorter than the four-cycle one and does not need cases. The sum is over the twelve roots of upper minus lower. Every pitch class appears exactly once as an upper note and exactly once as a lower note, so all twelve of them cancel — and what is left is the octave that has to be added each time the interval runs off the top of the set and wraps round. Stepping up by k semitones from each of the twelve notes wraps exactly k times. So the total is 1,200k, and nothing about the tuning enters the argument at any point.
That makes the essay’s negative claim much wider than the headline states it. The mean size of every interval is fixed at 100k cents in every temperament, so the mean departure from just is fixed for every interval at once: −11.7 cents on the minor second, −15.6 on the minor third, +13.7 on the major third, +9.8 on the tritone, −2.0 on the fifth, +15.6 on the major sixth. No temperament is on average closer to just on any interval than any other. The choice is always and only a distribution.
The errors also come in equal and opposite pairs — the major third at +13.686 and the minor sixth at −13.686, the minor third at −15.641 and the major sixth at +15.641 — which is not a further fact but the same one seen twice, since an interval and its octave complement must sum to 1,200 in each of the twelve keys.
What is actually being chosen
If the total is fixed, then a temperament is not choosing a quantity. It is choosing a shape: which of the twelve keys carry the error that has to exist.
Read as a table it becomes clear what each system optimises.
Equal temperament wins the worst key and the spread, and it wins them by putting every key at the mean. That is a minimax solution: the maximum error is as small as it can be, and it is as small as it can be precisely because nothing is allowed to be better than average either. There is no key in which equal temperament sounds good, and that is not a criticism of it — it is the definition of what it does.
Every well temperament wins the best key and pays for it in the worst. Kirnberger III has one key with a pure major third, at zero cents, and one at 21.5. Werckmeister’s best is 3.9 and its worst is 21.5. What they are buying is a home key that is better than equal temperament can ever be, at the cost of remote keys that are worse than equal temperament can ever be.
Pythagorean tuning is the one row where the mean column is different, and the reason is worth stating because it is the only exception to the invariant’s practical form. Its twelve thirds are not all sharp of pure: four of them are 1.95 cents flat. The signed total is still 4,800, exactly as the proof requires, but the average of the absolute values is 14.99 rather than 13.686, because errors of opposite sign no longer cancel in the way the mean assumes. Pythagorean tuning is the only system here that manages to be worse than average on average.
Those four almost-pure thirds are the schisma, and they are the subject of a later rung.
The claim, tested rather than asserted
A one-line proof is the kind of thing that is easy to state and easy to have got wrong, so the site’s own machinery was pointed at it before any of the above was written.
The function that builds a temperament here takes a list of twelve narrowings, one per link of the chain, and refuses any list that does not sum to a Pythagorean comma — a chain that does not close describes an instrument that cannot be played in every key, and that refusal has been in the code since the essay about where to hide the comma. Handed a list of twelve arbitrary numbers scaled to sum correctly, it produces a temperament nobody has ever proposed and nobody would want.
Its twelve major thirds total 4,800.000000 cents.
That is the test the claim needed, because it is the case the proof is about: the invariant is not a property of the historical systems, which might share it by some accident of how they were designed. It is a property of closure, and a temperament chosen at random has it as firmly as Vallotti does. The five systems in the figure above are five points sampled from a space in which every point has the same total.
What the assertion caught, in passing, is that the statement has to be made about the signed total and not the mean absolute error. Those two are the same number for five of the six systems here and differ by 1.3 cents for the sixth, and writing the invariant in terms of the average error — which is how it is tempting to state it, because that is the quantity a reader cares about — would have made the essay’s headline claim false for Pythagorean tuning and true for everything else, with no obvious reason for the exception.
What the invariant does not forbid
It is a constraint on one interval at a time, and it is worth being precise about how little it rules out.
It does not say the thirds and the fifths are traded against each other, though they are: the meantone family makes that trade explicit, and both totals hold along the whole of it. It does not say a temperament cannot be better for a particular piece, which is the only sense in which any of them is better. And it does not say the twelve values have to be near each other — the spread column runs from zero to 21.5 cents, which is the full width the budget allows.
What it forbids is exactly one thing: a temperament that is closer to just intonation everywhere. That system is the one every proposal is implicitly offered as, and it is the one that cannot exist on twelve notes.
It is worth noticing what the general form of the invariant does not buy, because the eleven totals could be mistaken for eleven constraints. They are not independent: fixing the twelve fifths fixes the whole tuning, so once the fifths close, every other interval’s total follows and no further freedom has been removed. The eleven rows of the table are one fact written eleven ways, which is why a designer never feels the other ten as constraints — they are already satisfied by the time a temperament exists at all. What the eleven rows are good for is closing off eleven separate places somebody might have hoped to find slack.
The picture of a distribution
The distribution is more legible drawn round a circle than along a row, because the keys are related in a circle and a temperament’s shape is usually a fact about the circle.
The two figures enclose the same amount, and one of them has a direction. That is the whole difference between a regular and an irregular temperament, drawn.
Why the argument never ended
An argument in which one side has more of a good thing ends when somebody measures it. An argument about where to put a fixed quantity does not end, because the answer depends on a fact outside the arithmetic: which keys the music is going to use.
That makes the historical record read differently. The eighteenth-century writers who preferred a well temperament were not making a mistake about accuracy — they were declining to spread the error over keys their repertoire never visited, which is a perfectly rational thing to do with a fixed budget. And the nineteenth-century move to equal temperament was not the discovery of a better tuning. It was a decision that the repertoire would visit every key, taken in advance of the music that would do so.
The decision and the music arrive together, and each is usually offered as the reason for the other. What the invariant shows is that neither can be justified by the tuning being more in tune, because none of them is.
The bet is also, in the end, what decided the argument. A tuning that favours an arc is only worth its cost while the music stays inside the arc, and once the arc is the whole circle a distribution with a direction has nothing left to buy.
Drawn as deviations from equal temperament, degree by degree, four circulating temperaments are all within a few cents of equal everywhere — the whole disagreement is small — and all four contain exactly the same total error. What separates them is which of the twelve degrees is displaced and in which direction. That is why the argument never ended by measurement: an argument in which one side has more of a good thing ends when somebody measures it, and an argument about where to put a fixed quantity depends on a fact outside the arithmetic, which is which keys the music is going to use.
The one thing that does reduce the error
There is a way to lower the mean third error, and the invariant says exactly what it is: stop requiring that the chain close on twelve notes.
If the tuning has more than twelve pitch classes, the four-cycle argument no longer applies, because the thirds no longer partition the notes into closed cycles of three. Nineteen, thirty-one and fifty-three steps to the octave each do better than 13.686 cents on the thirds, and they do it by having more notes rather than by distributing more cleverly.
Thirty-one equal steps gives major thirds 0.8 cents from pure. That is not a better distribution of the twelve-note budget; it is a different budget, obtained by building an instrument with thirty-one keys to the octave, which several people did.
The same escape is available to anything without a keyboard. A singer or a violinist is not tuning twelve pitch classes and is not subject to the invariant at all; the constraint applies to instruments that commit to a finite set of pitches in advance, which is what a temperament is for.
Whose music, and when
The four temperaments in the table are documents with dates: Werckmeister III from 1691, Vallotti from 1754, Kirnberger III from 1779, Young II from 1799. They were published as instructions for tuning a keyboard, and each is a list of which fifths to narrow and by how much.
What none of them is, is a claim about a piece. The long argument about which temperament Bach intended for the collection of twenty-four preludes and fugues is an argument about evidence that does not exist in the sources, and it is not settled by this essay or by any other. What the invariant does settle is the shape of the question: whichever temperament was intended, it contained the same total error as every other candidate, and the choice between the candidates is a choice about which keys are favoured.
The one claim that can be made from the arithmetic alone is negative and firm. A performance in a well temperament is not more in tune than one in equal temperament. It is differently in tune, key by key, and a listener who prefers it prefers a distribution.
Where the model stops
The invariant is about cents and says nothing about audibility. Twelve thirds summing to 4,800 does not mean twelve equally objectionable thirds. A third 21.5 cents sharp in the bass is a different event from one 21.5 cents sharp two octaves up, because roughness depends on register, and the invariant is blind to that.
It assumes twelve pitch classes and pure octaves. Both are assumptions about the instrument. A real piano’s octaves are stretched, so its twelve “pitch classes” are not the same twelve at the top of the keyboard as at the bottom, and the four-cycle argument holds only within one octave at a time.
And the mean is the wrong statistic for a listener. Nothing in a listener’s experience averages the twelve keys. A piece is in one key, modulates to three or four others, and the thirds it uses are those. Weighting the twelve by how often the repertoire visits them is the calculation that would actually decide the question — and it needs an encoded corpus, which is the limitation this site records rather than one it can repair.
What the picture cannot show
It cannot show what a key sounds like. The spoke diagram gives the size of one interval above each of the twelve roots. A key’s character, if it has one, involves its fifths, its minor thirds, the register the music sits in and the instrument it is played on, and this figure has one of those.
And it cannot show the wolf. Every system drawn here circulates, so every one of its twelve thirds is playable. Quarter-comma meantone does not appear because its chain does not close, and it is the temperament for which the invariant genuinely does not hold: it has eight excellent thirds, four dreadful ones and a fifth that cannot be used, which is a different bargain from any of these.
Where the ladder goes next
Every gap in this ladder so far has had to be paid for and the last two rungs have been about where. The next one is about the exception — the one gap in the subject small enough that no system has ever needed to do anything about it, and the tuning that spends it deliberately because it is free.
Part 9 of 12
One essay in the series on the comma. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
CentsChain of fifthsCirculating temperamentEqual temperamentKey colourTemperamentWell temperament
- An open string pulls the quartet flat chain of fifths, equal temperament, key colour
- A comma is a polyrhythm that never closes chain of fifths, equal temperament
- A guitar tuned by harmonics hides a comma chain of fifths, equal temperament
- A note that is never at its pitch cents, temperament
- A standard is a point, and a performance is a band cents, temperament
- A wind instrument is a thermometer cents, temperament