A spectrum chooses its own scale
Assumes: Two notes and a ratio, which is the whole of consonance
Feeding two spectra into a roughness model and sweeping one past the other produces a curve with dips at the consonant intervals, and this site has described that result four times as the small whole numbers arriving unbidden. That description is true and it is the wrong causal story, which is a bad combination because nothing in the figure contradicts it.
The dips are not at the small whole numbers. They are wherever two partials of the lower note and the upper note land on the same frequency, and for a harmonic spectrum — where the partials are the small whole multiples — the two coincide. Change the spectrum and the coincidences move, and everything the model says about consonance moves with them.
The exact statement, which is stronger than the usual one
Take a harmonic spectrum and replace every partial n with s raised to log₂n, for some stretch factor s near two. At s = 2 nothing changes; at s = 2.1 every partial moves apart by a fixed proportion of its distance in cents from the fundamental.
Run the roughness model on that spectrum and its minima land at 5.33, 7.51 and 9.47 semitones, against the harmonic set’s 4.98, 7.02 and 8.84.
The ratios are 1.0703, 1.0698 and 1.0713. Log₂ 2.1 is 1.0704.
Every consonance has moved by exactly the stretch factor. Not approximately, and not only the ones near the fundamental: the whole interval structure the model recognises has been scaled by the same number the partials were scaled by, because the model’s only mechanism is coincidence and coincidence is preserved under a uniform stretch.
That is the sharp form of what William Sethares argued in the early 1990s, and it is a stronger claim than “consonance depends on timbre”. It says the relation between a scale and a spectrum is a matching, that either can be computed from the other, and that neither is prior.
The same result in the other direction
Compression is the test that the stretch is not a special case. Take s = 1.9 — partials pulled together rather than apart — and the minima land at 4.61, 6.50 and 8.19 semitones, against the harmonic 4.98, 7.02 and 8.84.
The ratios are 0.926, 0.926 and 0.926. Log₂ 1.9 is 0.9260.
So the relation is exact in both directions and over a range of about twenty per cent, which is far larger than the departures any real instrument shows. The consonances are the spectrum, rescaled, and there is no residue left over for arithmetic to explain.
Why the coincidence is the mechanism
Two tones a fifth apart, both harmonic. The lower note’s third partial is at 3f and the upper note’s second partial is at 2 × 1.5f, which is also 3f. They are the same frequency, they do not beat, and the pair contributes nothing to the roughness sum.
Move the upper note ten cents and those two partials are now 1.8 Hz apart at middle C, which is inside the region where the Plomp–Levelt model peaks, and the pair contributes a great deal. The dip is a hole punched in the roughness by an exact coincidence, and its depth is how much roughness would otherwise have been there.
So the depth of a dip depends on how strong the coinciding partials are. The fifth’s dip is 37 per cent of the curve’s range because it involves the second and third partials, which are loud. The major third’s is not a dip at all but a shoulder, because the coincidence at 5:4 involves the fifth partial, which in an ordinary string spectrum is faint.
Which partials coincide, interval by interval, is a count: the octave shares six of twelve, the fifth four, the fourth three, the major third two. Each shared partial is a dip’s worth of roughness that never happens, and the count is a fact about a harmonic spectrum rather than about the interval. That is why the fifth’s dip is 37 per cent of the curve’s range — it involves the second and third partials, which are loud — and why the major third’s is not a dip at all but a shoulder, since its coincidence at 5:4 involves the fifth partial, which in an ordinary string spectrum is faint.
What a real inharmonic spectrum picks out
A stretched harmonic series is a thought experiment. A bar is not.
A free bar vibrating in flexure has partials at 1, 2.756, 5.404, 8.933 and 13.34 times its fundamental — zeros of a transcendental equation rather than whole numbers, and the same kind of object as a membrane’s Bessel zeros. Run the roughness model over a pair of such tones and the minima fall at 6.94, 8.70, 9.82 and 11.66 semitones.
Two of those are strong: 11.66 at 15 per cent prominence and 8.70 at 14 per cent. Neither is a scale degree in twelve-tone equal temperament. 11.66 semitones is 66 cents flat of an octave and 8.70 is thirty cents from a minor sixth, which is six times the difference limen and unmistakable.
So a metallophone ensemble’s consonances are genuinely not at the keyboard’s intervals, and the model says why without being told anything about any tradition.
The octave is the case where the claim bites hardest
Every interval in this essay moves when the spectrum moves, including the one nobody argues about.
The octave’s dip is the deepest the model produces, and it is deepest because a harmonic tone an octave up shares every one of its partials with the tone below — the second with the first, the fourth with the second, and so on all the way up. That is the maximum possible coincidence and it is why an octave is the strongest consonance there is.
Stretch the spectrum and that total coincidence is gone. Every partial of the upper note is now slightly sharp of one below it, the sharpness grows with the partial number, and the frequency ratio at which the coincidence is best restored is no longer 2:1 but 2.1:1 — which is exactly the 12.85 semitones the figure reports.
The site has already met one half of this from the opposite side: a spectrum with no even partials has no 2:1 coincidence to offer, so the octave stops fusing and the 3:1 starts. That essay’s object was a scale with a different equave; this one’s is the general rule those two are both instances of.
The inverse problem, and why it is the interesting one
If a spectrum determines a set of consonances, then a set of consonances determines a spectrum — the design problem run backwards.
Given a scale, place the partials so that the model’s minima land on its degrees. For an equal division of the octave into n steps the answer is a spectrum whose partials are at powers of 2^(k/n) — that is, partials that are themselves members of the scale. That is why a harmonic spectrum suits a twelve-tone scale so well: the harmonic series’ first six partials are within fourteen cents of scale degrees, and the coincidences the model needs therefore fall on notes the instrument can play.
For a division into ten equal steps, or thirteen, a harmonic spectrum has nothing useful to offer and a purpose-built one does. That is a claim that can be tested by building the timbre, and it has been elsewhere — and it can be run here, on this site’s own roughness function, in a few lines.
Take the six harmonic partial amplitudes this essay uses throughout, and move each partial to the nearest degree of the target division: for ten steps that puts them at 1, 2, 3.031, 4, 4.925 and 6.063 instead of 1 to 6. Then sweep and find the minima.
| division | partials | mean distance of the minima from a scale degree |
|---|---|---|
| 10 | 1, 2, 3.031, 4, 4.925, 6.063 | 0.0 ¢ |
| 12 | 1, 2, 2.997, 4, 5.040, 5.993 | 0.0 ¢ |
| 13 | 1, 2, 3.064, 4, 4.951, 6.128 | 7.3 ¢ |
| 19 | 1, 2, 2.988, 4, 4.979, 5.975 | 0.2 ¢ |
It works, and it works exactly. The ten-step spectrum’s five strongest minima land at 3.60, 4.80, 7.20, 8.40 and 12.00 semitones, every one of which is a degree of ten-equal to the resolution of the sweep. The twelve-step spectrum tidies the harmonic minima from 3.87, 4.98, 7.02, 8.84 onto exactly 4, 5, 7 and 9. Nothing is fitted; the partials are rounded to scale degrees and the consonances follow.
The control, which says twelve is not the special one
The construction only means something with the comparison beside it, so: how far are the harmonic spectrum’s minima from the degrees of each division?
| division | harmonic spectrum’s mean error | as a fraction of one step |
|---|---|---|
| 10 | 21.4 ¢ | 18% |
| 12 | 6.6 ¢ | 6.6% |
| 13 | 26.0 ¢ | 28% |
| 19 | 4.6 ¢ | 7.3% |
Ten and thirteen are badly served, exactly as the essay claims — a fifth to a quarter of the way to the next degree, which is far outside anything a tuner would accept. Twelve is well served.
And so is nineteen, slightly better in cents. The harmonic spectrum’s consonances sit 4.6 cents from a nineteen-equal degree against 6.6 from a twelve-equal one. Read as a fraction of a step the order reverses — 6.6 per cent of twelve’s step against 7.3 per cent of nineteen’s — so which division a harmonic spectrum “suits best” depends on whether the question is how far the error is or how far it is relative to the grid.
That is worth having because the matching argument is usually made as though twelve were picked out. It is not picked out; it is one of a small set of divisions a harmonic spectrum lands on well, and the other members of that set are the ones this site found by scoring approximations to the simple ratios. Two criteria, arrived at independently, selecting the same short list — which is a better argument for twelve than uniqueness would have been, because it is true.
What the inverse problem does not do is choose the scale. It says: given this scale, here is a spectrum that makes it consonant. It is entirely silent on why any tradition would want ten steps rather than twelve, and the next rung of this ladder is about a related and larger silence.
Against the twelve equal steps the just ratios line up unevenly: the fifth misses by 2 cents and the fourth by 2, and the thirds and sixths miss by 14 to 16. A harmonic spectrum’s strong coincidences are at the fifth and the fourth, and those are precisely the two the twelve-fold division gets nearly exactly right — a matching between an instrument and a division, arrived at long before either was described this way. And read the site’s three stock spectra as three consonance systems and the general rule is visible in one comparison: the string is harmonic and its minima are the familiar ones, the clarinet’s near-absent even partials remove the coincidence that builds the fourth’s dip and leave the fifth standing, and the bell’s minima are nowhere in particular. One instrument’s consonant interval is another’s rough one.
Whose music, and when
The general claim is about a model. The specific claims need a repertoire attached, and two of them have one.
Gamelan. Javanese and Balinese ensembles are dominated by struck metal, whose partials are inharmonic in the way a bar’s are, and their tunings are not twelve-tone equal temperament and not built from fifths. The suggestion that these two facts are connected is old and it is not settled by this essay: the model shows that such a spectrum has consonances elsewhere, which makes the connection plausible and does not establish it. The tuning of any two gamelan differs by more than any theory of consonance would predict, which is the strongest evidence against reading the match too tightly.
The stretched octave. Listeners set octaves wide, with pure tones that have no partials at all, and a piano’s stiff strings stretch their own partials. Both facts are in the same direction as this essay’s arithmetic and neither is explained by it — the pure-tone result in particular cannot be, since the model needs partials and there are none.
And the deliberate case. Music written since about 1990 for spectra designed to make a chosen scale consonant is a small repertoire with named pieces, and it is the only music in this essay that exists because of the model rather than being described by it.
Where the model stops
Roughness is one component of consonance and this site has said so four times. The cross-cultural evidence is that roughness discrimination is universal and preference is not, so a computation over partial coincidences predicts one measurable thing and not the thing a listener would call consonance.
The model has no harmonicity term. A second family of models scores an interval by how well the combined spectrum fits a single harmonic template, and those models do not care about coincidence in the same way — for a stretched spectrum they would say something quite different, because a stretched spectrum fits no harmonic template at all. Which of the two is right about a gamelan is not a question this figure can answer.
And it is a model of two tones. Everything here is dyadic. Three notes have three pairs and the pairs interact, which is the finding of the triad essay and is not addressed by any curve drawn against one interval.
What the picture cannot show
It cannot show the strength of a dip in a way a listener would recognise. Prominence as a fraction of the curve’s range is a reasonable measure and it is arbitrary. A 37 per cent dip and a 3 per cent dip are both drawn as dips, and the second is at the edge of being an artefact of where the sweep started.
And the inverse-problem table compares five minima, not a whole curve. Taking the five most prominent dips and measuring their distance from the nearest scale degree is a summary that says nothing about the dips it discarded, nor about how deep any of them is — a division could in principle be matched perfectly on five shallow minima and badly on the rest. What makes the comparison usable is that the same five ranks come out for every spectrum tried, at prominences within a few points of each other, so the five being compared are the same five features seen at different positions. That is checked rather than assumed, and it would be the first thing to break on a spectrum less like a harmonic series than the ones here.
It cannot show that the stretched spectrum’s octave is not an octave. The stretched tone’s strongest minimum is at 12.85 semitones, and a listener presented with two such tones at that distance would not necessarily call them the same note — octave equivalence is a perceptual phenomenon that fusion supports, and a spectrum with no 2:1 coincidence supports it much less.
It cannot show loudness. The Plomp–Levelt weights are amplitude products, so a quiet dissonance and a loud one differ in the model only by a scale factor — and in a listener they do not, because the critical band itself widens with level.
And it cannot show the attack. Half of what identifies an instrument is over before the steady spectrum arrives, and a bar’s attack is most of its sound. The partial list this essay computes from is a description of a decaying steady state that a struck metallophone spends very little time in.
What this does to the site’s founding claim
The first rung of this ladder states that consonance is small whole numbers, and that claim now needs a qualification rather than a retraction.
The qualification is that the small whole numbers are doing their work inside the tone rather than between the tones. A harmonic spectrum is a set of partials at whole-number multiples; two such tones coincide when their frequency ratio is a ratio of small whole numbers, because that is what makes one set of multiples intersect another. The integers are in the instrument.
Take the integers out of the instrument and they leave the intervals with them. That is the whole of this rung, and it means the founding claim should be read as a claim about a pairing — harmonic spectra and small-integer intervals — of which neither half is a fact about music on its own.
It also means the claim is falsifiable in a way it did not look, and has in a small way been tested: build a spectrum whose partials are not integers, ask which intervals are smooth, and the answer is a different set of intervals. The model was not fitted to that case and it predicts it correctly.
The inverse construction above is the same falsification run forwards rather than backwards, and it is the sharper version. Rounding six partials to the degrees of a ten-note division moves every consonance onto that division exactly — from a model with no parameter that knows anything about ten, and from partials that were chosen by rounding rather than by search. The integers are in the instrument, and any integers will do. What twelve has is not a privilege but a coincidence of fit with the one spectrum most instruments happen to produce, and even that fit is shared with nineteen.
Where the ladder goes next
If a spectrum picks out a set of intervals, the obvious next question is whether it picks out a scale — a set of seven notes to play them on. The next rung runs that search over every seven-note selection from the twelve and finds that the model cannot answer, for a reason that is about what a scale is rather than about how good the model is.
Part 5 of 10
One essay in the series on consonance. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 13.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
InharmonicityPartialPlomp–Levelt curveRoughnessSensory dissonanceSpectrumTimbre
- The spectrum that was supposed to explain the gamelan inharmonicity, partial, roughness, sensory dissonance, spectrum
- A clarinet keeps what a string loses partial, roughness, spectrum, timbre
- The body is the filter partial, roughness, spectrum, timbre
- The spectrum that will not fuse inharmonicity, partial, spectrum, timbre
- Where the hammer lands partial, roughness, spectrum, timbre
- A beat has a depth, and six essays held it at one partial, sensory dissonance, spectrum