Roughness can be computed, and the answer looks like a scale
Here is an experiment that can be run without any music in it at all.
Take a model of how two nearby frequencies interfere in the ear. Take two complex tones with realistic partials. Fix the lower one, sweep the upper one slowly up through an octave, and at every point sum the interference between every pair of partials. Plot the total.
The curve has structure. It rises steeply from the unison, peaks somewhere around a semitone, and then descends through a series of dips. The dips are at the fourth, the fifth and the octave, with shallower ones at the thirds and the sixth.
Nobody told it to put them there. The input was a partial list and an interference model; the output is something that looks a great deal like a scale.
The model, in full
The mechanism is worth stating precisely, because the whole argument depends on it being modest.
Two pure tones close in frequency are not resolved separately by the ear. They excite overlapping regions of the basilar membrane, and the result is a fluctuating response that is perceived as roughness. How rough depends on how far apart they are relative to the critical bandwidth at that frequency.
The relationship, established by Plomp and Levelt in 1965 from listening tests with pairs of sine tones, has a characteristic shape:
- At zero separation, no roughness — a unison is smooth.
- Roughness rises steeply as separation increases, peaking at roughly a quarter of a critical bandwidth.
- It then falls away, reaching negligible values at about one full critical bandwidth.
Critical bandwidth is itself frequency-dependent: about 100 hertz below 500 hertz, and roughly a fixed fraction — near a minor third — above that. So the same interval is a different fraction of a critical band in different registers, and is correspondingly rougher or smoother.
For two complex tones, the total roughness is the sum over every pair of partials, weighted by their amplitudes. That is the entire computation, and it has exactly two inputs: a partial list and the Plomp–Levelt curve.
Why the dips land where they do
The reason the minima fall on simple ratios is almost embarrassingly direct.
Consider two tones a fifth apart, at frequencies and . The lower tone’s partials are at . The upper tone’s are at . Two of the upper tone’s first four partials — and — coincide exactly with partials of the lower tone.
Coinciding partials do not beat. They contribute zero roughness. And the remaining partials are far enough apart to be resolved separately, which also contributes nothing.
Now consider a tritone at roughly 1.414. The upper tone’s partials fall at 1.414, 2.83, 4.24, 5.66 — every one of them close to a partial of the lower tone without landing on it. Each of those near-misses is inside a critical band, each contributes roughness, and they add.
So the rule “simpler ratios are more consonant” is not a fundamental principle. It is a consequence of the fact that simpler ratios cause more partials to coincide, given partials at whole-number multiples.
The prediction that makes it a theory
An explanation that only accounts for what is already known is not worth much. This one makes a prediction, and the prediction is falsifiable and has been tested.
If the partials move, the consonant intervals move with them.
Feed the same computation a tone whose partials sit at 1, 2.1, 3.3, 4.7 times the fundamental instead of 1, 2, 3, 4, and the curve comes out with dips in entirely different places. Some of the classical intervals lose their minima; new minima appear at ratios that have no names.
This has been done both computationally and with synthesised sound since the 1970s, and it works. Intervals that are dissonant for harmonic tones are consonant for the stretched ones, and listeners agree — the effect does not require training to hear.
There is also a natural experiment. Indonesian gamelan ensembles are built from metallophones and gongs, whose partials are genuinely inharmonic; and gamelan tuning systems, sléndro and pélog, divide the octave in ways that no harmonic-series theory predicts and that fit those instruments’ actual spectra well. The instruments and the scale co-evolved, and the roughness model gets the relationship roughly right where the ratio theory gets it entirely wrong. The same reasoning explains why a bell cannot be tuned by beats.
What the curve gets right about real music
Several features of ordinary practice fall straight out of the curve, and are usually taught as rules with no reason attached.
The semitone is the worst interval, and it is used as a dissonance. The peak of the curve is at roughly a semitone, and every tradition that has a semitone treats it as maximally unstable.
Wide spacing in the bass. Critical bandwidth is proportionally much wider at low frequencies, so a third in the bass is inside one critical band and is genuinely rough, while the same third two octaves up is not. Every orchestration text says to space low chords widely; the reason is the width of a critical band, and the reason is almost never given.
The fourth’s ambiguity. The fourth has a deep roughness minimum — as deep as the fifth’s — and Western theory has treated it as a dissonance requiring resolution in certain contexts for six hundred years. That is a straightforward disagreement between the sensory account and the practice, and it is the clearest evidence that roughness is not the whole story — the fourth’s treatment depends entirely on which note is in the bass.
Reading the curve carefully
The shape rewards more attention than a glance gives it, and two of its features are easy to miss.
The wells are not equally deep. The octave’s minimum goes essentially to zero; the fifth’s is deep; the thirds’ are shallow depressions on a slope rather than wells. That ordering matches the historical order in which intervals were admitted as consonances — octave first, then fifth and fourth, then thirds only in the fifteenth century — which is a suspicious coincidence and probably not one.
The thirds are not minima at all, quite. With eight string-like partials, the major third sits on a shoulder rather than in a dip. That is a real feature of the computation and not an artefact, and it is the single most useful thing the figure shows: the consonance of the third is marginal on a purely sensory account, which is exactly why it took two thousand years to be accepted and why temperaments argued about it so bitterly.
Change the partial list — fewer partials, or weaker upper ones — and the thirds’ shoulders deepen into dips. A harpsichord, whose upper partials decay fast, gives a different curve from a modern piano. Some of the historical disagreement about whether thirds are consonant was a disagreement between instruments.
What a single partial does
The cleanest way to see the mechanism is to remove almost all of it.
Two pure sine tones have one partial each. There is exactly one pair to consider, and the roughness curve for sine tones is just the Plomp–Levelt function itself: a single hump near the unison and flat thereafter. No dip at the fifth, no dip at the octave, nothing.
And that is what sine tones sound like: intervals between them are strangely characterless, and listeners find them hard to identify. A sine-tone fifth is not obviously a fifth. Musicians consistently report this as disconcerting the first time, and it is exactly what the model predicts — with one partial each there is nothing to coincide, so the consonance of the interval has nothing to be made of.
This is the strongest single confirmation available, because it is a case where the ratio theory and the roughness theory make opposite predictions. The ratio 3:2 is just as simple for sine tones as for strings. The consonance is not.
Beating is the mechanism underneath
Roughness is not a primitive. It is beating, at a rate too fast to count.
Two partials one hertz apart produce an audible slow pulsing. Ten hertz apart, the pulsing is too fast to follow and becomes a buzz. Thirty hertz apart, the buzz fades and the two components begin to separate into distinct tones. The Plomp–Levelt curve is a summary of exactly that progression, expressed as a fraction of critical bandwidth rather than in hertz so that it applies at any register.
So the chain runs: two frequencies close together, unresolved by the cochlea, produce fluctuation; fluctuation at a few hertz is heard as pulsing and at a few tens of hertz as roughness; roughness summed over all partial pairs is sensory dissonance; and sensory dissonance is minimised at ratios that make partials coincide. Each link is a measured relationship, and none of them mentions music.
Where the number of partials matters
One parameter changes the picture more than any other, and it is not the interval.
An instrument with two partials produces a nearly flat roughness curve with one shallow octave dip. An instrument with twelve produces a curve with pronounced dips at every simple ratio down to 7:4. The more partials a tone has, the more sharply defined its consonances are — and the more of them there are.
That has a corollary worth sitting with. A bright instrument has more consonance structure than a dull one. A harpsichord, a bowed string or a reed makes the difference between a fifth and a tritone acute; a flute or a stopped organ pipe, whose spectra are nearly pure, makes it mild.
Composers know this without the theory: dense harmony is written for strings and reeds and thinned for flutes, and close dissonant spacing is tolerable in soft flute writing that would be intolerable on oboes. The rule of thumb precedes the explanation by about three hundred years.
What the curve gets wrong
The fourth is not an isolated failure. There is a whole category of musical fact that roughness cannot reach, and being clear about the boundary is more useful than defending the model past it.
Dissonance is partly expectation. The dominant seventh chord has substantial roughness and was, in eighteenth-century practice, a dissonance requiring resolution. By 1920 it was a stable chord to end a blues on. Nothing about its spectrum changed. What changed was what a listener expected to happen next, which is a fact about a repertoire and not about a cochlea, and which belongs to how progressions work rather than to acoustics.
Consonance is partly harmonicity. Two tones whose frequencies suggest a common fundamental sound fused, and that appears to be a separate cue from roughness. An octave and a twelfth are fused even when spaced widely enough that no partials are near enough to beat at all. Roughness predicts nothing there; harmonicity does.
Context reverses judgements. A chord that is rough in isolation can be the resolution of something rougher, and it will be heard as restful. Roughness is computed pairwise on a static sound; music is neither pairwise nor static.
The honest summary is that the roughness model explains sensory dissonance — the immediate grating quality of a sound — completely and quantitatively, and explains musical dissonance only partially. The two words are usually used interchangeably, and the confusion causes most of the arguments.
The instrument chooses the scale
The strongest consequence of the model is a reversal of the usual causal story, and it is worth stating baldly.
The usual story: musical scales derive from the harmonic series, which is a fact of physics, so scales are grounded in nature.
The model’s story: musical scales derive from the spectra of the available instruments. For strings and pipes, those spectra are harmonic, so the scales that result are built on small integer ratios. For inharmonic instruments, different scales result, and they are equally well grounded.
That is a much weaker claim than the usual one and a much more defensible one. It explains why so many traditions converged on the octave and the fifth — every string and every pipe produces them — and it explains why traditions with metallophones did not converge on the same thirds.
It also predicts something that can be checked: a culture whose principal instruments are inharmonic should have a scale that the harmonic series does not explain. Gamelan is that culture, and it does.
Whose music, and when
The mathematics is culture-free and the interpretation is not, and it is easy to overreach here.
Plomp and Levelt’s listening tests were conducted with Dutch subjects in the 1960s, and later work has repeated the sine-tone roughness measurements across populations with broadly consistent results — the critical-band mechanism is anatomical, and anatomy does not vary much. That part travels.
What does not travel is the preference. Roughness is a measurable property; whether roughness is wanted is an aesthetic decision, and traditions differ sharply. Bulgarian women’s choral singing uses sustained seconds as a central sonority. Balinese gamelan tunes paired instruments a few hertz apart to produce deliberate beating. Both are choices to use roughness rather than to minimise it, and describing either as out of tune is a category error.
The model says how rough a sound is. It does not say that less is better, and every claim that it does has smuggled in a preference.
Where the model stops
Pairwise only. Roughness is computed between pairs of partials and summed. Three or more simultaneous tones may interact in ways the pairwise sum misses, and the model’s extension to full chords is an approximation with known weaknesses.
Static. The computation assumes steady tones. Real notes have attacks and decays, and a rough sonority that lasts a tenth of a second is not heard as rough at all.
No level axis. Roughness grows with loudness and the model as usually stated does not include it. Every curve on this page is for one unspecified level.
One spectrum drawn. The hero figure uses a single string-like partial list. The essay’s central claim is that a different list gives a different curve, and only one curve appears — a real limitation of the picture, and the reason the argument leans on the gamelan case rather than on the drawing.
The model is old. Plomp and Levelt’s 1965 formulation has been refined many times since, and modern versions use auditory filterbanks rather than a single critical-band curve. The refinements move the details and not the conclusions.
The ladder from here
Later rungs: critical bandwidth, measured and modelled. The Plomp–Levelt curve in full derivation. Harmonicity as an independent cue. Roughness in chords rather than intervals. Stretched and compressed spectra, with the synthesis to hear them. Gamelan tuning and its instruments. Sethares’s method for deriving a scale from a spectrum. Dissonance as expectation, and the history of the dominant seventh. Bulgarian seconds, and roughness used deliberately. And amplitude, which every figure here has quietly held constant.
The most striking thing about the curve is how little goes into it. Two spectra and one psychophysical function, measured in a laboratory with sine tones by people who were not asking a musical question, and out comes a picture with a fifth and an octave in it.