Intervals and chords

Roughness can be computed, and the answer looks like a scale

Feed two spectra into a model of how nearby frequencies interfere, sweep one past the other, and the curve that comes out has dips exactly where the consonant intervals are. Nobody put them there.

Here is an experiment that can be run without any music in it at all.

Take a model of how two nearby frequencies interfere in the ear. Take two complex tones with realistic partials. Fix the lower one, sweep the upper one slowly up through an octave, and at every point sum the interference between every pair of partials. Plot the total.

Roughness across an octaveSensory dissonance between two complex tones as the upper one is swept through an octave, computed by summing the roughness between every pair of their partials. Nothing here is placed by hand. The deep wells land on the fourth, the fifth and the octave; the thirds sit on shoulders rather than in wells, which is a real feature of this model and not a defect of the drawing.semitones above the lower tone6/55/44/33/25/32/1CC♯DE♭EFF♯GA♭AB♭BCroughest at about a semitonesmooth at the simple ratios
Fig. 1 The result. Roughness between two complex tones as the upper one is swept through an octave, computed by summing the interference between every pair of their partials. Nothing in this curve is placed by hand — the vertical marks are added afterwards, at the simple ratios, to show where the features landed.

The curve has structure. It rises steeply from the unison, peaks somewhere around a semitone, and then descends through a series of dips. The dips are at the fourth, the fifth and the octave, with shallower ones at the thirds and the sixth.

Nobody told it to put them there. The input was a partial list and an interference model; the output is something that looks a great deal like a scale.

The model, in full

The mechanism is worth stating precisely, because the whole argument depends on it being modest.

Two pure tones close in frequency are not resolved separately by the ear. They excite overlapping regions of the basilar membrane, and the result is a fluctuating response that is perceived as roughness. How rough depends on how far apart they are relative to the critical bandwidth at that frequency.

The relationship, established by Plomp and Levelt in 1965 from listening tests with pairs of sine tones, has a characteristic shape:

  • At zero separation, no roughness — a unison is smooth.
  • Roughness rises steeply as separation increases, peaking at roughly a quarter of a critical bandwidth.
  • It then falls away, reaching negligible values at about one full critical bandwidth.

Critical bandwidth is itself frequency-dependent: about 100 hertz below 500 hertz, and roughly a fixed fraction — near a minor third — above that. So the same interval is a different fraction of a critical band in different registers, and is correspondingly rougher or smoother.

For two complex tones, the total roughness is the sum over every pair of partials, weighted by their amplitudes. That is the entire computation, and it has exactly two inputs: a partial list and the Plomp–Levelt curve.

Four spectra of the same noteThe amplitude of each partial for four timbres at the same pitch. These are the exact lists the sound buttons on this site synthesise from, so the picture and the sound are the same data.1pureone partial, nothing else12345678stringall partials, falling12345678clarineteven partials nearly absent123456789bellodd partials onlyamplitude
Fig. 2 The partial lists that go into the calculation. These are the same lists the sound buttons on this site synthesise from, so the picture, the computation and the noise are one object.

Why the dips land where they do

The reason the minima fall on simple ratios is almost embarrassingly direct.

Consider two tones a fifth apart, at frequencies ff and 32f\tfrac{3}{2}f. The lower tone’s partials are at f,2f,3f,4f,5f,6ff, 2f, 3f, 4f, 5f, 6f. The upper tone’s are at 1.5f,3f,4.5f,6f,7.5f,9f1.5f, 3f, 4.5f, 6f, 7.5f, 9f. Two of the upper tone’s first four partials — 3f3f and 6f6f — coincide exactly with partials of the lower tone.

Coinciding partials do not beat. They contribute zero roughness. And the remaining partials are far enough apart to be resolved separately, which also contributes nothing.

The first eight partials of a stringA string vibrating in one, two, three and more equal parts, with the frequency ratio and the nearest named note beside each. The seventh partial is a third of a semitone flat of anything on a keyboard, which is a fact about strings rather than about tuning.1130.8 HzC2261.6 HzC3392.4 HzG4523.3 HzC5654.1 HzE -14¢6784.9 HzG7915.7 HzB♭ -31¢81046.5 HzCpartialthe dots are the nodes — the places that do not move
Fig. 3 The partials of one note. When two notes sound together, two stacks like this one interleave, and whether their members land on each other or merely near each other is the whole of the calculation.

Now consider a tritone at roughly 1.414. The upper tone’s partials fall at 1.414, 2.83, 4.24, 5.66 — every one of them close to a partial of the lower tone without landing on it. Each of those near-misses is inside a critical band, each contributes roughness, and they add.

So the rule “simpler ratios are more consonant” is not a fundamental principle. It is a consequence of the fact that simpler ratios cause more partials to coincide, given partials at whole-number multiples.

The prediction that makes it a theory

An explanation that only accounts for what is already known is not worth much. This one makes a prediction, and the prediction is falsifiable and has been tested.

If the partials move, the consonant intervals move with them.

Feed the same computation a tone whose partials sit at 1, 2.1, 3.3, 4.7 times the fundamental instead of 1, 2, 3, 4, and the curve comes out with dips in entirely different places. Some of the classical intervals lose their minima; new minima appear at ratios that have no names.

This has been done both computationally and with synthesised sound since the 1970s, and it works. Intervals that are dissonant for harmonic tones are consonant for the stretched ones, and listeners agree — the effect does not require training to hear.

There is also a natural experiment. Indonesian gamelan ensembles are built from metallophones and gongs, whose partials are genuinely inharmonic; and gamelan tuning systems, sléndro and pélog, divide the octave in ways that no harmonic-series theory predicts and that fit those instruments’ actual spectra well. The instruments and the scale co-evolved, and the roughness model gets the relationship roughly right where the ratio theory gets it entirely wrong. The same reasoning explains why a bell cannot be tuned by beats.

What the curve gets right about real music

Several features of ordinary practice fall straight out of the curve, and are usually taught as rules with no reason attached.

The semitone is the worst interval, and it is used as a dissonance. The peak of the curve is at roughly a semitone, and every tradition that has a semitone treats it as maximally unstable.

Wide spacing in the bass. Critical bandwidth is proportionally much wider at low frequencies, so a third in the bass is inside one critical band and is genuinely rough, while the same third two octaves up is not. Every orchestration text says to space low chords widely; the reason is the width of a critical band, and the reason is almost never given.

The fourth’s ambiguity. The fourth has a deep roughness minimum — as deep as the fifth’s — and Western theory has treated it as a dissonance requiring resolution in certain contexts for six hundred years. That is a straightforward disagreement between the sensory account and the practice, and it is the clearest evidence that roughness is not the whole story — the fourth’s treatment depends entirely on which note is in the bass.

Four triads as stacked intervalsEach triad drawn as the semitone distances above its root, with the nearest simple frequency ratio beside each note. The chords differ only in the size of the two stacked thirds, and that difference is the whole of their character.C1/1E5/4G3/243major0–4–7C1/1E♭6/5G3/234minor0–3–7C1/1E♭6/5F♯33diminished0–3–6C1/1E5/4A♭8/544augmented0–4–8semitones above the root
Fig. 4 Four triads as stacked intervals. The major and minor triads sit in low-roughness territory; the diminished and augmented ones do not, and both are used as unstable sonorities that go somewhere.

Reading the curve carefully

The shape rewards more attention than a glance gives it, and two of its features are easy to miss.

The wells are not equally deep. The octave’s minimum goes essentially to zero; the fifth’s is deep; the thirds’ are shallow depressions on a slope rather than wells. That ordering matches the historical order in which intervals were admitted as consonances — octave first, then fifth and fourth, then thirds only in the fifteenth century — which is a suspicious coincidence and probably not one.

The thirds are not minima at all, quite. With eight string-like partials, the major third sits on a shoulder rather than in a dip. That is a real feature of the computation and not an artefact, and it is the single most useful thing the figure shows: the consonance of the third is marginal on a purely sensory account, which is exactly why it took two thousand years to be accepted and why temperaments argued about it so bitterly.

Change the partial list — fewer partials, or weaker upper ones — and the thirds’ shoulders deepen into dips. A harpsichord, whose upper partials decay fast, gives a different curve from a modern piano. Some of the historical disagreement about whether thirds are consonant was a disagreement between instruments.

What a single partial does

The cleanest way to see the mechanism is to remove almost all of it.

Two pure sine tones have one partial each. There is exactly one pair to consider, and the roughness curve for sine tones is just the Plomp–Levelt function itself: a single hump near the unison and flat thereafter. No dip at the fifth, no dip at the octave, nothing.

The same partials, drawn as pressureEach spectrum summed into the wave it actually produces, over two cycles. The shapes are strikingly different and the ear has almost no access to that difference — what it hears is the list of partials, not the shape they add up to.pure1 partialstring8 partialsclarinet8 partials2 cycles · each normalised by its own peak
Fig. 5 A pure tone against two complex ones. The pure tone has one partial and nothing to interfere with anything; the others have eight, and every pair of those eight is a term in the roughness sum.

And that is what sine tones sound like: intervals between them are strangely characterless, and listeners find them hard to identify. A sine-tone fifth is not obviously a fifth. Musicians consistently report this as disconcerting the first time, and it is exactly what the model predicts — with one partial each there is nothing to coincide, so the consonance of the interval has nothing to be made of.

This is the strongest single confirmation available, because it is a case where the ratio theory and the roughness theory make opposite predictions. The ratio 3:2 is just as simple for sine tones as for strings. The consonance is not.

Beating is the mechanism underneath

Roughness is not a primitive. It is beating, at a rate too fast to count.

Two partials one hertz apart produce an audible slow pulsing. Ten hertz apart, the pulsing is too fast to follow and becomes a buzz. Thirty hertz apart, the buzz fades and the two components begin to separate into distinct tones. The Plomp–Levelt curve is a summary of exactly that progression, expressed as a fraction of critical bandwidth rather than in hertz so that it applies at any register.

220 Hz against 223 HzTwo tones 3 hertz apart, added. The rapid oscillation is their average; the slow swelling is their difference, heard as 3 beats a second and used by every tuner who has ever worked by ear.00.511.52seconds3 beats per second — the difference, exactlythe carrier is drawn slower than it sounds, or it would be a solid band
Fig. 6 Two tones a few hertz apart, and their sum. Speed the difference up by a factor of ten and the swelling stops being countable and starts being a texture. That texture, summed over every pair of partials, is what the dissonance curve plots.

So the chain runs: two frequencies close together, unresolved by the cochlea, produce fluctuation; fluctuation at a few hertz is heard as pulsing and at a few tens of hertz as roughness; roughness summed over all partial pairs is sensory dissonance; and sensory dissonance is minimised at ratios that make partials coincide. Each link is a measured relationship, and none of them mentions music.

Where the number of partials matters

One parameter changes the picture more than any other, and it is not the interval.

An instrument with two partials produces a nearly flat roughness curve with one shallow octave dip. An instrument with twelve produces a curve with pronounced dips at every simple ratio down to 7:4. The more partials a tone has, the more sharply defined its consonances are — and the more of them there are.

Four spectra of the same noteThe amplitude of each partial for four timbres at the same pitch. These are the exact lists the sound buttons on this site synthesise from, so the picture and the sound are the same data.1pureone partial, nothing else12345678stringall partials, falling12345678clarineteven partials nearly absent123456789bellodd partials onlyamplitude
Fig. 7 Four partial lists, from one partial to nine. The number of partials is the parameter that decides how much structure the roughness curve has, which means that how strongly a timbre distinguishes consonance from dissonance is a property of the instrument.

That has a corollary worth sitting with. A bright instrument has more consonance structure than a dull one. A harpsichord, a bowed string or a reed makes the difference between a fifth and a tritone acute; a flute or a stopped organ pipe, whose spectra are nearly pure, makes it mild.

Composers know this without the theory: dense harmony is written for strings and reeds and thinned for flutes, and close dissonant spacing is tolerable in soft flute writing that would be intolerable on oboes. The rule of thumb precedes the explanation by about three hundred years.

What the curve gets wrong

The fourth is not an isolated failure. There is a whole category of musical fact that roughness cannot reach, and being clear about the boundary is more useful than defending the model past it.

Dissonance is partly expectation. The dominant seventh chord has substantial roughness and was, in eighteenth-century practice, a dissonance requiring resolution. By 1920 it was a stable chord to end a blues on. Nothing about its spectrum changed. What changed was what a listener expected to happen next, which is a fact about a repertoire and not about a cochlea, and which belongs to how progressions work rather than to acoustics.

Consonance is partly harmonicity. Two tones whose frequencies suggest a common fundamental sound fused, and that appears to be a separate cue from roughness. An octave and a twelfth are fused even when spaced widely enough that no partials are near enough to beat at all. Roughness predicts nothing there; harmonicity does.

Context reverses judgements. A chord that is rough in isolation can be the resolution of something rougher, and it will be heard as restful. Roughness is computed pairwise on a static sound; music is neither pairwise nor static.

The honest summary is that the roughness model explains sensory dissonance — the immediate grating quality of a sound — completely and quantitatively, and explains musical dissonance only partially. The two words are usually used interchangeably, and the confusion causes most of the arguments.

The instrument chooses the scale

The strongest consequence of the model is a reversal of the usual causal story, and it is worth stating baldly.

The usual story: musical scales derive from the harmonic series, which is a fact of physics, so scales are grounded in nature.

The model’s story: musical scales derive from the spectra of the available instruments. For strings and pipes, those spectra are harmonic, so the scales that result are built on small integer ratios. For inharmonic instruments, different scales result, and they are equally well grounded.

The major scale as a cycleThe twelve semitones drawn as a cycle, with the notes of the scale filled in. The gaps between filled positions are the step pattern, and reading them round the circle is what makes the scale's asymmetry obvious.CC♯DE♭EFF♯GA♭AB♭B2 · 2 · 1 · 2 · 2 · 2 · 1steps, in semitonesthe short chords are the semitones — two of them, unevenly spaced
Fig. 8 A scale drawn as a selection from twelve equal positions round a cycle. That there are twelve positions to select from is a consequence of the harmonic series in a specific way — through the chain of fifths — and a tradition with other instruments would have a different necklace with a different number of beads. Which of the twelve get chosen is a separate question again.

That is a much weaker claim than the usual one and a much more defensible one. It explains why so many traditions converged on the octave and the fifth — every string and every pipe produces them — and it explains why traditions with metallophones did not converge on the same thirds.

It also predicts something that can be checked: a culture whose principal instruments are inharmonic should have a scale that the harmonic series does not explain. Gamelan is that culture, and it does.

Whose music, and when

The mathematics is culture-free and the interpretation is not, and it is easy to overreach here.

Plomp and Levelt’s listening tests were conducted with Dutch subjects in the 1960s, and later work has repeated the sine-tone roughness measurements across populations with broadly consistent results — the critical-band mechanism is anatomical, and anatomy does not vary much. That part travels.

What does not travel is the preference. Roughness is a measurable property; whether roughness is wanted is an aesthetic decision, and traditions differ sharply. Bulgarian women’s choral singing uses sustained seconds as a central sonority. Balinese gamelan tunes paired instruments a few hertz apart to produce deliberate beating. Both are choices to use roughness rather than to minimise it, and describing either as out of tune is a category error.

The model says how rough a sound is. It does not say that less is better, and every claim that it does has smuggled in a preference.

Where the model stops

Pairwise only. Roughness is computed between pairs of partials and summed. Three or more simultaneous tones may interact in ways the pairwise sum misses, and the model’s extension to full chords is an approximation with known weaknesses.

Static. The computation assumes steady tones. Real notes have attacks and decays, and a rough sonority that lasts a tenth of a second is not heard as rough at all.

No level axis. Roughness grows with loudness and the model as usually stated does not include it. Every curve on this page is for one unspecified level.

One spectrum drawn. The hero figure uses a single string-like partial list. The essay’s central claim is that a different list gives a different curve, and only one curve appears — a real limitation of the picture, and the reason the argument leans on the gamelan case rather than on the drawing.

The model is old. Plomp and Levelt’s 1965 formulation has been refined many times since, and modern versions use auditory filterbanks rather than a single critical-band curve. The refinements move the details and not the conclusions.

The ladder from here

Later rungs: critical bandwidth, measured and modelled. The Plomp–Levelt curve in full derivation. Harmonicity as an independent cue. Roughness in chords rather than intervals. Stretched and compressed spectra, with the synthesis to hear them. Gamelan tuning and its instruments. Sethares’s method for deriving a scale from a spectrum. Dissonance as expectation, and the history of the dominant seventh. Bulgarian seconds, and roughness used deliberately. And amplitude, which every figure here has quietly held constant.

The most striking thing about the curve is how little goes into it. Two spectra and one psychophysical function, measured in a laboratory with sine tones by people who were not asking a musical question, and out comes a picture with a fifth and an octave in it.