Intervals and chords

Roughness can be computed, and the answer looks like a scale

Feed two spectra into a model of how nearby frequencies interfere, sweep one past the other, and the curve that comes out has dips exactly where the consonant intervals are. Nobody put them there.

Assumes: Two notes and a ratio, which is the whole of consonance · Beats are arithmetic that anybody can hear

Here is an experiment that can be run without any music in it at all.

Take a model of how two nearby frequencies interfere in the ear. Take two complex tones with realistic partials. Fix the lower one, sweep the upper one slowly up through an octave, and at every point sum the interference between every pair of partials. Plot the total.

Roughness across an octave. Sensory dissonance between two complex tones as the upper one is swept through an octave, computed by summing the roughness between every pair of their partials. Nothing here is placed by hand. The wells this spectrum produces sit on 4/3, on 3/2, on 5/3, found by scanning the curve rather than by marking them. And the tempered minor third sits 198 cents below the nearest well, and the tempered major third sits 98 cents below the nearest well, which is a real feature of this model and not a defect of the drawing.
Fig. 1 The result. Roughness between two complex tones as the upper one is swept through an octave, computed by summing the interference between every pair of their partials. Nothing in this curve is placed by hand — the vertical marks are added afterwards, at the simple ratios, to show where the features landed.

The curve has structure. It rises steeply from the unison, peaks somewhere around a semitone, and then descends through a series of dips. The dips are at the fourth, the fifth and the octave, with shallower ones at the thirds and the sixth.

Nobody told it to put them there. The input was a partial list and an interference model; the output is something that looks a great deal like a scale.

The model, in full

The mechanism is worth stating precisely, because the whole argument depends on it being modest.

Two pure tones close in frequency are not resolved separately by the ear. They excite overlapping regions of the basilar membrane, and the result is a fluctuating response that is perceived as roughness. How rough depends on how far apart they are relative to the critical bandwidth at that frequency.

The relationship, established by Plomp and Levelt in 1965 from listening tests with pairs of sine tones, has a characteristic shape:

  • At zero separation, no roughness — a unison is smooth.
  • Roughness rises steeply as separation increases, peaking at roughly a quarter of a critical bandwidth.
  • It then falls away, reaching negligible values at about one full critical bandwidth.

Critical bandwidth is itself frequency-dependent — which is why a third is rougher in the bass than the same third an octave up: about 100 hertz below 500 hertz, and roughly a fixed fraction — near a minor third — above that. So the same interval is a different fraction of a critical band in different registers, and is correspondingly rougher or smoother.

For two complex tones, the total roughness is the sum over every pair of partials, weighted by their amplitudes. That is the entire computation, and it has exactly two inputs: a partial list and the Plomp–Levelt curve.

Four spectra of the same note. The amplitude of each partial for 4 timbres at the same pitch — pure, string, clarinet, bell. These are the exact lists the sound buttons here synthesise from, so the picture and the sound are the same data.
Fig. 2 The partial lists that go into the calculation. These are the same lists the sound buttons here synthesise from, so the picture, the computation and the noise are one object.

Why the dips land where they do

The reason the minima fall on simple ratios is almost embarrassingly direct.

Consider two tones a fifth apart, at frequencies ff and 32f\tfrac{3}{2}f. The lower tone’s partials are at f,2f,3f,4f,5f,6ff, 2f, 3f, 4f, 5f, 6f. The upper tone’s are at 1.5f,3f,4.5f,6f,7.5f,9f1.5f, 3f, 4.5f, 6f, 7.5f, 9f. Two of the upper tone’s first four partials — 3f3f and 6f6f — coincide exactly with partials of the lower tone.

Coinciding partials do not beat. They contribute zero roughness, and what makes two partials one note is the same coincidence read from the other side. And the remaining partials are far enough apart to be resolved separately, which also contributes nothing.

Two spectra of the same note. The amplitude of each partial for 2 timbres at the same pitch — pure, string. These are the exact lists the sound buttons here synthesise from, so the picture and the sound are the same data.
Fig. 3 The two stacks the calculation is over. A pure tone has one line; a string has eight or more, at whole-number multiples of its fundamental. When two notes sound together two stacks like these interleave, and whether their members land on each other or merely near each other is the whole of the arithmetic. Two tones a fifth apart put the upper tone’s second and fourth partials exactly on the lower tone’s third and sixth; a tritone puts every one of the upper tone’s partials close to a partial of the lower one without landing on it, and each near-miss is a pair inside a critical band contributing roughness.

Now consider a tritone at roughly 1.414. The upper tone’s partials fall at 1.414, 2.83, 4.24, 5.66 — every one of them close to a partial of the lower tone without landing on it. Each of those near-misses is inside a critical band, each contributes roughness, and they add.

So the rule “simpler ratios are more consonant” is not a fundamental principle. It is a consequence of the fact that simpler ratios cause more partials to coincide, given partials at whole-number multiples.

The prediction that makes it a theory

An explanation that only accounts for what is already known is not worth much. This one makes a prediction, and the prediction is falsifiable and has been tested.

If the partials move, the consonant intervals move with them.

Feed the same computation a tone whose partials sit at 1, 2.1, 3.3, 4.7 times the fundamental instead of 1, 2, 3, 4, and the curve comes out with dips in entirely different places. Some of the classical intervals lose their minima; new minima appear at ratios that have no names.

This has been done both computationally and with synthesised sound since the 1970s, and it works — a spectrum chooses its own scale is the general statement of it. Intervals that are dissonant for harmonic tones are consonant for the stretched ones, and listeners agree — the effect does not require training to hear.

There is also a natural experiment. Indonesian gamelan ensembles are built from metallophones and gongs, whose partials are genuinely inharmonic; and gamelan tuning systems, sléndro and pélog, divide the octave in ways that no harmonic-series theory predicts and that fit those instruments’ actual spectra well. The instruments and the scale co-evolved, and the roughness model gets the relationship roughly right where the ratio theory gets it entirely wrong. The same reasoning explains why a bell cannot be tuned by beats.

What the curve gets right about real music

Several features of ordinary practice fall straight out of the curve, and are usually taught as rules with no reason attached.

The semitone is the worst interval, and it is used as a dissonance. The peak of the curve is at roughly a semitone, and every tradition that has a semitone treats it as maximally unstable.

Wide spacing in the bass. Critical bandwidth is proportionally much wider at low frequencies, so a third in the bass is inside one critical band and is genuinely rough, while the same third two octaves up is not. Every orchestration text says to space low chords widely; the reason is the width of a critical band, and the reason is almost never given.

The fourth’s ambiguity. The fourth has a deep roughness minimum — as deep as the fifth’s — and Western theory has treated it as a dissonance requiring resolution in certain contexts for six hundred years. That is a straightforward disagreement between the sensory account and the practice, and it is the clearest evidence that roughness is not the whole story — the fourth’s treatment depends entirely on which note is in the bass.

Four triads as stacked intervals: the major and minor sit in the smooth region, the diminished and augmented do not — which is the ordering the curve produces and the ordering four centuries of practice already had.

Reading the curve carefully

The shape rewards more attention than a glance gives it, and two of its features are easy to miss.

The wells are not equally deep. The octave’s minimum goes essentially to zero; the fifth’s is deep; the thirds’ are shallow depressions on a slope rather than wells. That ordering matches the historical order in which intervals were admitted as consonances — octave first, then fifth and fourth, then thirds only in the fifteenth century — which is a suspicious coincidence and probably not one.

The thirds are not minima at all, quite. With eight string-like partials, the major third sits on a shoulder rather than in a dip. That is a real feature of the computation and not an artefact, and it is the single most useful thing the figure shows: the consonance of the third is marginal on a purely sensory account, which is exactly why it took two thousand years to be accepted and why temperaments argued about it so bitterly.

The obvious next thought is that a different instrument would rescue the third: change the partial list, and the shoulder deepens into a dip. That is worth running rather than asserting, and this site’s own model does not agree with the obvious form of it. Truncating the string spectrum to six partials, to five, to four, and steepening the roll-off to one over n squared, all leave the third exactly where it was — a shoulder — and take wells away rather than adding them. Sixteen partials and twenty do the same. The claim that fewer or weaker upper partials deepens the third is simply false here.

What does produce a well on the third is a strong fifth partial, which is a different thing and a sharper one. The well at 5:4 is built by the coincidence of the lower tone’s fifth partial with the upper tone’s fourth, and in a one-over-n spectrum that fifth partial is a fifth of the fundamental’s amplitude, which is not enough to dig a well with. Raise it to nine tenths and both thirds appear at once: a minimum on the just major third at 386 cents with a prominence of 5.5% of the curve’s range, and one on the just minor third at 316 cents at 2.2%. So the historical disagreement about whether thirds are consonant may well have been a disagreement between instruments — but the property that decides it is where the energy sits in the spectrum, not how much of it there is.

What a single partial does

The cleanest way to see the mechanism is to remove almost all of it.

Two pure sine tones have one partial each. There is exactly one pair to consider, and the roughness curve for sine tones is just the Plomp–Levelt function itself: a single hump near the unison and flat thereafter. No dip at the fifth, no dip at the octave, nothing.

Three spectra of the same note. The amplitude of each partial for 3 timbres at the same pitch — pure, string, clarinet. These are the exact lists the sound buttons here synthesise from, so the picture and the sound are the same data.
Fig. 4 A pure tone against two complex ones, which is where a single partial’s contribution is legible. The pure tone has one partial and can only be rough with the fundamental of whatever it is played against, so its curve is nearly flat with one shallow octave dip. A clarinet’s even partials are almost absent, so half the pairs that would have coincided are missing and the fourth’s dip goes with them. The more partials a tone has, the more sharply defined its consonances are — and the more of them there are.

And that is what sine tones sound like: intervals between them are strangely characterless, and listeners find them hard to identify. A sine-tone fifth is not obviously a fifth. Musicians consistently report this as disconcerting the first time, and it is exactly what the model predicts — with one partial each there is nothing to coincide, so the consonance of the interval has nothing to be made of.

This is the strongest single confirmation available, because it is a case where the ratio theory and the roughness theory make opposite predictions. The ratio 3:2 is just as simple for sine tones as for strings. The consonance is not.

Beating is the mechanism underneath

Roughness is not a primitive. It is beating, at a rate too fast to count.

Two partials one hertz apart produce an audible slow pulsing. Ten hertz apart, the pulsing is too fast to follow and becomes a buzz. Thirty hertz apart, the buzz fades and the two components begin to separate into distinct tones. The Plomp–Levelt curve is a summary of exactly that progression, expressed as a fraction of critical bandwidth rather than in hertz so that it applies at any register.

And beating is the mechanism underneath: two tones a few hertz apart swell and fade at their difference, and speeding that difference up past about twenty a second turns the swelling into a texture rather than a count. Roughness is that texture, summed over every pair of partials close enough to produce it.

So the chain runs: two frequencies close together, unresolved by the cochlea, produce fluctuation; fluctuation at a few hertz is heard as pulsing and at a few tens of hertz as roughness; roughness summed over all partial pairs is sensory dissonance; and sensory dissonance is minimised at ratios that make partials coincide. Each link is a measured relationship, and none of them mentions music.

Where the number of partials matters

One parameter changes the picture more than any other, and it is not the interval.

An instrument with two partials produces a nearly flat roughness curve with one shallow octave dip. An instrument with twelve produces a curve with pronounced dips at every simple ratio down to 7:4. The more partials a tone has, the more sharply defined its consonances are — and the more of them there are.

Two spectra of the same note. The amplitude of each partial for 2 timbres at the same pitch — pure, clarinet. These are the exact lists the sound buttons here synthesise from, so the picture and the sound are the same data.
Fig. 5 Four partial lists, from one partial to nine. The number of partials is the parameter that decides how much structure the roughness curve has, which means that how strongly a timbre distinguishes consonance from dissonance is a property of the instrument.

That has a corollary worth sitting with. A bright instrument has more consonance structure than a dull one. A harpsichord, a bowed string or a reed makes the difference between a fifth and a tritone acute; a flute or a stopped organ pipe, whose spectra are nearly pure, makes it mild.

Composers know this without the theory: dense harmony is written for strings and reeds and thinned for flutes, and close dissonant spacing is tolerable in soft flute writing that would be intolerable on oboes. The rule of thumb precedes the explanation by about three hundred years.

What the curve gets wrong

The fourth is not an isolated failure. There is a whole category of musical fact that roughness cannot reach, and being clear about the boundary is more useful than defending the model past it.

Dissonance is partly expectation. The dominant seventh chord has substantial roughness and was, in eighteenth-century practice, a dissonance requiring resolution. By 1920 it was a stable chord to end a blues on. Nothing about its spectrum changed. What changed was what a listener expected to happen next, which is a fact about a repertoire and not about a cochlea, and which belongs to how progressions work rather than to acoustics.

Consonance is partly harmonicity. Two tones whose frequencies suggest a common fundamental sound fused, and that appears to be a separate cue from roughness. An octave and a twelfth are fused even when spaced widely enough that no partials are near enough to beat at all. Roughness predicts nothing there; harmonicity does.

Context reverses judgements. A chord that is rough in isolation can be the resolution of something rougher, and it will be heard as restful. Roughness is computed pairwise on a static sound; music is neither pairwise nor static.

The honest summary is that the roughness model explains sensory dissonance — the immediate grating quality of a sound — completely and quantitatively, and explains musical dissonance only partially. The two words are usually used interchangeably, and the confusion causes most of the arguments.

The instrument chooses the scale

The strongest consequence of the model is a reversal of the usual causal story, and it is worth stating baldly.

The usual story: musical scales derive from the harmonic series, which is a fact of physics, so scales are grounded in nature.

The model’s story: musical scales derive from the spectra of the available instruments. For strings and pipes, those spectra are harmonic, so the scales that result are built on small integer ratios. For inharmonic instruments, different scales result, and they are equally well grounded.

Three spectra of the same note. The amplitude of each partial for 3 timbres at the same pitch — bell, string, pure. These are the exact lists the sound buttons here synthesise from, so the picture and the sound are the same data.
Fig. 6 The reversal, in one comparison. A string’s partials are whole-number multiples, so the intervals at which two of them coincide are the small-integer ratios, and a tradition with strings and pipes arrives at a scale built on those. A bell’s are multiples of nothing, so there is no ratio at which two bells’ partials line up and no reason for such a tradition to arrive at the same intervals. The usual story is that scales derive from the harmonic series and are therefore grounded in nature; the model’s story is that scales derive from the spectra of the available instruments, which is a weaker claim and a much more defensible one — and it predicts that a culture whose principal instruments are inharmonic will have a scale the harmonic series does not explain.

That is a much weaker claim than the usual one and a much more defensible one. It explains why so many traditions converged on the octave and the fifth — every string and every pipe produces them — and it explains why traditions with metallophones did not converge on the same thirds.

It also predicts something that can be checked: a culture whose principal instruments are inharmonic should have a scale that the harmonic series does not explain. Gamelan is that culture, and it does.

Whose music, and when

The mathematics is culture-free and the interpretation is not, and it is easy to overreach here.

Plomp and Levelt’s listening tests were conducted with Dutch subjects in the 1960s, and later work has repeated the sine-tone roughness measurements across populations with broadly consistent results — the critical-band mechanism is anatomical, and anatomy does not vary much. That part travels.

What does not travel is the preference. Roughness is a measurable property; whether roughness is wanted is an aesthetic decision, and traditions differ sharply. Bulgarian women’s choral singing uses sustained seconds as a central sonority. Balinese gamelan tunes paired instruments a few hertz apart to produce deliberate beating. Both are choices to use roughness rather than to minimise it, and describing either as out of tune is a category error.

The model says how rough a sound is. It does not say that less is better, and every claim that it does has smuggled in a preference.

Where the model stops

Pairwise only. Roughness is computed between pairs of partials and summed. Three or more simultaneous tones may interact in ways the pairwise sum misses, and the model’s extension to full chords is an approximation with known weaknesses.

Static. The computation assumes steady tones. Real notes have attacks and decays, and a rough sonority that lasts a tenth of a second is not heard as rough at all.

No level axis. Roughness grows with loudness and the model as usually stated does not include it. Every curve on this page is for one unspecified level.

One spectrum drawn. The hero figure uses a single string-like partial list. The essay’s central claim is that a different list gives a different curve, and only one curve appears — a real limitation of the picture, and the reason the argument leans on the gamelan case rather than on the drawing.

The model is old. Plomp and Levelt’s 1965 formulation has been refined many times since, and modern versions use auditory filterbanks rather than a single critical-band curve. The refinements move the details and not the conclusions.

The ladder from here

Later rungs: critical bandwidth, measured and modelled. The Plomp–Levelt curve in full derivation. Harmonicity as an independent cue. Roughness in chords rather than intervals. Stretched and compressed spectra, with the synthesis to hear them. Gamelan tuning and its instruments. Sethares’s method for deriving a scale from a spectrum. Dissonance as expectation, and the history of the dominant seventh. Bulgarian seconds, and roughness used deliberately. And amplitude, which every figure here has quietly held constant.

The most striking thing about the curve is how little goes into it. Two spectra and one psychophysical function, measured in a laboratory with sine tones by people who were not asking a musical question, and out comes a picture with a fifth and an octave in it.

Part 2 of 10

One essay in the series on consonance. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 61.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Critical bandwidthInharmonicityPlomp–Levelt curveRoughnessSensory dissonance