Intervals and chords

Two notes and a ratio, which is the whole of consonance

Sound two tones together and the pair either settles or does not. What decides it is the ratio of their frequencies, and the rule is that simpler ratios settle — which is two and a half thousand years old and still not quite an explanation.

Sound two tones at once and one of two things happens. Either they fuse into something that sounds like a single object with a colour, or they sit side by side and grate. The distinction is immediate, involuntary and shared across listeners who have never met.

What decides it is the ratio between their frequencies.

Two tones in the ratio 3 to 2Two sine tones whose frequencies are in the ratio 3 to 2, and their sum. The combined pattern repeats every time both waves return to the start together, which for a simple ratio is soon.one frame = one repeat of the combined wavelower — 2 per frameupper — 3 per frametheir sumthe pattern repeats after 2 cycles of the lower tone
Fig. 1 Two sine tones whose frequencies stand in the ratio three to two, and their sum below. The combined pattern repeats every time both waves return to their starting phase together, which for a simple ratio happens soon and for a complex one happens late.

Frequencies in the ratio 2:1 give an octave, which fuses so completely that most musical cultures treat the two notes as versions of the same note. 3:2 gives a fifth, 4:3 a fourth, 5:4 a major third. Keep going and the ratios get more complicated as the intervals get less settled, until 45:32 gives a tritone, which does not settle at all.

It is also the reason the tuning systems exist: a system is an attempt to make as many ratios as possible simple at once. This is the oldest quantitative result in music, and probably in acoustics. It is attributed to Pythagoras, it is demonstrable on a stretched string with nothing but a moveable bridge, and it is correct.

It is also not, on its own, an explanation.

The classical argument, and its gap

The ancient reasoning runs: simple ratios produce a combined waveform that repeats quickly, and the ear prefers patterns it can grasp. The figure above draws exactly that. For a 3:2 pair the whole pattern completes in two cycles of the lower tone. For a 45:32 pair it takes forty-five.

The argument is appealing, and it survives in textbooks. It has a fatal problem: the ear is very nearly deaf to the shape of a waveform.

The same partials, drawn as pressureEach spectrum summed into the wave it actually produces, over two cycles. The shapes are strikingly different and the ear has almost no access to that difference — what it hears is the list of partials, not the shape they add up to.pure1 partialstring8 partialsclarinet8 partials2 cycles · each normalised by its own peak
Fig. 2 Three spectra summed into the waves they actually produce. The shapes are strikingly different, and the difference between them is largely inaudible — what the ear hears is the list of partials each contains, not the shape they add up to.

Take a complex tone and shift the phase of its partials. The waveform changes completely; a shape that looked like a sawtooth becomes something unrecognisable. The sound is essentially unchanged. This was established by Ohm in 1843 and confirmed exhaustively since, and it is one of the foundational facts of hearing — the same fact that makes a spectrum a better description of a sound than a waveform: the ear performs something close to a frequency analysis and discards most phase information.

If the ear cannot hear waveform shape, then “the combined waveform repeats quickly” cannot be why a fifth is consonant. The periodicity is real, the consonance is real, and the causal link asserted between them does not hold.

What actually happens

The working explanation is different, and it is about the partials rather than about the fundamentals.

The first eight partials of a stringA string vibrating in one, two, three and more equal parts, with the frequency ratio and the nearest named note beside each. The seventh partial is a third of a semitone flat of anything on a keyboard, which is a fact about strings rather than about tuning.1130.8 HzC2261.6 HzC3392.4 HzG4523.3 HzC5654.1 HzE -14¢6784.9 HzG7915.7 HzB♭ -31¢81046.5 HzCpartialthe dots are the nodes — the places that do not move
Fig. 3 A string vibrating in one, two, three and more parts simultaneously, with the frequency of each partial. Every note from a string, a pipe or a voice is a stack like this, and what sounds together when two notes sound together is two of these stacks.

Every real musical tone is a stack of partials at whole-number multiples of its fundamental. When two such tones sound together, what actually reaches the ear is two stacks — and the interesting question is how the partials of one relate to the partials of the other.

For an octave, every partial of the upper tone lands exactly on a partial of the lower. Nothing new is introduced at all; the upper tone reinforces partials the lower one already had, which is why the octave fuses so completely that it is barely heard as two notes.

For a fifth at 3:2, the upper tone’s partials fall on every third partial of the lower. Two-thirds of them coincide exactly and the rest sit in gaps.

For a tritone at 45:32, almost nothing coincides. Partials land near each other without landing on each other — and a pair of partials that are near but not equal beats, and beating at the wrong rate is roughness.

Roughness across an octaveSensory dissonance between two complex tones as the upper one is swept through an octave, computed by summing the roughness between every pair of their partials. Nothing here is placed by hand. The deep wells land on the fourth, the fifth and the octave; the thirds sit on shoulders rather than in wells, which is a real feature of this model and not a defect of the drawing.semitones above the lower tone6/55/44/33/25/32/1CC♯DE♭EFF♯GA♭AB♭BCroughest at about a semitonesmooth at the simple ratios
Fig. 4 Sensory dissonance across an octave, computed by summing the roughness between every pair of partials of two complex tones as the upper one is swept upward. Nothing here is placed by hand. The deep wells fall on the simple ratios because that is where the partials coincide.

That curve is the modern answer, and its shape was not assumed: it comes out of a model of how two nearby frequencies interact in the ear, applied to every pair of partials and summed. The simple ratios come out as minima because coinciding partials do not beat.

The circularity that is not one

There is an objection that has to be met, because it is a good one. If consonance is explained by partials coinciding, and partials are at whole-number multiples of the fundamental, has anything been explained? The whole numbers went in and the whole numbers came out.

The answer is that the model makes a prediction the classical account cannot: change the partials and the consonant intervals move.

Four spectra of the same noteThe amplitude of each partial for four timbres at the same pitch. These are the exact lists the sound buttons on this site synthesise from, so the picture and the sound are the same data.1pureone partial, nothing else12345678stringall partials, falling12345678clarineteven partials nearly absent123456789bellodd partials onlyamplitude
Fig. 5 Four spectra of the same pitch. The partial lists differ enormously, and the roughness between any two tones is computed from these lists — so an instrument with an unusual spectrum has an unusual set of consonant intervals.

Take a tone whose partials are stretched — not at 1, 2, 3, 4 times the fundamental but at 1, 2.1, 3.3, 4.6 — and compute the dissonance curve again. The minima move. They no longer fall at 3:2 and 5:4; they fall wherever the stretched partials happen to coincide, and an interval that is consonant for that timbre is not one anybody would name.

This has been done, both computationally and with synthesised sound, and the effect is real and immediately audible. It also has a natural experiment attached: the Indonesian gamelan is built around metallophones and gongs, whose partials are genuinely inharmonic, and gamelan tuning systems — sléndro and pélog — divide the octave in ways that no theory built on the harmonic series predicts and that fit the instruments’ actual spectra well.

So consonance is not a fact about small integers. It is a fact about spectra, which for strings and pipes happen to be built on small integers.

The interval is the thing, not the notes

One consequence of consonance being a matter of ratios is easy to state and surprisingly hard to internalise: an interval is a relationship, and it is unchanged by moving both notes.

C to G and F-sharp to C-sharp are both 3:2. They sound like the same interval, they behave the same way in a harmony, and a listener asked to compare them will say they are identical — even though not one frequency is shared between them. The same is true of a melody: transpose it and it remains recognisably the same melody, because what a listener retained was the sequence of ratios rather than the sequence of pitches.

Notes on the keyboardA piano keyboard with the notes under discussion marked. The keyboard is used throughout this site because it shows distance rather than name, and distance is what the theory is about.CEG3 notes sounding
Fig. 6 Two notes marked on a keyboard. The keyboard is used throughout this site in preference to a stave because it shows distance rather than name, and distance — measured in ratios, drawn in semitones — is what every argument here is actually about.

This is why the keyboard, rather than the stave, is the natural display for this subject. A stave says which note; a keyboard says how far. And it is why scales and chords are usually written as interval patterns rather than as lists of pitches: the pattern is the object, and the pitches are one of twelve instances of it.

The independence is not total. Very low intervals are rougher than the same interval higher up, and very high ones lose definition, so the ratio account is exact only in the middle of the range. But over the two or three octaves most music lives in, it holds well enough that transposition is musically free.

Why the octave is different from everything else

The octave deserves separating out, because it is not just the simplest ratio but a qualitatively different relationship.

At 2:1, every partial of the upper tone coincides with a partial of the lower — the upper tone adds no frequency the lower did not already contain. It reinforces the even partials and introduces nothing. That is a stronger statement than “the partials mostly line up”, and it is unique to the octave.

The perceptual result is octave equivalence: notes an octave apart are treated as the same note, given the same name, and used interchangeably in harmonic contexts. Nearly every musical culture with a pitch system does this, and no other interval gets the same treatment anywhere.

The circle of fifthsThe twelve pitch classes arranged so that each is a fifth above the last. Keys next to each other differ by one sharp or flat, which is why the notes of a key form a contiguous arc rather than a scattering.CaGeDbAf♯Ec♯Ba♭F♯e♭C♯b♭A♭fE♭cB♭gFdone step= one fifth= one sharpouter ring: major keys · inner ring: their relative minors
Fig. 7 The twelve pitch classes, arranged by fifths. The fact that this diagram has twelve positions rather than an infinite number of them is octave equivalence at work: every C is one point, whatever octave it is in.

It is worth registering how much of the theory this assumption carries. Pitch class, chord inversion, the circle of fifths, the whole apparatus of harmonic analysis — all of them assume that the octave collapses. Take the assumption away and none of the diagrams work.

The ratio nobody agrees about

If simple ratios are consonant and complex ones are not, there should be a clean ordering. There nearly is, and the exception is instructive.

The minor seventh is 16:9 in Pythagorean terms, 9:5 in just intonation, and 7:4 if the seventh partial is admitted. Those are 996, 1018 and 969 cents — spread over half a semitone, and all three have been called the correct value by serious theorists.

The 7:4 version, the harmonic seventh, is genuinely smoother than either of the others; it is what a barbershop quartet sings on a dominant seventh chord, and the resulting chord has a fused quality that no keyboard can produce. It is also 31 cents flat of the equal-tempered minor seventh, which is a quarter of a semitone — enough that a keyboard player and a barbershop singer performing the same chord are not performing the same chord.

Whether the seventh partial is “allowed” was argued for centuries, and the argument was never about acoustics. It was about whether a ratio involving 7 counts as simple.

The simple ratios, and the twelve equal stepsOne octave laid out in cents. Above the line, the frequency ratios of small whole numbers, where they actually fall; below it, the twelve equal steps. The two sets almost never coincide.6/5minor third5/4major third4/3fourth3/2fifth8/5minor sixth5/3major sixthCC♯DE♭EFF♯GA♭AB♭BC+16-14-2+2+14-16the ratios of small whole numberstwelve equal steps of exactly 100 cents1200 cents to the octave
Fig. 8 One octave in cents with the simple ratios where they fall. The clustering matters: several ratios sit close enough together that a single keyboard key has to serve for all of them, and which one a listener expects depends on the context the chord arrives in.

Ratios are not the only thing the ear does

Two further complications matter, because a theory of consonance that ignores them predicts the wrong music.

Roughness is not the same as dissonance. The curve above measures sensory roughness, which is a property of the sound. What a listener calls dissonant also involves expectation: in the eighteenth century a dominant seventh chord was a dissonance requiring resolution, and by the twentieth it was a stable sonority to end a blues on. The chord did not change. Nothing about its roughness changed. Its function did.

Harmonicity is a separate cue. Two tones a fifth apart share a plausible common fundamental — they could both be partials of a note an octave below the lower one — and the auditory system appears to reward that, independently of roughness. This is why an octave and a twelfth sound fused even when they are far enough apart that no partials are close enough to beat at all.

Any complete account needs both, and current models use both.

Four triads as stacked intervalsEach triad drawn as the semitone distances above its root, with the nearest simple frequency ratio beside each note. The chords differ only in the size of the two stacked thirds, and that difference is the whole of their character.C1/1E5/4G3/243major0–4–7C1/1E♭6/5G3/234minor0–3–7C1/1E♭6/5F♯33diminished0–3–6C1/1E5/4A♭8/544augmented0–4–8semitones above the root
Fig. 9 Four triads as stacked intervals, with the nearest simple ratio beside each note. The major and minor triads are built from ratios that are simple; the diminished and augmented ones are not, and they behave completely differently in the music that uses them.

Whose music, and when

The claim that simple ratios sound consonant is about as close to a human universal as this subject gets, and the qualifications are still substantial.

The octave is essentially universal. The fifth is very widely privileged. Beyond that, agreement thins fast. The major third — 5:4 — was regarded as a dissonance in European theory until the fifteenth century, and medieval cadences resolve away from thirds onto bare fifths. Nothing physical changed in 1450; a repertoire changed, and the theory followed.

Meanwhile, traditions built on inharmonic instruments developed interval sets that a ratio-based theory does not predict, and traditions with continuous pitch — Indian classical music, maqam-based traditions — treat interval size as a continuous expressive variable rather than as a set of targets.

The safe statement is narrow: for tones with harmonic partials, roughness is minimised near simple frequency ratios, and listeners across cultures prefer low roughness for sustained simultaneous tones. Everything beyond that is a claim about a repertoire, and needs the repertoire named.

The unison, which is the limiting case

Push the ratio to 1:1 and the two tones become one. That sounds like a degenerate case with nothing to say, and it is where several of the subject’s threads meet.

A unison is maximally consonant — every partial coincides with every partial, and the roughness is zero. Move one tone by a few cents and the roughness does not rise gradually from zero; it rises very steeply, because the coinciding partials are the ones most exposed to small changes. A unison a few cents out is more obviously wrong than a fifth a few cents out, which is why tuning by ear starts with unisons and why an orchestra tunes to one note rather than to a chord.

Move further and the roughness peaks around a semitone, then falls as the ratio simplifies again. The whole octave is therefore a single excursion out of the unison’s well and back into the octave’s, with the named intervals sitting in the dips along the way.

That framing has one useful consequence: it makes clear that the consonant intervals are not points on a list but local minima of a continuous function, and that a system with more than twelve notes to the octave would find more of them. Nineteen and thirty-one tone systems do exactly that, and the extra consonances they reach are real rather than theoretical.

Where the model stops

Sine tones do not work. Two pure sine tones a fifth apart have no partials to coincide, and the roughness model predicts — correctly — that the interval is far less distinctive than it is with real instruments. Most listeners find sine-tone intervals oddly characterless, which is a nuisance for demonstrations and a good confirmation of the theory.

The model is for sustained tones. Roughness needs time to establish. A rapid arpeggio does not produce it, which is why a figuration can outline a chord that would be intolerable if sustained, and why notating an arpeggio and a chord identically hides something.

Level matters. Roughness grows with loudness, and an interval that is acceptable quietly can be unpleasant loudly. None of the figures here has an amplitude axis.

Register matters. The same interval is far rougher in the bass than in the treble, because the ear’s frequency resolution is coarser at low frequencies. This is why arrangers space chords widely at the bottom and closely at the top — a rule of thumb that is entirely a consequence of critical bandwidth, and is usually taught without the reason.

The dissonance curve draws one spectrum. The figure above uses a string-like partial list. A different instrument gives a visibly different curve, and the essay’s argument depends on that being true — but only one of them is drawn.

The ladder from here

Later rungs: roughness computed properly, and the Plomp–Levelt model in detail. Critical bandwidth, and why register changes everything. Beats as the mechanism. The missing fundamental, and pitch without energy at the pitch. Harmonicity as a second cue. Stretched and compressed spectra, and consonance in artificial timbres. Gamelan tuning, and instruments whose spectra chose their scales. Dissonance as expectation rather than sensation. And the tritone, which is the most interesting interval in the Western system precisely because it is the least consonant.

The Pythagorean experiment — a stretched string, a moveable bridge, and the discovery that the pleasant divisions are the simple ones — is repeatable in five minutes with a rubber band. It is the oldest quantitative experiment in any science that anyone still bothers to reproduce, and it still produces the same answer.