Two notes and a ratio, which is the whole of consonance
Sound two tones at once and one of two things happens. Either they fuse into something that sounds like a single object with a colour, or they sit side by side and grate. The distinction is immediate, involuntary and shared across listeners who have never met.
What decides it is the ratio between their frequencies.
Frequencies in the ratio 2:1 give an octave, which fuses so completely that most musical cultures treat the two notes as versions of the same note. 3:2 gives a fifth, 4:3 a fourth, 5:4 a major third. Keep going and the ratios get more complicated as the intervals get less settled, until 45:32 gives a tritone, which does not settle at all.
It is also the reason the tuning systems exist: a system is an attempt to make as many ratios as possible simple at once. This is the oldest quantitative result in music, and probably in acoustics. It is attributed to Pythagoras, it is demonstrable on a stretched string with nothing but a moveable bridge, and it is correct.
It is also not, on its own, an explanation.
The classical argument, and its gap
The ancient reasoning runs: simple ratios produce a combined waveform that repeats quickly, and the ear prefers patterns it can grasp. The figure above draws exactly that. For a 3:2 pair the whole pattern completes in two cycles of the lower tone. For a 45:32 pair it takes forty-five.
The argument is appealing, and it survives in textbooks. It has a fatal problem: the ear is very nearly deaf to the shape of a waveform.
Take a complex tone and shift the phase of its partials. The waveform changes completely; a shape that looked like a sawtooth becomes something unrecognisable. The sound is essentially unchanged. This was established by Ohm in 1843 and confirmed exhaustively since, and it is one of the foundational facts of hearing — the same fact that makes a spectrum a better description of a sound than a waveform: the ear performs something close to a frequency analysis and discards most phase information.
If the ear cannot hear waveform shape, then “the combined waveform repeats quickly” cannot be why a fifth is consonant. The periodicity is real, the consonance is real, and the causal link asserted between them does not hold.
What actually happens
The working explanation is different, and it is about the partials rather than about the fundamentals.
Every real musical tone is a stack of partials at whole-number multiples of its fundamental. When two such tones sound together, what actually reaches the ear is two stacks — and the interesting question is how the partials of one relate to the partials of the other.
For an octave, every partial of the upper tone lands exactly on a partial of the lower. Nothing new is introduced at all; the upper tone reinforces partials the lower one already had, which is why the octave fuses so completely that it is barely heard as two notes.
For a fifth at 3:2, the upper tone’s partials fall on every third partial of the lower. Two-thirds of them coincide exactly and the rest sit in gaps.
For a tritone at 45:32, almost nothing coincides. Partials land near each other without landing on each other — and a pair of partials that are near but not equal beats, and beating at the wrong rate is roughness.
That curve is the modern answer, and its shape was not assumed: it comes out of a model of how two nearby frequencies interact in the ear, applied to every pair of partials and summed. The simple ratios come out as minima because coinciding partials do not beat.
The circularity that is not one
There is an objection that has to be met, because it is a good one. If consonance is explained by partials coinciding, and partials are at whole-number multiples of the fundamental, has anything been explained? The whole numbers went in and the whole numbers came out.
The answer is that the model makes a prediction the classical account cannot: change the partials and the consonant intervals move.
Take a tone whose partials are stretched — not at 1, 2, 3, 4 times the fundamental but at 1, 2.1, 3.3, 4.6 — and compute the dissonance curve again. The minima move. They no longer fall at 3:2 and 5:4; they fall wherever the stretched partials happen to coincide, and an interval that is consonant for that timbre is not one anybody would name.
This has been done, both computationally and with synthesised sound, and the effect is real and immediately audible. It also has a natural experiment attached: the Indonesian gamelan is built around metallophones and gongs, whose partials are genuinely inharmonic, and gamelan tuning systems — sléndro and pélog — divide the octave in ways that no theory built on the harmonic series predicts and that fit the instruments’ actual spectra well.
So consonance is not a fact about small integers. It is a fact about spectra, which for strings and pipes happen to be built on small integers.
The interval is the thing, not the notes
One consequence of consonance being a matter of ratios is easy to state and surprisingly hard to internalise: an interval is a relationship, and it is unchanged by moving both notes.
C to G and F-sharp to C-sharp are both 3:2. They sound like the same interval, they behave the same way in a harmony, and a listener asked to compare them will say they are identical — even though not one frequency is shared between them. The same is true of a melody: transpose it and it remains recognisably the same melody, because what a listener retained was the sequence of ratios rather than the sequence of pitches.
This is why the keyboard, rather than the stave, is the natural display for this subject. A stave says which note; a keyboard says how far. And it is why scales and chords are usually written as interval patterns rather than as lists of pitches: the pattern is the object, and the pitches are one of twelve instances of it.
The independence is not total. Very low intervals are rougher than the same interval higher up, and very high ones lose definition, so the ratio account is exact only in the middle of the range. But over the two or three octaves most music lives in, it holds well enough that transposition is musically free.
Why the octave is different from everything else
The octave deserves separating out, because it is not just the simplest ratio but a qualitatively different relationship.
At 2:1, every partial of the upper tone coincides with a partial of the lower — the upper tone adds no frequency the lower did not already contain. It reinforces the even partials and introduces nothing. That is a stronger statement than “the partials mostly line up”, and it is unique to the octave.
The perceptual result is octave equivalence: notes an octave apart are treated as the same note, given the same name, and used interchangeably in harmonic contexts. Nearly every musical culture with a pitch system does this, and no other interval gets the same treatment anywhere.
It is worth registering how much of the theory this assumption carries. Pitch class, chord inversion, the circle of fifths, the whole apparatus of harmonic analysis — all of them assume that the octave collapses. Take the assumption away and none of the diagrams work.
The ratio nobody agrees about
If simple ratios are consonant and complex ones are not, there should be a clean ordering. There nearly is, and the exception is instructive.
The minor seventh is 16:9 in Pythagorean terms, 9:5 in just intonation, and 7:4 if the seventh partial is admitted. Those are 996, 1018 and 969 cents — spread over half a semitone, and all three have been called the correct value by serious theorists.
The 7:4 version, the harmonic seventh, is genuinely smoother than either of the others; it is what a barbershop quartet sings on a dominant seventh chord, and the resulting chord has a fused quality that no keyboard can produce. It is also 31 cents flat of the equal-tempered minor seventh, which is a quarter of a semitone — enough that a keyboard player and a barbershop singer performing the same chord are not performing the same chord.
Whether the seventh partial is “allowed” was argued for centuries, and the argument was never about acoustics. It was about whether a ratio involving 7 counts as simple.
Ratios are not the only thing the ear does
Two further complications matter, because a theory of consonance that ignores them predicts the wrong music.
Roughness is not the same as dissonance. The curve above measures sensory roughness, which is a property of the sound. What a listener calls dissonant also involves expectation: in the eighteenth century a dominant seventh chord was a dissonance requiring resolution, and by the twentieth it was a stable sonority to end a blues on. The chord did not change. Nothing about its roughness changed. Its function did.
Harmonicity is a separate cue. Two tones a fifth apart share a plausible common fundamental — they could both be partials of a note an octave below the lower one — and the auditory system appears to reward that, independently of roughness. This is why an octave and a twelfth sound fused even when they are far enough apart that no partials are close enough to beat at all.
Any complete account needs both, and current models use both.
Whose music, and when
The claim that simple ratios sound consonant is about as close to a human universal as this subject gets, and the qualifications are still substantial.
The octave is essentially universal. The fifth is very widely privileged. Beyond that, agreement thins fast. The major third — 5:4 — was regarded as a dissonance in European theory until the fifteenth century, and medieval cadences resolve away from thirds onto bare fifths. Nothing physical changed in 1450; a repertoire changed, and the theory followed.
Meanwhile, traditions built on inharmonic instruments developed interval sets that a ratio-based theory does not predict, and traditions with continuous pitch — Indian classical music, maqam-based traditions — treat interval size as a continuous expressive variable rather than as a set of targets.
The safe statement is narrow: for tones with harmonic partials, roughness is minimised near simple frequency ratios, and listeners across cultures prefer low roughness for sustained simultaneous tones. Everything beyond that is a claim about a repertoire, and needs the repertoire named.
The unison, which is the limiting case
Push the ratio to 1:1 and the two tones become one. That sounds like a degenerate case with nothing to say, and it is where several of the subject’s threads meet.
A unison is maximally consonant — every partial coincides with every partial, and the roughness is zero. Move one tone by a few cents and the roughness does not rise gradually from zero; it rises very steeply, because the coinciding partials are the ones most exposed to small changes. A unison a few cents out is more obviously wrong than a fifth a few cents out, which is why tuning by ear starts with unisons and why an orchestra tunes to one note rather than to a chord.
Move further and the roughness peaks around a semitone, then falls as the ratio simplifies again. The whole octave is therefore a single excursion out of the unison’s well and back into the octave’s, with the named intervals sitting in the dips along the way.
That framing has one useful consequence: it makes clear that the consonant intervals are not points on a list but local minima of a continuous function, and that a system with more than twelve notes to the octave would find more of them. Nineteen and thirty-one tone systems do exactly that, and the extra consonances they reach are real rather than theoretical.
Where the model stops
Sine tones do not work. Two pure sine tones a fifth apart have no partials to coincide, and the roughness model predicts — correctly — that the interval is far less distinctive than it is with real instruments. Most listeners find sine-tone intervals oddly characterless, which is a nuisance for demonstrations and a good confirmation of the theory.
The model is for sustained tones. Roughness needs time to establish. A rapid arpeggio does not produce it, which is why a figuration can outline a chord that would be intolerable if sustained, and why notating an arpeggio and a chord identically hides something.
Level matters. Roughness grows with loudness, and an interval that is acceptable quietly can be unpleasant loudly. None of the figures here has an amplitude axis.
Register matters. The same interval is far rougher in the bass than in the treble, because the ear’s frequency resolution is coarser at low frequencies. This is why arrangers space chords widely at the bottom and closely at the top — a rule of thumb that is entirely a consequence of critical bandwidth, and is usually taught without the reason.
The dissonance curve draws one spectrum. The figure above uses a string-like partial list. A different instrument gives a visibly different curve, and the essay’s argument depends on that being true — but only one of them is drawn.
The ladder from here
Later rungs: roughness computed properly, and the Plomp–Levelt model in detail. Critical bandwidth, and why register changes everything. Beats as the mechanism. The missing fundamental, and pitch without energy at the pitch. Harmonicity as a second cue. Stretched and compressed spectra, and consonance in artificial timbres. Gamelan tuning, and instruments whose spectra chose their scales. Dissonance as expectation rather than sensation. And the tritone, which is the most interesting interval in the Western system precisely because it is the least consonant.
The Pythagorean experiment — a stretched string, a moveable bridge, and the discovery that the pleasant divisions are the simple ones — is repeatable in five minutes with a rubber band. It is the oldest quantitative experiment in any science that anyone still bothers to reproduce, and it still produces the same answer.