Perception and the listener

Counting produced the hierarchy

Ask listeners how well each of the twelve notes fits after a passage in C major and the answers are not a smooth gradient. They fall into four groups with no overlap at all: the tonic, then the rest of the tonic triad, then the rest of the scale, then everything else — categories the subject had names for centuries before anybody ran the experiment.

Assumes: Seven of the twelve, chosen unevenly

Play a listener a short passage that establishes C major, then a single note, and ask how well that note fits what came before. Repeat for all twelve. The result is a profile, it was published by Carol Krumhansl and Edward Kessler in 1982, and it is one of the most-used measurements in the subject.

The probe-tone profile, major key. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.
Fig. 1 The ratings for each of the twelve pitch classes after a major-key context. The bars are the measurement. The shading is not: it marks the tonic, the rest of the tonic triad, the rest of the scale, and the remaining five notes — four categories this subject had long before the experiment existed. The profile separates all four with no overlap whatever.

The structure is a partition, not a gradient

The obvious reading of that figure is that some notes fit better than others, which would be unremarkable. The actual structure is stronger and more specific.

Sort the twelve ratings and the order is: C, G, E, F, A, D, B, then the five accidentals. Now apply three cuts, defined without reference to the data:

Cut one: the tonic against everything else. C is rated 6.35. The next highest is G at 5.19. No other note comes near.

Cut two: the rest of the tonic triad against the rest of the scale. The lowest triad member is E at 4.38. The highest non-triad scale member is F at 4.09. The gap is clean.

Cut three: the scale against the chromatic notes. The lowest scale member is B at 2.88. The highest non-scale note is F♯ at 2.52. Clean again.

Three cuts, none of them fitted, and no exceptions in either mode. It is worth being clear about why that is a result rather than a coincidence: the cuts are defined by music theory, the numbers are supplied by listeners, and nothing in the experimental design encouraged them to agree. And the same three work for the minor profile: tonic 6.33, then the minor triad’s members, then the rest of the scale, then the rest — with the tightest margin in the whole result at the third cut, where the lowest scale member is 3.34 and the highest non-scale note is 3.17.

That is what makes the result more than a survey of preferences. The listener’s ratings recover the tonic, the triad and the scale — three objects defined by music theory, in a system the listener has never been taught to describe — purely from a judgement about how well a note fits.

The probe-tone profile, major against minor. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.
Fig. 2 Major and minor profiles side by side. The tonic is rated almost identically in both, which is expected. Where they part is the third: E is the second-highest note in major and E♭ is in minor, and each is near the bottom of the other’s profile. The listener’s map of a key is anchored on the tonic and organised by the mode’s own third, which is the note that names it.

Where the hierarchy comes from

If the profile recovers the theory, the question is what the listener learned it from — and the answer that has held up is: from counting.

Krumhansl reported that the major profile correlates with the duration-weighted frequency of occurrence of each pitch class in tonal repertoires at around r = 0.9. Notes that are common in the music are rated as fitting; notes that are rare are rated as not fitting; and the correlation is high enough that the two measurements are close to interchangeable.

That is a strong claim about a mechanism, and its strength is that it requires nothing except exposure. Nobody teaches a listener that the fifth degree is more stable than the sixth. Nobody needs to. The statistics of the music do it, over years, without instruction and without the listener being aware that any learning is going on.

The evidence for it being learning rather than acoustics is that the profiles are specific to the tradition. Profiles collected from listeners in North Indian classical music match the statistics of that repertoire, with different notes elevated, and the same experiment run on a rāga a listener does not know produces a profile closer to their own tradition’s than to the rāga’s. What the probe-tone task measures is the listener’s accumulated statistics, and different listeners have accumulated different ones.

The one place acoustics might still be doing work

An honest treatment has to note that the ranking is also roughly the ranking of consonance with the tonic, which the site has computed several times from an entirely different starting point.

The tonic is a unison with itself, the fifth is 3:2, the major third 5:4 — the top three ratings are the three simplest ratios available. That is not nothing, and it is the standard objection to a purely statistical account: perhaps the statistics of the repertoire are themselves a consequence of acoustics, and the listener’s profile is tracking one through the other.

The probe-tone profile, minor key. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.
Fig. 3 The same experiment in a minor key, which is where the acoustic objection has to be tested rather than assumed. The tonic triad here is 1, ♭3, 5, and the ♭3 is 6:5 — a less simple ratio than the major third it replaces, and less simple than several notes outside the scale. A listener whose hierarchy came from ratio simplicity would therefore have to rate a minor key’s own third below notes that do not belong to the key at all. The measured profile puts it second of twelve, and the three cuts fall exactly where the mode says they should.

Set that beside the lattice this site has drawn elsewhere — the pitch classes arranged by their just relationships, fifths on one axis and thirds on the other — and the agreement in the major case is real: the notes nearest the tonic on the lattice are the notes at the top of the major profile. What the minor case shows is that the agreement is a coincidence of the major mode rather than a mechanism, because the same lattice, unchanged, predicts the wrong second-place note the moment the mode changes.

There is an arithmetic version of that objection and it can be answered rather than argued with. Correlate the twelve ratings against two acoustic quantities — the roughness of each pitch class against the tonic, and the simplicity of its just ratio — and then do it again with the three notes the objection is built on removed.

whole profile with the tonic, fifth and modal third removed
major, ratio simplicity r = 0.802 0.509
major, roughness 0.576 0.227
minor, ratio simplicity 0.690 0.361
minor, roughness 0.385 0.129

The acoustic correlation is almost entirely the three notes that were already conceded. Ratio simplicity accounts for sixty-four per cent of the major profile’s variance and twenty-six per cent once those three are gone; roughness accounts for thirty-three per cent and then five. So the honest statement of the objection is that acoustics explains the top of the hierarchy and has very little to say about the ordering of the other nine — which is where the interesting structure is, since the second and third cuts both live down there.

The minor profile makes the same point harder. Its acoustic correlations are lower everywhere, and they have to be: the minor third is 6:5 where the major is 5:4, so the note that anchors the mode is the less simple of the two, and a listener whose hierarchy came from ratios would rate a minor key’s own third below several notes outside its scale. They do not.

What settles it, as far as anything does, is the cross-cultural evidence: profiles differ between traditions in ways that track repertoire statistics rather than acoustic distance. A note that is acoustically close to the tonic but rare in a tradition is rated low by listeners in that tradition. Acoustics may well explain why certain statistics arose in the first place; it does not explain the profile a particular listener has, and on the numbers above it does not explain most of one.

The triads each scale builds on its own degrees. Every chord that can be stacked in thirds on each degree of each scale, with its quality worked out from the intervals the scale actually supplies. The quality of the chord on each degree is a consequence of each scale's own step pattern rather than of a convention.
Fig. 4 The seven triads of a major key, each built on one degree of the scale. Read this against the profile and the second cut in the hierarchy stops looking arbitrary: the three notes rated highest are exactly the members of the chord built on the first degree, and the notes rated next are the ones that complete the other six chords. What the listener has learned is not a list of notes but a set of chords, and the profile is the shadow those chords cast on a twelve-note histogram.

That reading resolves something the raw partition leaves open. Why should the triad be a level in the hierarchy at all? A statistical account predicts that common notes are rated high, and it does not obviously predict that the notes of one particular chord form a group. They do because the triad is the harmonic unit of the repertoire the statistics came from: its members occur more often and, more importantly, occur in stronger metrical positions and for longer durations, which is exactly what the duration weighting picks up.

What the hierarchy is for

A hierarchy of stability is not decorative. It does work, and the work is predictive.

It finds the key. The standard algorithm for key-finding — Krumhansl and Schmuckler’s — correlates the pitch-class distribution of a passage against all twenty-four profiles and takes the best match. It is simple, it works well enough to be a baseline that later methods are measured against, and it is a direct claim that a key is a statistical object. Where it fails is instructive: it does badly on passages that establish a key by order rather than by content — a cadence names a key in four chords that a histogram of the same four chords does not distinguish from several others — and every improvement on it since has been an attempt to put the order back in.

It predicts melodic expectation. A note high in the hierarchy is expected, arrives without effort, and closes a phrase. A note low in it is a surprise, and surprise is an expressive resource. That is the mechanism behind every claim about tension and resolution in tonal music, restated as something measurable — and it is why a progression is a path with a destination rather than a sequence of equally good places to be. The tonic is not the end of a progression because a rule says so; it is the end because it is the note the accumulated statistics of a lifetime rate at 6.35 against everything else’s 5.19 and below.

And it makes distance between keys a computable quantity. Correlate two keys’ profiles and the result is a similarity; do it for all twenty-four and the structure that comes out is the circle of fifths with relative majors and minors attached — a result derived from listener ratings rather than from arithmetic, which is a genuinely surprising place for it to come from.

The order it produces is worth printing, because it is not quite the order the circle predicts. Against C major the nearest keys are A minor at 0.650, then F and G major tied at 0.591, then E minor at 0.536 and C minor at 0.512. The relative minor is nearer than either neighbouring fifth, and by a clear margin — so what the listener ratings give is not the circle with the relatives hung off it but a structure in which the relative is the closest neighbour of all, with the dominant and subdominant behind it and exactly equal to each other. That equality is itself a check: nothing in the profile knows which direction round the circle is which, and the two come out identical to three figures.

Drawn on the circle of fifths those five numbers are almost the familiar picture and not quite it. The neighbouring keys sit where the theory puts them and the relatives hang off them as they always have, but the relative is nearer than the neighbour, which no version of the diagram has ever indicated — every printed circle gives the fifth pride of place and treats the relative minor as an attachment. The listener’s ratings reverse that, and they reverse it by a margin larger than the gap between the fifth and the next key out. The two constructions agree on the ordering of everything except the one relation the diagram was drawn to display.

The arithmetic version of the same map is the count of shared notes, and it is worth putting the two side by side because they are close and not identical. Neighbouring keys share six of their seven notes; keys a tritone apart share two; and the ordering that count produces reproduces the profile correlations closely enough that the listener’s map of key distance can fairly be called the arithmetic of shared notes, arrived at without anybody counting anything deliberately.

The agreement between the two is close but not exact, and the discrepancy is informative. Correlating profiles weights notes by how much they matter, not merely by whether they are present, so a key that shares six notes with C major but disagrees about which of them is the tonic scores lower than the raw count suggests. The relative minor is the clearest case: A minor shares all seven notes with C major and its profile correlation is high but well short of one, because the two profiles disagree about where the peak is. Shared content and shared organisation are different quantities, and the listener is tracking the second.

Which computation produced the numbers

The profiles are quoted; the analysis of them is done here.

The two twelve-number vectors are Krumhansl and Kessler’s published values and are not computable from anything. Everything else in the figures is computed from them: the group means, the three cuts, and — crucially — the assertion that the four groups do not overlap. The site’s gate checks that assertion in both major and minor, so a mistyped digit anywhere in either vector that broke the nesting would fail the check rather than quietly producing a figure that draws four tidy bands over data that does not support them.

The correlation with corpus statistics is quoted rather than recomputed, because this site does not carry a corpus. That is a real limitation and it is worth stating plainly: the essay’s central causal claim — that the hierarchy is learned by counting — rests on a correlation the site cannot reproduce, and the figures illustrate the structure of the profile rather than its origin.

The experiment done to an infant

The learning account makes a prediction about when the hierarchy arrives, and it has been tested in the least verbal population available.

Infants cannot rate anything on a scale. What they can do is look longer at a loudspeaker playing something unexpected, and that preference is a usable measure. Run with tone sequences that either conform to or violate the statistics of a scale, it shows sensitivity to scale structure in the first year, and sensitivity to key-specific structure appearing later — somewhere in the second year and consolidating through the preschool years.

The ordering matters. Broad structural properties come early and the specific hierarchy of a specific tradition comes later, which is what a statistical-learning account predicts and what a nativist account does not. The same experiments run with artificial scales show infants learning those statistics too, given enough exposure, which is about as direct a demonstration as the method allows.

Whose music, again. These studies are overwhelmingly on infants in Western households listening to Western music, so what they establish is that the mechanism operates, not that its output is universal. The one thing they establish robustly about content is the negative: the hierarchy is not present at birth.

What the picture cannot show

The context that established the key. The profile depends on what was played first, and different context passages give measurably different profiles. A cadence produces a sharper hierarchy than a scale does.

Time. The profile is a static object and expectation is not: what a listener expects depends on the last three notes as much as on the key, and a model with only a key-level hierarchy predicts the same expectation everywhere in a phrase. Later models add a short-term component learned from the piece in progress alongside the long-term one learned over a lifetime, and both are needed.

Individual variation and training. Musicians produce sharper profiles than non-musicians. The shape is the same; the contrast is greater — which is what more exposure to the same statistics would produce, and is one more small piece of evidence for the counting account.

And it says nothing about why the repertoire has those statistics. The profile explains how a listener came to have a hierarchy. It has no opinion at all about why the music that trained them was written the way it was, and treating the profile as an explanation of the music rather than of the listener inverts the whole argument.

Nor whether the same listener would agree with themselves in a different register. The profile is collected in one octave and treated as a fact about pitch classes, which assumes octave equivalence — and the octave is not exactly 2:1 for a listener, which is a small crack in an assumption the whole method rests on.

And the task is a judgement, not a behaviour. Asking how well a note fits is asking for an introspective rating on a scale, which is a much more artificial act than listening. The results correlate well with less artificial measures, which is why the method has survived, but the correlation is not identity.

Whose music, and how far it goes

This is a claim about listeners enculturated in Western tonal music, tested with Western tonal contexts. The finding that generalises is the method: give a listener a context and ask about a probe, and what comes back is that listener’s statistics.

The finding that does not generalise is the profile itself. Applied to a tradition where a scale is not a set of pitches — where a rāga’s identity includes which notes are approached from where, and where two rāgas with identical pitch sets are different rāgas — a twelve-number vector is the wrong shape of object entirely. It records how often each note occurred and discards everything about order, which for that repertoire is most of the information.

That is a limitation of the measurement rather than of the listener, and it is a good illustration of a general hazard: an experiment that produces a number for every tradition may still be asking a question that only one tradition has an answer to.

Where the ladder goes next

This anchor is one rung old — the youngest on the site — and the obvious next one is time — the hierarchy above is a distribution over notes, and expectation in real music is a distribution over what comes next, conditioned on what just happened. That is a different and larger object, and the models that handle it are the ones that made statistical learning a serious account rather than a suggestive one.

It is also worth noting what the anchor now sits beside. A hierarchy learned by counting is one of three things on this site that a listener supplies rather than receives — the others being the metre, which is inferred from a preference-rule scoring, and the interval categories, which are what make a tempered third still a third. All three are learned, all three are specific to a tradition, and all three are invisible to any measurement made on the sound.

Sideways, the same argument arrives at consonance and does more damage there. If a preference for certain notes is accumulated from exposure, what about a preference for certain intervals? Consonance turns out to be half learned, and separating the half that is not from the half that is required going a long way from any conservatory.

Part 1 of 11

One essay in the series on Tonal-expectation. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down, the 8 sharing most with it of 36.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

ExpectationKey-findingProbe-toneStabilityStatistical learningTonal hierarchy