Concept

Expectation — where it appears

What a listener predicts will happen next, built from exposure to a repertoire and from the material of the passage in hand. Every model of it here predicts the next event, which is why none of them has a state corresponding to a piece ending.

Named by 27 essays across 5 fields — each of them below, with the objects they name alongside it.

The probe-tone profile, major key. How well each of the twelve pitch classes was rated as fitting, after a context establishing the key — Krumhansl and Kessler, 1982. The shading is not part of the measurement: it is the tonic, the rest of the tonic triad, the rest of the scale and the remaining five notes, which are categories this subject had before anybody ran the experiment. The profile separates all four without overlap.

Counting produced the hierarchy

Ask listeners how well each of the twelve notes fits after a passage in C major and the answers are not a smooth gradient. They fall into four groups with no overlap at all: the tonic, then the rest of the tonic triad, then the rest of the scale, then everything else — categories the subject had names for centuries before anybody ran the experiment.

perception · Tonal-expectation
Two ways to fill eight bars, and only one of them accelerates. The period against the sentence, drawn as the lengths of their constituent units against position on a grid of eight bars. The ratio beside each row is the mean unit length in its second half divided by the mean in its first: the period at 1.00, the sentence at 0.67. A ratio below one is an acceleration — the unit shortening as the phrase approaches its arrival — and a ratio of one is a plan whose unit never changes length.

One of these eight-bar phrases accelerates

The sentence and the period both occupy eight bars, both end with a cadence, and both are recognised by ear rather than counted. What separates them is arithmetic. One halves its unit halfway through and the other does not, and the difference comes out as a single ratio — 0.67 against 1.00 — computed from nothing but the lengths of the parts.

form · Phrase
Five signals, computed separately, and no total. The five components of closure for 6 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.

What makes an ending an ending

Cadences are ranked. The authentic one is strong, the plagal weaker, the deceptive weaker still — and the solver here measured the quantity that ranking is usually explained by and found it says something else entirely. What survives is not a weaker version of the ranking but a different kind of object, with five components and no total.

form · Closure
Five signals, computed separately, and no total. The five components of closure for 4 chord pairs. The first three are computed from the chords alone; the last two are properties of where the goal lands and how long it is held. There is no total column: the components are not commensurable and the ordering of these cadences depends on which is weighted.

An ending that exists so a bigger one can

Half of the cadences in tonal music are built to fail. A phrase that stopped convincingly at bar four would be a piece four bars long, so the ending at bar four is engineered to arrive and not to settle — and the components it withholds are exactly the ones its partner at bar eight supplies. Closure is nested, and the nesting is what turns two phrases into one thing.

form · Closure
How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; I – IV – V – vi totals 7.3 bits over 3 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire.

The chord that did not come

A deceptive cadence is described as a surprise, and the explanation offered is that the wrong chord arrived. Measured against the tonal hierarchy already in use, the wrong chord is the second best-fitting triad in the key — and two of its three voices do exactly what they would have done in the right one. The surprise is not statistical. It is one voice, and it is the bass.

harmony · Tonal-expectation
The period, as the piece goes by. The strongest lag of thirty-two-bar AABA computed on only the bars heard so far, against how many bars that is. The final answer is 4 bars; it is revised 6 times on the way, and is not reached for the last time until bar 29 of 32, which is 91 per cent of the way through and 64 seconds at 108 beats a minute. Nothing about the boundary operator is involved: this is the global statistic, and it is the half of the form that a first hearing cannot have.

The form a first hearing cannot have

Every figure so far was computed with the whole piece in hand. Run the same methods over only the bars already heard and one of the two methods survives intact — the boundary operator turns out to be causal at a fixed delay of a few bars — while the other collapses. The period of a piece is not knowable until the piece is nearly over, and in two of the six schemes here not until its last bar.

form · Repetition
The ranking is settled either side of one narrow band. Remembered repetition — each bar's best match to an earlier bar, discounted by exp(−Δt/τ) with Δt in seconds — for 6 schemes at 108 beats a minute, against the decay constant τ on a logarithmic axis. The order of the schemes changes only between 8 and 13 seconds; outside that band it is fixed, so an estimate of τ wrong by any amount that stays outside it leaves the ranking alone.

A return has to be remembered

A stripe four bars off the diagonal and a stripe twenty-four bars off it are the same ink and are not the same experience. Convert the lag axis to seconds, discount every comparison by how long ago it was, and the ranking of these six schemes by how repetitive they are changes — and the decay constant and the tempo turn out to enter the arithmetic as one number rather than two.

perception · Repetition
thirty-two-bar AABA, as a strip of time. thirty-two-bar AABA laid out one cell per bar, coloured by section, with the roman numeral in each bar. the A section's turnaround is the ii-V every variant keeps; the bridge is a chain of applied dominants. At 108 beats a minute in 4/4 the whole of it lasts 71 seconds. Cut into 4 repeat units of 8 bars, 3 pairs of units agree on more than 50 per cent of their bars. 2 of them are not identical, and 2 of those 2 differ in a run of bars ending at the last bar of the unit; the changed bars are marked in orange.

Where a repeat is changed

Cut every scheme into its own repeat unit, compare each unit with every other, and ask where a repeat stops agreeing with what it repeats. The answer is that it stops at the end, in every case the corpus contains — and the number of cases the corpus contains depends entirely on where the threshold for "a repeat" is put. Moving it by nothing at all takes the count from two to twenty and the finding with it.

form · Repetition
One pattern, four metres. The same 16-step onset pattern read under 4 candidate metres, each scored by a preference rule set: 3 for a strong position that carries an onset, -2 for one that does not, -1 for an onset that lands off every strong position. The scores are downbeat on step 1 0, downbeat on step 2 -12, downbeat on step 3 -12, downbeat on step 4 0, so downbeat on step 1 and downbeat on step 4 tie and the model does not choose. Nothing about the sound differs between these readings; the bar line is supplied by the listener.

The beat that is never sounded

A listener who has heard four bars of a groove and then hears two bars with the downbeats taken out does not move the downbeat. This site's rule set does, every time, on every pattern tried — and the direction it moves in says exactly what kind of model would be needed instead.

perception · Metre induction
A tracker following a tempo that will not stay still. The tracker's error as a fraction of a beat, against beat number, for 3 rates of tempo change with a period-correction gain of 0.2. It settles at 0.025 of a beat behind at 0.5 per cent a beat, 0.092 of a beat behind at 2.0 per cent a beat, 0.206 of a beat behind at 5.0 per cent a beat. It does not lose the beat; it lags, by very nearly the rate divided by the period-correction gain, and the lag reaches a quarter of a beat at 6.3 per cent a beat — at which point the tracker is nearer the wrong onset than the right one.

A metre has to be able to change its mind

Replace the scoring function with a phase-corrected oscillator and the two failures reported earlier separate. The beat now survives two silent bars, drifting 23 milliseconds of a 560-millisecond beat. But it does not lose a moving tempo so much as lag behind it, by the rate over the correction gain — and a cadential ritardando that halves the tempo in eight beats is faster than the model can follow.

form · Metre induction
The price of a tonic. Every note of Dorian is given the same duration except its tonic, which is lengthened; the horizontal axis is the share of the total that goes to it. The key-finder answers with the parent key until 25.0 per cent of the time is spent on the modal tonic, and with D minor above it. At the left-hand edge every note has equal weight, which is the pitch-class set itself — and with every weight identical the correlation is not merely low but undefined, because a flat histogram has no variance to correlate with anything.

What a tonic costs in seconds

The standard key-finding algorithm cannot be run on a pitch-class set at all — a flat histogram has no variance and the correlation is undefined. Give it durations and it answers with the parent key for all seven modes identically, and it takes between 15.8 and 30.0 per cent of the total time spent on one note before it names that note instead.

perception · Modes
How many bars a key change takes to be heard. A twelve-bar progression that moves to G major at bar 6, read by the same correlation against all twenty-four profiles, with a window of 3, 4 and 8 bars. With 3 bars of history the new key is never the answer at all. With 4 bars of history the answer is G major from bar 7, one bar late, and it holds it from there. With 8 bars of history the answer is G major from bar 9, 3 bars late, and it holds it from there. The pivot bar is ambiguous by construction — it belongs to both keys, which is what makes it a pivot — so the lag is not a defect of the algorithm but a statement about how much evidence a key is.

How much evidence a modulation needs

Run a key-finder bar by bar over a progression that moves to the dominant at bar six. With four bars of history the answer becomes the new key at bar seven and holds. With three bars it never gets there at all, and reports E minor and B minor on the way. The window decides the lag as much as the music does.

perception · Key-relations
The boundary operator run over only what has been heard. Foote's checkerboard novelty on thirty-two-bar AABA at a kernel width of 4 bars, computed twice: once with the whole piece available, and once using only the bars heard up to and including each bar. The kernel reaches 4 bars forward, so every cell it needs has been heard 3 bars after its centre — the retrospective curve replotted 3 bars to the right lands on the causal one, and the operator turns out to be causal at a fixed delay rather than blind. The dashed verticals are the encoding's real section boundaries and are not an input.

An ending that can be heard coming

Two measurements are both called hearing an ending coming and they point in opposite directions. By the halfway mark of an ordinary form almost nothing new arrives — and the cost of coding each bar has not fallen at all. Neither statistic says anything is about to stop, because no statistic over content can: predicting the next event well is not predicting that there will not be one.

form · Closure
Four endings, and the loudness each produces from the page alone. Short-term loudness through the closing 6 bars of a thirty-two bar scheme, computed from the part count of each bar with no performance data of any kind — the parts are realised every way their ranges allow, every partial is placed in its critical band, and the sum is run through the two loudness smoothers. thins to one arrives at 0.764 of the running impression; full final chord arrives at 0.952 of the running impression; unchanged arrives at 1.000 of the running impression; thins then full arrives at 0.929 of the running impression. The result worth the figure is that full final chord is not the loudest: adding parts to a final chord adds power and almost no loudness, because the extra parts land in critical bands the chord already occupies. An ending is made loud by contrast with what preceded it, not by thickness.

A final chord is not made loud by adding to it

An earlier essay on closure said the loudest cue an ending has needs a corpus rather than an arithmetic. The arithmetic was built one essay ago, so it does not. Run four ending textures through it and two things come out backwards: a final chord three parts thicker than the rest arrives *quieter* against the running impression than the passage it ends, and a texture that drops a part a bar does not get quieter at all until the bar where there is one part left.

form · Closure
What the notes in between do to the anchor the interval is measured against. How finely a 7-semitone interval can be judged when its two notes are separated by other notes rather than by silence, under the two published accounts. Confirming material restates the key and refreshes the shared reference, so the correlation climbs from 0.5 toward a ceiling and the limen falls to 4.81 cents. Overwriting material competes for the same memory, so the correlation decays to 0.04 and the limen rises to 9.28. By 8 notes the two accounts differ by 4.5 cents, which is 47 per cent of the limen with no anchor at all — and no experiment here distinguishes them.

The notes in between

Every figure until now is about two notes with nothing between them, and a melody is notes with other notes between them. Two published accounts of what the intervening material does predict opposite signs — one says the key is restated and the shared reference is refreshed, the other says each note competes for the same memory and it decays. By eight notes they differ by four and a half cents, which is nearly half the limen the interval would have with no anchor at all.

intervals · Pitch-acuity
How surprising each chord is, in bits. Each step's information content, −log₂ of the probability the root-motion weights used here give it. a perfect cadence totals 6.4 bits over 3 steps; a deceptive cadence totals 7.3 bits over 3 steps; I – IV – V – vi totals 7.3 bits over 3 steps. The single most surprising move drawn is IV to V at 2.7 bits, which is 42 per cent of everything its passage spends. The eight weights are ordinal and stipulated rather than counted, so these are the numbers that ordering implies and not a measurement of any repertoire.

Surprise is a number

The chord that did not come was described rather than measured. Its measure is the information content of what did arrive, and a model of the probability has been to hand since the key-finding essays — eight root-motion weights, ordinal and stipulated. Reading them as a distribution prices a deceptive cadence at 2.71 bits against a perfect one's 1.85, and turns up the fact that the largest of the eight had never been read by anything.

harmony · Tonal-expectation
The surprise of each chord, against the uncertainty it arrived into. The information content of each step — minus the log of its probability under a distribution that multiplies the root-motion weight by how well the destination triad's notes fit the key — with the entropy of the moment before it drawn behind. I – IV – V – I: I→IV 1.71 bits, IV→V 2.80 bits, V→I 1.60 bits, against a mean uncertainty of 2.63; I – IV – V – vi: I→IV 1.71 bits, IV→V 2.80 bits, V→vi 2.61 bits, against a mean uncertainty of 2.63. A surprise larger than the entropy it arrived into is an outcome the model was not expecting even given how uncertain it was; one below it is an outcome the model had already mostly bet on. An earlier essay produced the first of those numbers and had no way to produce the second, because a set of preferences is not a distribution and only a distribution has an entropy.

A chord, given a key and a predecessor

A chord's improbability has been priced from its root motion alone, which left one multiplication unmade: a chord is also improbable because its notes do not fit the key, and that number has been available since the probe-tone profile. Multiplied and renormalised, the two give a conditional distribution — and a distribution has an entropy, which is the quantity a surprise has to be read against and which a list of preferences cannot supply.

harmony · Tonal-expectation
Expectation as a curve, and what a change costs where it lands. Every quantity so far is attached to a chord change: a list of surprises, one per event. A listener's expectation is continuous — it sharpens through a bar and collapses when the change arrives — and the two ingredients for it are already here, the harmonic rhythm and the metrical beat weights. The curve is the hazard: given that the chord has not changed yet, the chance that it changes on this beat. It runs from 0.043 on the weakest beat to 0.290 on the downbeat, a ratio of 6.7, against 0.125 if every beat were alike. The marked beats are where the changes actually arrive, and their timing bill is 3.6 bits against 6.0 for a listener with no metre — so these changes are 1.7 times cheaper to expect than a metreless listener would find them. That term is new: an earlier essay prices which chord arrived and this prices when, and a listener meets the sum.

Expectation is a curve, not a list

Every quantity so far is attached to a chord change: a list of surprises, one per event. A listener's expectation is continuous, sharpening through a bar and collapsing when the change arrives — and the two ingredients for it were already here, in two other accounts. What comes out is a second surprise, for when a chord arrives rather than for which one it is.

harmony · Tonal-expectation
How much of the reading comes from what has not happened yet. Every margin reported earlier is two-sided: the best path through a key at a bar is the best score into it plus the best score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line, because the dynamic program already had both halves separately. The mean margin falls from 3.90 bits with hindsight to 1.79 without it, so 54 per cent of this passage's certainty is retrospective. The two passes never disagree about which key is best here, so the hindsight buys confidence rather than a different answer. This is the quantity every earlier essay has assumed and none has measured.

How much of the reading arrives late

Every margin reported earlier is two-sided: the best path through a key at a bar is the score into it plus the score onward from it, and the second half uses bars a listener has not heard. Dropping that term is one line. On a thirty-two-bar song it removes more than half the certainty, and on a passage built to be ambiguous it changes the key named at nine bars out of eleven.

scales · Key-relations
The number nobody has moves the size and not the order. The mean total surprise per chord change, against how much the two surprises share. At zero they are independent and the total is their sum; at one they are the same event and the total is the larger of the two. The mean falls by a factor of 1.53 across that whole range, which is the size of the thing a corpus would settle. The ordering of the events by total surprise does not move at all until the very end: 5 of the 6 correlations swept give exactly the ordering independence gives, and only perfect dependence changes it, by 3 places out of 7. The most surprising event in the passage is the same one at every correlation. So the corpus three separate accounts have recorded wanting would change what this figure reports and not what it concludes.

Two surprises and one event

A chord change is surprising twice over — in which chord it is, and in when it comes — and a listener meets one event. Adding two surprises needs to know how much they share, which is a fact about a repertoire nobody has. Sweeping it instead: the total moves by half, the ordering does not move at all, and the most surprising moment in a passage is the same one whatever the answer turns out to be.

harmony · Tonal-expectation
The same eight notes are four times as much to read. How many bits each note of a line carries, taken as minus the log of the probability of the interval that reached it, under the distribution of melodic steps measured over the tunes used throughout. A scale costs 1.76 bits a note and a wide leaps costs 7.02 — a factor of 4.0 at the same number of notes on the page. Every quantity computed until now counts notes, and the page cannot tell these apart: eight quavers are eight quavers of horizontal space whichever line they spell.

A reader does not read notes

Eleven earlier essays count notes, and the page cannot tell one line of eight quavers from another. A reader can: a scale of eight is one object where eight leaps are eight. Measured against the melodic interval distribution, the same eight notes are four times as much to read — and the eye–hand span, the best-measured quantity in the reading literature, is four notes of a tune and one of a leaping line.

scales · Notation
A third of the timing surprise is paid for by nothing happening. Every beat of a bar of 4/4 at 1 chord change a bar, with what it costs a listener in expectation. The lower block is the arrival — the chance the chord changes here times what that change costs to be surprised by. The upper block is the hold, which is paid when the chord does not change and is charged at every beat rather than at every event. Over the bar the arrivals come to 2.61 bits and the non-arrivals to 1.29, so the term never spent before is 33 per cent of the total rather than most of it. The two together are exactly the binary entropy of each beat's own hazard, which is why the bar's whole timing bill is 3.90 bits and cannot be raised by rearranging where the changes fall.

The surprise of nothing happening

A hazard charges a listener twice — once when the chord changes and once, quietly, at every beat it does not. The second term was expected to dominate, because there are more beats than changes. It is a third of the bill at one chord a bar and never reaches a half at any rate a metre survives, because the cost of a beat where nothing happens is second-order small.

harmony · Tonal-expectation
What a chord change costs where it actually lands. Every beat of a bar of 4/4 at 1 change a bar, priced as a listener meets it: the pale block is what has already been paid waiting through the beats the chord did not come on, and the dark block is the arrival itself. Their sum is what it costs to be surprised by a change here. The last column is the remaining case, never priced before: no change in the bar at all, at 1.61 bits and a probability of 0.33. The 9 costs are a proper distribution — 1.000 — which is the check that this is one model rather than two. And the spread is the finding: the arrival term alone puts a factor of 6.7 between the best and worst beat, and counting the waiting makes it 19.

Where the chord actually lands

Every timing surprise so far is evaluated on a downbeat: the curve spans a factor of 6.7 and every number is read at its peak. Charge the waiting as well as the arrival and the costs over a bar become a proper distribution, the spread between the best and worst beat rises to a factor of 19.5, and a third of the probability sits on a bar in which nothing changes at all.

harmony · Tonal-expectation
Thirty cents out of tune is heard as 8 on a 125 ms note and 25 on a 1 s one. How far out of tune a note sounds against how far out of tune it is, at 4 note lengths, at 440 hertz. The key is treated as a prior over pitch: a mixture of Gaussians on the twelve scale degrees, weighted by Krumhansl and Kessler's probe-tone profile and given the width the degree account already uses. The likelihood is the note's own effective limen, which for a short note is the Fourier bound 1/2T. The estimate is the posterior mean, and the shrinkage toward a prior is one line of arithmetic. At 1 s a thirty-cent mistuning is heard as 25.0 cents and at 125 ms as 7.8. Every curve turns back up near the middle of the semitone, because past there the nearest degree is the other one and the pull reverses. The buttons sound at A4, which is the pitch the figure is computed at, at the shortest note length it draws.

A short note is heard more in tune than it is

Four earlier essays treat a key as something that reduces the noise in a pitch judgement. Treat it instead as a prior and the prediction changes kind: not a smaller error but a systematic bias, pulling a short note toward the nearest scale degree by an amount the Fourier bound sets. Thirty cents out of tune on an eighth-of-a-second note is heard as eight. And the part the debt got wrong is the part that matters — the bias does not vanish on a long note. It stops at 17 per cent at A4 and at 48 per cent at A2, because the likelihood's width has a floor that no duration removes.

intervals · Pitch-acuity
Eighty-one chords the expectation model cannot tell apart. Every voicing of a dominant seventh on G inside the three octaves above its own root, placed by how rough it is and how far its outer voices are apart, and coloured by which member of the chord is at the bottom. The roughness runs from 0.269 to 1.449, a factor of 5.4, computed from each voicing's own spectrum under Plomp and Levelt's roughness model. The identity surprise the expectation model assigns is 3.51 bits for every one of the 81, because it is a function of a scale degree and its predecessor and there is no register anywhere in it. What separates them is spacing rather than inversion: roughness falls as the outer voices spread apart, correlating -0.42 with the span, and is indifferent to which member of the chord is at the bottom at 0.02. The seventh in the bass is not what makes a chord rough; a fourth and a third packed together at the bottom of the range is.

Eighty-one chords, one number

A dominant seventh has eighty-one arrangements inside three octaves and their roughness spans a factor of five and a half. The tonal-expectation model gives every one of them the same 3.51 bits, because its states are scale degrees and there is no register anywhere in them. Conditioning the surprise on the voicing costs no corpus — and the arithmetic says the conditioning belongs beside the probability rather than inside it, for three reasons that can each be computed.

harmony · Tonal-expectation
Level does not dilute the register's roughness, it multiplies it. The mean roughness of the I – vi – IV – V – I arrivals at 4 registers, each relative to the register as written, read three ways. Level-free, the bass is 8.6 times rougher than the treble. With every note at 70 dB it is 8.6 times, the same factor, because one level rescales every pair alike. With each chord played at the level that makes it as loud as the written register's chords — 82.8 dB −2 octaves, 75.5 dB −1 octave, 70.0 dB as written, 66.8 dB +1 octave — the bass is 343 times rougher than the treble, because roughness grows with the square of the pressure and the bass needs more of it to be heard at the same loudness.

A rough arrival is rough because of its spacing

The pair the expectation essays report for every chord — how surprising it was, how rough its voicing is — has no level in it. Putting level back in answers the question it left open, and not the way it was framed. At one written dynamic the arrivals keep their order from 40 to 90 dB at three registers of four, and the bass stays 8.6 times rougher than the treble. Made equally loud, the bass has to be played 12.8 dB harder, and it is 343 times rougher: level does not explain the register's roughness away, it multiplies it.

harmony · Tonal-expectation
Counted over what arrives, the balanced bass is not the roughest register. The mean roughness of the I – vi – IV – V – I arrivals with each chord played as loud as the written register's, relative to the written register, counted over every partial and over the partials that stand above what the rest of the chord masks. Every partial: 70 −2 octaves, 7.48 −1 octave, 1.00 as written, 0.20 +1 octave. Delivered partials only: 6e-9 −2 octaves, 2.83 −1 octave, 1.00 as written, 0.19 +1 octave. Over every partial the lowest register is 343 times rougher than the highest; over what arrives it is the smoothest of the four, and the roughest is −1 octave, 2.8 times the written register.

A bass chord low enough to balance has already hidden its tenor

Played as loud as the written register, a progression two octaves down is 343 times rougher than the same progression an octave up — if every partial on the page is counted. Count only the partials that stand above what the rest of the chord masks and that register is the smoothest of the four, with nothing left that beats. The balance is not what does it: the extra thirteen decibels move no voice by more than two partials. The register had already buried the tenor at the written dynamic.

perception · Tonal-expectation

Named alongside it

The objects these essays reach for when they reach for this one.

CadenceInformationTonal hierarchyClosureMetreSurpriseKey-findingProbe-toneHarmonic rhythmRepetitionSegmentationCritical bandwidth

All concepts