Concept

Identification — where it appears

The task of naming which category a stimulus belongs to, as opposed to hearing that two stimuli differ. It is much coarser than discrimination — a listener can hear a five-cent change long before they can say which interval either version was.

Named by 12 essays across 4 fields — each of them below, with the objects they name alongside it.

Three ways a category boundary could move, and how far each moves it. The predicted shift of one boundary against how strong the context is, for three mechanisms. Expectation alone — a listener who thinks one category 20 times more likely than the other — moves the optimal boundary by σ²·ln(odds)/Δ, which with the eleven-cent noise used here is 2.8 cents at ten to one and 3.6 at 20. Re-learning the centres from a context 30 cents away moves it by half of that, 15 cents. Selective adaptation moves it the OTHER way. The two directions are what an experiment would separate, and no absolute calibration is needed to do it.

The boundary that barely moves

Every identification figure here has fixed category centres, and the essay before this one ended by admitting that real boundaries are supposed to move with context. Three mechanisms could move one, and their predictions are an order of magnitude apart and in two different directions. Expectation on its own — a listener who thinks one interval twenty times more likely than the other — is worth three and a half cents.

perception · Categorical-hearing
Every result so far, against the listener's own noise. Three findings drawn against the one parameter all of them assume: how finely the listener resolves a pitch. At 11 cents — a trained listener, and the value every earlier essay used — the octave holds 6 nameable categories, twelve equal ones are named right 91 per cent of the time, and a 20-to-one expectation moves a boundary by 3.6 cents. At 35 cents it is 2 categories, 72 per cent, and 37 cents. The capacity falls roughly as one over sigma and the shift rises as its square, so the three curves separate rather than moving together.

The listener the model was never run for

Five earlier essays rest on one number — how finely a listener resolves a pitch — and every one of them used a trained listener's eleven cents. The model's dependence on it is not gentle: the capacity goes as its reciprocal and the expectation shift as its square, so an untrained listener at thirty-five cents has two nameable categories per octave rather than six, and a foreign tuning system is not mis-transcribed by them but absorbed.

scales · Categorical-hearing
Every resolution claim here, against the number it rests on. How many equal steps of the octave can be named at 95 per cent accuracy, against the internal noise the model gives a listener. The laboratory value this collection quotes everywhere is 11 cents, which gives 6 nameable categories — the "about seven per octave" every claim here has been repeating. The laboratory measures it on isolated intervals and music never presents one, so the effective value inside a piece is smaller by an amount nobody has measured: at 4 cents it is 18, which is the chromatic scale, and the conclusion changes from "the ear has fewer boxes than the notation" to "it has exactly as many". Nothing here measures it. This is what it is worth if it moves.

The number every claim here has been quoting

Sigma is the internal noise a listener's pitch judgements carry, it is measured in a laboratory on isolated intervals, and music never presents an isolated interval. Every resolution claim here rests on the laboratory value. Sweep it and the headline finding moves: at eleven cents the ear has about six nameable categories per octave, and at six it has twelve — which is the chromatic scale, and turns 'the ear has fewer boxes than the notation' into 'it has exactly as many'.

scales · Categorical-hearing
6 step patterns, 2 category counts, and only the count matters. Each scale drawn as its own steps across the octave, with the share of trials a listener whose internal noise is 11 cents names correctly — and, beside it, the same figure for a scale of equally spaced degrees of the same size. The two agree to three decimal places on every row, though the narrowest step here is 100 cents on the diatonic major, tempered and the widest is 300. Naming is lost at boundaries, an n-degree scale has n of them wherever they are put, and each costs the same as long as no category is narrow enough for the noise to carry an estimate clean across it — 9.1 standard deviations, at the worst here.

A boundary costs the same wherever it is put

Every capacity figure drawn until now cuts the octave into equal categories, and no scale in the world is equal. Putting the real step patterns through the same model returns exactly the same accuracy to four decimal places — because naming is lost at boundaries, an n-degree scale has n of them wherever they are, and each costs the mean absolute value of the noise. That turns the most-quoted result about hearing from a search into one line: 1504 times one minus the criterion, over sigma.

scales · Categorical-hearing
A struck note's two ends are the same for every loss law. The partial levels of a string spectrum struck at 80 decibels on 130.8 hertz, and what is left of it when the fundamental itself falls under the threshold of hearing, for three laws relating a partial's decay rate to its number. The left panel is every one of them: a loss law cannot change the spectrum at the instant of the strike, because no time has passed. The other three are every one of them too: whatever the law, the note ends with nothing above the threshold. So both ends of the slide are shared, and everything that distinguishes an exponent of 0.5 from an exponent of 1 from an exponent of 2 is in the middle.

The middle nobody could have guessed

A struck note has no steady state, only a slide from one spectrum to another — so the question is what the middle carries that the ends do not. The answer is exact rather than statistical: every loss law in the family leaves the strike with the same spectrum and ends in the same silence, so both endpoints carry precisely nothing about which of them it is. The whole difference is 41.3 decibels, and it peaks 0.38 seconds in, seven per cent of the way through the note.

timbre · Envelope
Where a bar of 25 and a bar of 9 first disagree. The additive metre 2+2+2+3+2+2+2+3+2+2+3 — 25 units, an onset at the head of every group — with the accents it predicts drawn above the accents predicted by reading it as a repeating bar of 9, which is the cut of it that agrees longest. The two rows are identical for 24 consecutive steps and differ for the first time at step 25, where the shorter reading expects an accent and the metre does not supply one. Nothing before that step distinguishes the two hypotheses, so a listener who has not heard 25 consecutive steps has no evidence either way — whatever they are disposed to hear.

A twenty-five is a nine until its last unit

Every account of long additive metres says they are heard as groups of shorter ones, and the metre-induction model had never been pointed at the claim. Pointed at it, the model does not prefer the group — it prefers the long bar outright, and would go on preferring it more the longer anybody listened. What it cannot do is start: the evidence that separates a bar of twenty-five from a bar of nine does not exist until the whole bar has been heard, and at the tempo an unequal metre is best played at the psychological present holds sixteen units.

rhythm · Additive metre
A listener who knows every metre recognises none of them inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps from the downbeat before every other one of the 1820 arrangements of twos and threes has been contradicted by the stream, on onsets alone, with the long beats accented and with the downbeat accented, against the 16 steps a present of 3.5 seconds holds. onsets alone: recognised within the present for 0 of 1820; long-beat accent: recognised within the present for 0 of 1820; downbeat accent: recognised within the present for 85 of 1820. The dashed line is the present.

Knowing every metre is slower than knowing none

A long aksak bar cannot be told from its shorter cuts by induction before one step into its last beat, and no accent carried by the notes moves that floor. The obvious escape is a listener who knows the repertoire and recognises the metre instead. Recognition among all 1,820 arrangements of twos and threes never beats the floor, is never quicker than induction, and is slower for half the metres: a nine induced in 9 steps is recognised in 27. What breaks the floor is a small repertoire that leaves out the metre's own longest cut — with the cut known, no repertoire of any size does.

rhythm · Additive metre
Come in part-way with the downbeat accented, and no bar of sixteen units or more is recognised inside the present. For every bar length from nine units to twenty-five, the fewest and the most steps a listener who knows every arrangement of twos and threes needs to recognise the metre and where its bar begins, with the downbeat accented: coming in at a sample of steps inside the bar, against hearing it from its written downbeat. On onsets alone, or with the long beats accented, a metre entered part-way is never told from its rotations. 9: from inside the bar 10 to 17, 16 of 20 inside the present; from the downbeat 10 to 10; 10: from inside the bar 11 to 19, 12 of 20 inside the present; from the downbeat 11 to 11; 11: from inside the bar 12 to 21, 9 of 18 inside the present; from the downbeat 12 to 12; 12: from inside the bar 13 to 23, 8 of 24 inside the present; from the downbeat 13 to 13; 13: from inside the bar 14 to 25, 8 of 28 inside the present; from the downbeat 14 to 14; 14: from inside the bar 15 to 27, 4 of 28 inside the present; from the downbeat 15 to 15; 15: from inside the bar 16 to 29, 4 of 32 inside the present; from the downbeat 16 to 16; 16: from inside the bar 17 to 31, 0 of 32 inside the present; from the downbeat 17 to 17; 17: from inside the bar 18 to 32, 0 of 36 inside the present; from the downbeat 18 to 18; 18: from inside the bar 19 to 32, 0 of 36 inside the present; from the downbeat 19 to 19; 19: from inside the bar 20 to 37, 0 of 40 inside the present; from the downbeat 20 to 20; 20: from inside the bar 21 to 35, 0 of 40 inside the present; from the downbeat 21 to 21; 21: from inside the bar 22 to 40, 0 of 44 inside the present; from the downbeat 22 to 22; 22: from inside the bar 23 to 39, 0 of 44 inside the present; from the downbeat 23 to 23; 23: from inside the bar 24 to 44, 0 of 48 inside the present; from the downbeat 23 to 24; 24: from inside the bar 23 to 39, 0 of 48 inside the present; from the downbeat 23 to 25; 25: from inside the bar 22 to 44, 0 of 52 inside the present; from the downbeat 23 to 25. In all, 61 of 590 entries are recognised within the 16 steps of a 3.5-second present.

A dancer who comes in late needs the downbeat marked

Every window for recognising an aksak metre so far started at its written downbeat. A dancer joining a dance already going has not heard the downbeat, and the arithmetic of that is blunt: a metre entered part-way is, onset for onset, each of its own rotations heard from their downbeats, and the rotations are metres too — 2+2+3 and 3+2+2 are counted differently. So on onsets, and with the long beats accented, no metre is ever told from its rotations. Only an accented downbeat tells them apart, and with it a listener who knows thirty metres recognises 54 per cent of them inside the present from a random entry, against 1 per cent without.

rhythm · Additive metre
A bow shows the loss law a blow conceals. The same string on 130.8 hertz under three loss laws, drawn twice each: struck, and held by a continuous drive. The pale marks are the spectrum a blow produces, and they are identical in all three rows — a strike is the source spectrum and has no loss in it yet, which is why both endpoints of a struck note were found to carry nothing about the law. The solid marks are where each partial settles when a drive balances its own loss, at drive over loss, so the steady spectrum rolls off as the source's roll-off plus the exponent. At an exponent of 0.5 the held spectrum's centroid sits at 167 hertz, 5.7 semitones under the strike's 232; At an exponent of 1 the held spectrum's centroid sits at 144 hertz, 8.2 semitones under the strike's 232; At an exponent of 2 the held spectrum's centroid sits at 133 hertz, 9.6 semitones under the strike's 232. The quantity that is invisible at both ends of a struck note is the slope of a bowed one, for as long as the bow moves.

A bow holds the number a blow hides

Two struck notes with different loss laws are identical at the strike and identical at the end, which is why separating them at all meant looking in the middle. Drive the same two strings continuously and the loss law stops being a rate and becomes a slope: each partial settles at its drive over its own loss, so the exponent adds to the source's roll-off and sits in the spectrum for as long as the bow moves. It is 18.7 decibels of separation available from the first instant, against 41.3 that a blow delivers after four tenths of a second and then takes away.

timbre · Envelope
From partial 3 the room is the slower of the two. Decay rates in nepers a second for each partial of a note on 130.8 hertz, in a concert hall. The rising curve is the string's own loss, which grows as the partial number to the power 1. The flat-ish curve is the room's, from its reverberation time at that partial's frequency. A reverberant field is the source convolved with the room, so a partial's tail falls at the SLOWER of the two — the heavy line — and the room keeps returning energy the string has stopped making. From partial 3, at 392 hertz, the room is in charge: 6 of the note's 8 partials are held up by the room rather than let go by the string. Those are exactly the partials the string was losing fastest, which is why the room does not merely lengthen the note.

The room is the slower of the two

A reverberant field is the source convolved with the room, so a partial's tail falls at the slower of the two rates rather than at their sum — and the room is slower for exactly the partials the string is losing fastest. Half a note's colour is gone in 0.163 seconds in no room at all, 0.313 in a concert hall and 1.441 in a stone church. The destination is identical in all three, because a room cannot hold a partial up above the fundamental it is also holding. What a hall takes away is the rate, and the rate was the whole of the identity cue.

timbre · Envelope
The drain does not stop when the key does. The note does.. Semitones of colour gone, against time, for a note on 130.8 hertz left to ring and for the same note released after 0.4 seconds onto a damper of 0.15 seconds. The two curves lie on each other until the key comes up, and the damped one then ends: the note is inaudible at 0.52 seconds with 8.7 of the free note's 9.9 semitones delivered. A damper adds one loss to every partial alike, so it adds the same number to every decay rate and leaves every DIFFERENCE between rates exactly as it was — the spectrum at each instant is the ringing spectrum shifted bodily down by 400 decibels a second. The colour goes on draining at its own rate the whole time. What the damper takes away is not the drain but the seconds.

A damper changes the clock, not the colour

A damper is an extra loss on the string rather than a second decay, so it adds the same number of nepers a second to every partial — and adding a constant to every rate leaves every difference between rates exactly where it was. The damped spectrum at any instant is the ringing spectrum at that instant shifted bodily down, to machine precision. The colour goes on draining at its own rate; the note simply runs out of seconds, and how many it gets is written on the page as a note value and a tempo.

timbre · Envelope
A damper is a loss on the string, so a room can overrule it. Decay rate in nepers a second against partial number, for a note on 130.8 hertz in a room of 2 seconds. The rising line is the string's own loss, 1.15 nepers a second at the fundamental and growing as the partial number to the power 1. The line above it is that plus the damper's 46.1, which is what the string does once the key comes up. The flat line is the room. What a listener receives is the SLOWER of the damped string and the room, because a hall goes on radiating what the string has already given it — and here the room is slower on 8 of 8 partials, from the fundamental upward. The composition proposed earlier — take the slower of the string and the room, then add the damper to whichever won — would put the damper outside the minimum, where nothing can overrule it, and would predict a note 2.54 seconds shorter than ringing where the arithmetic here predicts 0.90.

A damper cannot reach into the room

The essay before this one proposed the arithmetic for a damped note in a hall: take the slower of the string's rate and the room's, then add the damper's to whichever won. The composition is wrong, and it is wrong in the one place that decides the answer. A damper is a loss on the string, so it belongs inside the minimum where a room can overrule it — and past about three seconds of reverberation it is overruled on every partial, so the damper removes no audible seconds of note at all.

timbre · Envelope

Named alongside it

The objects these essays reach for when they reach for this one.

BrightnessDecayEnvelopeSpectral centroidCategorical perceptionAdditive metreAksakCategory boundaryDampingInferenceMetreMicrotonality

All concepts