Perception and the listener

The spectrum that will not fuse

A partial about one per cent off its harmonic is heard as a sound of its own rather than as part of a note. Apply that criterion to a whole spectrum instead of to one mistuned component and it becomes a count: a piano string keeps nine of its ten partials, a bell keeps seven of eight, a bar keeps two of six. The physics of inharmonicity has had an essay here for a long time. This is what it sounds like.

Assumes: What makes two partials one note · The piano is tuned wrong on purpose

The second rung of this ladder asked what makes a set of simultaneous components one note rather than several, and found that the strongest single cue is harmonicity: a partial that is a whole multiple of the fundamental belongs, and one that is about one per cent off does not — it pops out of the note and is heard as a separate whistle at its own pitch.

That rung’s last paragraph pointed at where the question goes:

What a listener does with a spectrum that will not fuse — a bell, a gong, a detuned pair — is a question about inharmonicity, which already has an essay on the physics and not yet one on what it sounds like.

The physics essays are there. A stiff string’s partials are sharp; a bell’s ratios are a founder’s target; a bar’s are the zeros of a Bessel-like function and nothing like a series. All three are drawn on this site as pitch — how far each partial is from where a harmonic would be. Read through the fusion criterion they become a different quantity, and the different quantity is a count.

How much of each spectrum a listener can assemble into one noteEach partial of each spectrum at the harmonic number it is nearest, against the whole-number series that fuses the most of them, with anything more than 1 per cent out marked as heard separately. an ideal string keeps 10 of 10; a piano string keeps 9 of 10; a bell keeps 7 of 8; a bar keeps 2 of 6; a kettledrum keeps 3 of 5. The fundamental is capped at a tenth of the top partial, and the cap is load-bearing rather than tidy: a bell's ratios are all whole multiples of a tenth, so an unconstrained search finds a fundamental twenty-five harmonics down, calls every partial exact, and reports that a bell fuses perfectly. Nothing that high is resolved and the low harmonics of it are not there.each partial, at the harmonic number it is nearestan ideal string10 of 10 fusea piano string9 of 10 fuse+1%a bell7 of 8 fuse+20%a bar2 of 6 fuse-78%-38%+21%+4%a kettledrum3 of 5 fuse-2%-4%filled: fuses into the notehollow: heard as a sound of its own
Fig. 1 Every partial of five spectra, at the harmonic number it is nearest, against the whole-number series that fuses the most of them. Filled means within one per cent and part of the note; hollow means heard as a sound of its own. An ideal string keeps all ten; a piano string keeps nine; a bell keeps seven of eight; a bar keeps two of six.

Why this is a count and not a distance

The obvious way to measure how inharmonic a spectrum is would be an average — the mean absolute deviation of its partials from the nearest harmonics, say, in cents or in per cent. That number exists and it is not the one a listener has.

The fusion criterion is a threshold, not a slope. A partial 1.2 per cent sharp is heard out and a partial 40 per cent sharp is heard out, and they are heard out equally: both are simply not part of the note. So a spectrum with one wildly wrong partial and nine exact ones is, perceptually, a note plus one extra sound — and a spectrum with all ten partials 1.5 per cent out is ten separate sounds and no note at all. Those two have similar averages and completely different perceptual descriptions.

The count is therefore the right summary, and it produces an ordering the average does not. A piano string’s partials go steadily sharper — 0.7 per cent flat at the fundamental, 1.3 per cent sharp at the tenth, under the fit that centres them — so it is exactly the shape that keeps almost everything and loses the top. A bell’s are scattered, so it keeps most of its partials and loses one badly.

How sharp each partial of a real string is. The departure of each partial from the whole-number multiple it is supposed to be, in cents, for two strings — an ideal string, a small upright at middle C. The coefficient of each is computed from the stiffness of a steel wire of a stated length and gauge; nothing is fitted. The sharpest of them is 22 cents sharp by the eighth partial.
Fig. 2 The piano string’s deviations against a perfectly flexible one in the form the physics gives them: every partial sharp of its harmonic, by an amount growing as the square of the partial number. Read as pitch, this is why octaves are stretched. Read through the fusion threshold, it is a note whose first nine components belong and whose tenth has left.

The fundamental has to be findable, and this is where a naive search fails

Computing the count needs a fundamental to count against, and finding it is where the arithmetic goes wrong if it is left unconstrained.

A bell’s partial ratios are the founder’s targets: 0.5, 1.0, 1.2, 1.5, 2.0, 2.5, 3.0 and 4.0. Search freely for the fundamental that makes the most of those whole multiples and the search returns 0.1, against which every one of them is exact — 5, 10, 12, 15, 20, 25, 30, 40 — and reports that a bell fuses perfectly.

It does not, and the reason is not arithmetic. A fundamental forty harmonics below the top partial is not a fundamental a listener can use. Nothing that high in a series is resolved by the ear, the low harmonics of that fundamental are entirely absent, and the whole apparatus by which a listener supplies a missing fundamental works on the first handful of harmonics rather than on the fortieth.

So the search is capped at ten harmonics from top to bottom, which is the range the residue ladder says the mechanism works over, and ties go to the higher fundamental for the same reason. With the cap the bell’s answer is 0.5 — the hum note — and seven of its eight partials fuse against it, with the tierce at 1.2 standing twenty per cent out.

That cap is a perceptual constraint doing work inside what looks like a numerical procedure, and it is the kind of place a computation quietly stops being about music.

It is worth noticing that the failure it prevents is not a small error but an inversion. Uncapped, the search rewards a spectrum for being more inharmonic, because a scattered set of ratios is more likely to have a low common divisor than an orderly one — a bell scores perfectly and an ideal string, whose partials are already whole multiples, cannot do better than perfect. So the unconstrained measure ranks a bell at least as fusible as a string, which is the reverse of the thing being measured. A cap does not merely tighten the answer; it is what makes the measure point the right way.

What a competition decides when the two cues do not agree. Each spectrum with its two cue readings and the grouping the competition chooses. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest. Where they disagree there is no rule in this collection, so the published apparatus is used instead: every way of splitting the partials into one stream or two is scored for the partials each cue says it has wrongly grouped and wrongly separated, and the cheapest wins. an ideal string — harmonicity 100 per cent, common fate 20, and the competition says one stream; a piano string — harmonicity 90 per cent, common fate 20, and the competition says a cut after partial 2; a bell — harmonicity 88 per cent, common fate 13, and the competition says a cut after partial 1; a bar — harmonicity 33 per cent, common fate 67, and the competition says a cut after partial 4; a kettledrum — harmonicity 60 per cent, common fate 100, and the competition says one stream. The exchange rate between the two cues is the number nobody here can supply, so what is reported beside each is how many decades of it leave the answer unchanged.
Fig. 3 What the cap is protecting, drawn as the competition it makes possible. Harmonicity asks whether a partial is near enough a whole multiple to belong; common fate asks whether it decays at the same rate as the rest; and where the two disagree there is no rule, so every way of splitting the partials into one stream or two is scored and the cheapest wins. That arbitration only means anything once the fundamental is constrained. Uncapped, the search rewards a spectrum for being more inharmonic — a scattered set of ratios is likelier to have a low common divisor than an orderly one — so a bell scores perfectly and an ideal string cannot do better than perfect. The cap does not tighten the answer; it is what makes the measure point the right way.

The tierce, which is the whole of what a bell is

Seven of eight is a high score and the one that is missing is the one that matters.

A bell’s tierce sits at 1.2 times the prime — a minor third above it — and the whole character of a bell is that minor third sounding against the rest of the spectrum. Twenty per cent off is a very long way from one per cent: it is not a slightly ill-fitting partial, it is a second note.

So the census says a bell is a note plus a minor third, which is exactly how a bell is described by everybody who has ever listened to one and is not what a spectrum-as-a-list picture makes obvious. The other seven partials — hum, prime, quint, nominal and the rest — are within a per cent of multiples of the hum and fuse into a single tone. The tierce does not join them.

That also explains something the bell essay recorded and did not account for: why founders tune the tierce so carefully when it is the partial furthest from any harmonic. If it were going to fuse, its exact placement would be a detail of timbre. Because it will not fuse, its placement is a melodic decision — it is a note, it will be heard as a note, and a bell whose tierce is out of tune is out of tune in the ordinary sense of the phrase.

The piano, which is the interesting middle

The census’s most useful entry is the one that is neither harmonic nor obviously not.

A piano string keeps nine of ten. Its partials are all slightly sharp and the sharpness grows as the square of the partial number, so there is a specific partial at which the note stops and a set of separate whistles begins — and where that partial is varies enormously across the instrument, because the inharmonicity coefficient does.

Solving for it — the deviation from a centred harmonic fit reaches one per cent at partial √(0.04/B + 1), which returns the hero figure’s nine-of-ten at the coefficient the piano figure uses — gives the crossing at every note on the instrument:

B first partial to leave
A0 5.0 × 10⁻³ 3.0
A1 1.2 × 10⁻³ 5.9
A2 2.7 × 10⁻⁴ 12.3
A4 1.2 × 10⁻³ 5.9
A6 9.4 × 10⁻³ 2.3
C8 3.8 × 10⁻² 1.4

Two of those numbers are worse than the essay assumed. The tenor’s crossing is at the twelfth partial, not above the twentieth — reaching twenty would need a coefficient a quarter of the smallest this instrument has — so even at its best a piano note keeps about a dozen components and loses the rest.

And the top octave’s crossing is below the second partial. At C8 it is 1.4 and at A7 1.6, which means the second partial has already left before it arrives: a top-octave piano note is a fundamental alone, with every other component of its spectrum heard as a separate sound. That is stronger than barely a note; by this criterion it is a sine tone accompanied by a cloud of unrelated whistles, and the whistles are what a listener calls the attack’s brightness.

The bass is nearly as bad and the essay’s U-shape says why. A0’s crossing is at 3.0, between A6’s 2.3 and A5’s 3.6 — so the bottom note of a piano assembles about as much of itself into a note as a note four octaves above it does. The instrument is a harmonic tone only in its middle two octaves, and at both ends it is a fundamental with a handful of partials.

That is worth setting against the usual account of why the extremes of a piano sound as they do. The bass is normally explained by the wound strings and the treble by the short decay, and both explanations are about those notes. This one is a single quantity with a U in it, and it predicts that the two ends should be alike in a specific respect — few components fused — which is not something either of the usual explanations predicts and which is a thing a listener could be asked about.

The crossing formula also says what a maker could do about it, and the answer is discouraging. The crossing goes as one over the square root of the coefficient, so halving B buys a factor of 1.41 in the number of partials that fuse — and halving B on a treble string means a longer or a thinner one, which is exactly what the case and the tension will not allow. To take C8 from 1.4 partials to a mere four would need the coefficient down by a factor of eighteen, which is not a refinement of a piano but a different instrument.

That is not the usual account of why a piano’s treble sounds thin. The usual account is about radiated power and about the short decay, both of which are true. This adds a third: there is less note up there, in the sense that fewer components are being assembled into one.

Twenty milliseconds, and it stops being one note. The same spectrum with an asynchrony imposed on its upper partials, from nothing to sixty milliseconds. Up the axis is how many different hypotheses win over the swept range of exchange rates — three means the answer depends on a number nobody has, and one means it does not. Below 20 milliseconds the winner is one stream at every exchange rate that matters; at and above it the winner is a cut, and the arbitration becomes completely settled. The step is as sharp as it is because the cue is modelled as a threshold rather than as a graded quantity, which is the shape of the published criterion and not a finding. What is a finding is where it sits: 20 milliseconds is the spread between the attacks of two orchestral instruments told to play together.
Fig. 4 The same criterion asked of the other cue, because a threshold that depends on an exchange rate nobody has measured is not a threshold. Up the axis is how many different hypotheses win across the whole swept range of exchange rates between the two cues: one means the answer does not depend on the number, and three means it does. Below twenty milliseconds of onset difference the winner is one stream at every rate that matters; at and above twenty the winner is a cut. So the twenty-millisecond figure is a real breakpoint rather than an artefact of a weighting — which is the same kind of check the fundamental’s cap needed, applied to the second cue instead of the first.

What the criterion is, and how soft it is

One per cent is a published figure and it is a soft edge being used as a hard one, so it is worth showing what happens on either side.

At half a per cent, the piano string loses its ninth partial as well as its tenth and the bar loses one of the two it had. At three per cent, the kettledrum picks up a partial and the bell keeps its seven. The ordering across the five spectra is unchanged at every threshold tried, which is the useful robustness: the census is not a good estimate of the number, and it is a reliable statement about which spectra fuse better than which.

Partial 3, mistuned by 3%A 10-partial tone on 220 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 3.0 per cent, which is 19.8 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune.19.8 Hz off12345678910partial numberamplitude1% is enoughto hear the partialout of the note;3% and it stopscontributing to thepitch of the wholeMoore, Peters &Glasberg, 1985
Fig. 5 The criterion itself, from an earlier essay: one partial of a harmonic tone taken off its multiple and heard out of the note. The census above is this criterion applied to every partial of a spectrum at once, and the threshold in this figure is the one the counting uses.

What this says about the instruments that use it

The census sorts the percussion family in a way that matches what those instruments are used for, which is at least suggestive.

A bar keeps two of six. A glockenspiel or a xylophone bar has partials at roughly 1, 2.76, 5.4 and up, none of which is a whole multiple of anything, and the census finds a fundamental that explains almost nothing. Which is why a xylophone has a definite pitch and almost no tone — the pitch comes from the fundamental alone, everything above it is a separate collection of sounds, and the instrument is used for lines rather than for chords. What a drum is doing instead is the essay about the modes themselves.

A kettledrum keeps three of five. Its membrane modes are deliberately shifted by the air enclosed in the kettle so that four of them approximate 2:3:4:5, which is why a timpano has a pitch at all where a plain drum does not — and the census says that the design works well enough for three modes and not for five.

And an ideal string keeps all ten, which is why every sustained pitched instrument in the orchestra is a string or a tube.

That last is worth stating as the general form: the instruments that carry harmony are the ones whose spectra fuse, and the ones that do not fuse are used for colour, for punctuation and for melody. A chord is a set of notes that a listener has to keep apart while hearing them together, and a spectrum that will not fuse has already used up the mechanism that keeps them apart.

Partial 3, 30 ms early. A 10-partial tone on 220 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 0.0 per cent, which is 0.0 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune.
Fig. 6 And the cue that has nothing to do with frequency at all. This partial is exactly in tune — zero per cent off the harmonic grid — and it starts thirty milliseconds before the rest of the note, which is enough to hear it out as a whistle of its own. Harmonicity is satisfied and the note still splits. That matters for the census above, because it means a spectrum’s fusibility is not a property of the spectrum: the same partial list, struck so that its components arrive together, holds together, and struck so that they do not, does not. Which is why the instruments whose components arrive over tens of milliseconds — struck bars, bells, plucked wound strings — are also the ones the census scores worst.

Which computation produced the numbers

The threshold is this site’s own constant for a partial being heard out of a complex, which is about one per cent of the partial’s own frequency.

For each spectrum, candidate fundamentals are formed by dividing every partial by each of the first six integers. Candidates that would put the top partial more than ten harmonics up are discarded. Each surviving candidate is scored by how many partials fall within the threshold of a whole multiple of it, and the best is taken, ties going to the higher fundamental.

The five spectra are: a plain harmonic series; a stiff string with a coefficient typical of a tenor piano note; a bell’s founder’s ratios; the first modes of a free bar; and a kettledrum’s. The last three are measured targets and measurements rather than derivations, and they are the only numbers in this essay that are not computed — which is the same distinction the instruments field states as a rule and is worth restating whenever a founder’s target sits beside a wave equation.

Where the model stops

Harmonicity is one cue and there are several. The second rung listed them: common onset, common modulation, common fate under a vibrato, and harmonicity. A bell’s tierce shares an onset and a decay with the rest of the spectrum, and those cues push it toward fusion while harmonicity pushes it out. The census counts one term of a sum, and a partial that fails the harmonicity test is not thereby guaranteed to be heard out.

Amplitude is nowhere in it. A partial thirty decibels down and one per cent sharp is not heard out, because it is not heard — and where the threshold of hearing sits under each partial is a curve of its own. Every spectrum here is counted as though its partials were equally loud, and none of them is.

And decay is nowhere in it either. A bell’s partials decay at wildly different rates — the hum lasts for many seconds and the upper partials for a fraction of one — so the census is a description of a moment rather than of a bell. Which moment matters is a question the envelope ladder has the machinery for and this rung has not used.

What the picture cannot show

How many sounds a listener reports. The census counts partials that fail a criterion; it does not follow that each failing partial is one perceived sound. Two failing partials close together might well fuse with each other, which would make a bell a note plus one extra sound rather than a note plus several, and nothing here distinguishes those.

And it cannot show the strike note. The most-discussed thing about a bell is that its perceived pitch is often at neither the hum nor the prime but at a strike note a listener supplies, and this site has an essay about the arithmetic of it. The census’s chosen fundamental is not that object — it is the series that explains the most partials, and the strike note is what a residue mechanism extracts from a subset. They are two different computations on one spectrum and they need not agree.

Where this ladder goes next

Three rungs. The ear builds objects and sometimes offers a choice; the same question asked on the simultaneous axis rather than the successive one; and now a spectrum’s inharmonicity read as a perceptual count rather than as a pitch error.

The rung after it is the one the missing cue names. Every spectrum in this census is a static list, and the strongest fusion cue after harmonicity is common fate — components that move together belong together, and components that do not, do not. A bell’s partials decay at rates differing by an order of magnitude, so they are emphatically not moving together, and a vibrato applied to a harmonic tone moves every partial in exact proportion. That is a second census over the same five spectra with time in it, the site has the decay rates for three of them, and the two counts would very likely disagree about which spectrum holds together best.

Part 3 of 8

One essay in the series on auditory scene. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Auditory scene analysisBellFusionInharmonicityMistuned partialPartialSpectrumTimbre