Intervals and chords

Every member of a beat family is the same depth

A mistuned octave's beats have been counted on their rate and their place, with the depth recorded as the thing left out and a prediction that the shallow upper members would take the count from four to two. The depth turns out not to fall at all: on any power-law spectrum every member of a family has exactly the modulation index its interval's own ratio gives, at every register, on every wire. The count does fall to two, and the thing that takes it there is the criterion that essay was already using.

Assumes: Beats are arithmetic that anybody can hear · Three beats at most, and only in the middle of the keyboard

Three beats at most counted the members of a mistuned octave’s beat family twice — once by whether each fluctuation is separable in the cochlea from the partials beside it, and once by whether it is slow enough to wait for — and ended by naming the quantity it had counted them without. A fluctuation is audible only if it is deep enough as well — beats are arithmetic anybody can hear only while there is something to hear — and the depth of the third and fourth members at A3 was expected to be one or two orders of magnitude below the first, because the amplitudes that make them fall away up the spectrum. Applying a threshold to that fall would take the answer from four members to two.

The depth is now computed, and the first half of that expectation is wrong by about the largest margin an expectation can be wrong by. It does not fall.

Every beat in a family is the same depth, and none is near the threshold. The 8 members of the beat family a octave mistuned by 6.0 cents makes at A3 on a real string, each placed at the rate it beats and at the modulation index a listener's filter delivers there. The rising curve is the published detection threshold for amplitude modulation, which is flat at 0.03 below about fifty fluctuations a second and rises above it. The flat dashed line at 0.667 is the depth the pair has in isolation, and it is the same for every member: on a spectrum falling as one over n, the k-th member pairs partials 2k and 1k, whose ratio is 2.00 whatever k is. What the filled points show is the smaller effect that does depend on the member — the partials on either side of the coincidence leak into the same filter, add level without adding fluctuation, and dilute the index from 0.662 to 0.405. Inside the countable rate window the narrowest margin over the threshold is a factor of 16.7, on member 4. The depth criterion removes nothing.
Fig. 1 The eight members of the family a 6-cent octave makes at A3, each at the rate it beats and at the modulation index its own auditory filter delivers. The rising curve is the published detection threshold; the flat dashed line is the depth the beating pair has in isolation, which is 0.667 for every one of the eight. Inside the countable window the narrowest margin over the threshold is a factor of 16.7.

The depth does not know which member it is

The arithmetic is two lines and it was available the whole time.

A beat has a depth fixed what the depth of a beat is: two components of amplitude aa and bb make an envelope swinging between a+ba+b and ab|a-b|, so the modulation index — the quantity a demodulator reports and the quantity every published detection threshold is stated on — is 2min(a,b)/(a+b)2\min(a,b)/(a+b). It depends on the two amplitudes only through their ratio. Scale both by any factor and the index is unchanged, because a fluctuation of a tenth of a quiet thing is the same fluctuation as a tenth of a loud one.

Now take the family. The kk-th member of a mistuned m:nm{:}n pairs the lower note’s partial kmkm with the upper note’s partial knkn. On a spectrum whose partials fall as 1/ns1/n^s those two amplitudes are (km)s(km)^{-s} and (kn)s(kn)^{-s}, and their ratio is

(km)s(kn)s  =  (mn)s\frac{(km)^{-s}}{(kn)^{-s}} \;=\; \left(\frac{m}{n}\right)^{s}

in which kk has cancelled. Every member of every family therefore has the same modulation index, and it is

index  =  21+(m/n)s\text{index} \;=\; \frac{2}{1 + (m/n)^{s}}

a number set by the interval and the spectral slope and by nothing else at all. For an octave on a one-over-nn string spectrum it is two thirds, which is a dip of 9.5 decibels — the exact figure in the octave row of the table the seventh rung computed from a single pair of partials, arrived at here from eight pairs at once.

How steep a spectrum would have to be before an octave's beat vanished. The modulation index of a beat between two partials of a mistuned interval, against the slope of the spectrum the partials come from. Each curve is the closed form 2/(1 + (m/n)^s), and each dot is one member of that interval's beat family computed from the two amplitudes it actually pairs — 32 in all, at two slopes, every one of them on its own interval's curve, because the k-th member pairs partials mk and nk and their ratio is (m/n)^s whatever k is. The horizontal line is the published detection threshold at a slow rate, 0.03. On a one-over-n string spectrum the octave sits at 0.667, twenty-two times above it; the curve does not reach the threshold until a slope of 6.0, which is a second partial 36 decibels below the first and steeper than any string, plucked, struck or bowed. The fifth and the major third never reach it at all within the range drawn.
Fig. 2 The index against the spectrum’s own slope, with one dot for every member of three intervals’ families at two slopes. Thirty-two dots land on six points, because within a family the member number cancels. The octave’s curve does not reach the detection threshold until a slope of 6.0, which is a second partial 36 decibels below the first.

That is a stronger claim than the seventh rung was in a position to make. Its table of seven intervals was computed from the coincidence a tuner is said to listen to — the lower note’s pp-th partial against the upper note’s qq-th — and treated as one row per interval. It is one row per family. A minor third’s beat is 20.8 decibels deep at its 6:5 coincidence and 20.8 decibels deep at its 12:10 and its 18:15, and the ordering of the seven intervals is the ordering of every member of all seven. That matters to the tuner counting seconds against a watch, who is choosing a member as well as an interval and has been told only about the first.

What falls as a power of the member number, and what does not

Something does fall away up a family, and it is worth being exact about what.

The absolute size of the fluctuation falls. Member kk’s beat swings by 2min(a,b)2\min(a,b), which for an octave on a 1/n1/n spectrum is 1/k1/k — the eighth member’s fluctuation is an eighth the size of the first’s in the pressure waveform. The energy in the beating term falls faster still, as the product of the two amplitudes, which is 1/2k21/2k^2. This is the direction every partial beats at its own rate is looking when it climbs a detuned pair’s spectrum: the rates rise and the components get quieter, and it is easy to read the second as the beats fading out.

That product is a real quantity and this collection computes it, and the trouble is that it is a product of two amplitudes, which is not a depth of anything. It was read as one. The prediction that the third and fourth members would come in one to two orders of magnitude below the first is a reading of that product with a steeper spectrum in mind than a string has: with partials falling as 1/n21/n^2 the product falls as k4k^{-4}, and the fourth member is then 256 times smaller than the first. The number is right about the product. The product is the wrong object.

A quieter beat that is fully modulated is still fully modulated. This is the whole of it, and it is worth stating as a principle rather than as a correction, because the same confusion is available in every ladder that puts a threshold on a fluctuation. The ear’s amplitude-modulation threshold is a ratio — about three per cent of the carrier, and the carrier here is the member’s own two partials, which have gone quiet together. The beat a tuner can actually use already found the same thing from the other end: it swept the threshold and reported that on a string spectrum the first interval is not lost until the threshold is 26.5 times the published one. A quantity with that much room in it does not become binding because the components got smaller.

The figure above says how much room there is. An octave’s index reaches 0.03 only at a spectral slope of 6.0 — a second partial 36 decibels down, which is not a string, not a plucked string, not a struck one, and not a bowed one. The instrument that does approach it is the one the seventh rung already named: a clarinet’s second partial is about 28 decibels down, its octave’s index is 0.077, and that is only 2.6 times the threshold. The depth criterion has teeth on a stopped pipe and none at all on wire.

The depth that does change is the depth inside one filter

There is a second depth, and it is the one the ladder should have been asking about.

A modulation index of two thirds is what the pair delivers on its own. A listener does not receive the pair on its own; a listener receives the output of an auditory filter centred near the coincidence, and that filter passes whatever else is nearby. Those neighbours are not modulated at this member’s rate. They arrive as level, and level in the denominator of a ratio makes the ratio smaller.

5 other partials in the filter, and 23 per cent of the depth gone. The auditory filter centred on member 4 of the beat family a octave mistuned by 6.0 cents makes at A3, and everything inside it. Above: the rounded-exponential filter at 1777 hertz, whose equivalent rectangular bandwidth is 217 hertz. Below: every partial of both notes in that neighbourhood, drawn as an open stem at its own amplitude and as a solid one at the amplitude the filter passes. 5 of them arrive at more than a fiftieth of the pair's own level. The two thick stems are the pair that beats, partial 8 of the lower note against partial 4 of the upper; they are the only two components in the filter fluctuating at 11.1 a second. The rest arrive steady and sum to 0.112, which is 23 per cent of the level in the channel and takes the modulation index from 0.667 to 0.513. Of the fluctuation the filter delivers, 40 per cent is at this member's own rate.
Fig. 3 The auditory filter at the 8:4 coincidence of a 6-cent octave at A3, and everything inside it. The two thick stems are the pair that beats at 11.1 a second; five other partials pass at more than a fiftieth of the pair’s level and sum to 0.112 against the pair’s 0.375. The index falls from 0.667 to 0.513 — not quite a quarter of it, rather than the two orders of magnitude expected.

The filter is Patterson’s rounded exponential, with its sharpness tied to the same equivalent rectangular bandwidth this collection’s resolvability machinery has used since it first asked which harmonics of a note are individually available: p=4f/ERB(f)p = 4f/\mathrm{ERB}(f), gain (1+pg)epg(1+pg)e^{-pg} at a fractional offset gg. At the 8:4 coincidence, 1777 hertz, that bandwidth is 217 hertz and the neighbours are 220 apart, so they sit right at its skirt and a little of each gets through.

The dilution across the whole family at A3 runs 0.662, 0.630, 0.575, 0.513, 0.454, and so on down — a fall of not quite a quarter over the four members the tenth rung admitted, against the factor of a hundred it expected. Every one of those numbers is between thirteen and twenty-two times the threshold at its own rate. The depth threshold removes nothing anywhere on a piano.

A buried beat is not a shallow one

So the first result of running the depth is that the recorded expectation was wrong. The second is more interesting, and it arrives from noticing what that pedestal actually is.

The neighbouring partials do not sit in the filter doing nothing. They beat — with each other, and with the pair, at rates of their own. The filter’s output is not one fluctuation diluted by a steady background; it is several fluctuations at several rates, and the question a listener faces is not how deep this member’s beat is but how much of what arrives at that place is this member’s beat.

That quantity is exact and needs no approximation. The squared envelope of a sum of components has a constant term plus one cosine per pair, at the difference of their frequencies and with amplitude twice the product of theirs, and the terms add up to the square of the summed amplitudes exactly. So the share of the filter’s fluctuation belonging to the member’s own rate can be read straight off.

The share falls through a half before the gap falls through a bandwidth. For each member of a octave's beat family at three registers, how much of the fluctuation its own auditory filter delivers is at that member's own rate — the rest belonging to the neighbouring partials, which beat with each other at rates of their own. The dashed line at 0.5 is the criterion: more than half the fluctuation in the channel is this member's. The open triangles mark where the same family's coincidence gap falls through one filter width, which is the criterion used throughout to call a member resolved. The triangles sit to the right of the crossings in every case: at A3 the share criterion bites at member 3.4 and the bandwidth criterion at 4.1. A buried member is not a shallow one — its depth is untouched — it is one whose filter is delivering somebody else's fluctuation as well, which is what roughness is.
Fig. 4 For each member of an octave’s family at three registers, how much of the fluctuation its own filter delivers is at that member’s rate. The dashed line is a half. The open triangles mark where the same family’s coincidence gap falls through one filter width, which is the criterion used earlier — and they sit to the right of the crossings in every case.

And here is the thing worth the essay. The criterion the tenth rung used to decide whether a member is separately available is a depth criterion in disguise. “The gap to the nearest neighbour exceeds one filter width” is a binary line drawn across this continuous quantity, and writing the same question in the depth’s own units gives a stricter answer at every register: at A3 the share falls through a half at member 3.4 while the gap falls through one bandwidth at 4.1; at A1 the two are 1.7 and 2.6; at A5, 3.8 and 4.1.

The two thresholds this ladder has been treating as separate — a place and a depth — are one threshold. A member the tenth rung called buried is not shallow at all: its own fluctuation is untouched, at two thirds. What has happened is that its filter is delivering somebody else’s fluctuation as well, which is the definition of roughness rather than of quiet.

So the count at A3 is two

Applying the share criterion where the bandwidth criterion sat changes the answer, and it changes it to the number the recorded expectation named.

Of 8 beats at A3, 3 can be attended to. Every member of the beat family a 2:1 mistuned by 6.0 cents makes on a real string at A3, placed by how separable its coincidence is from the partials beside it — in auditory filter widths, across — and by how fast it beats, up. The shaded region is the set a listener can receive: wider than one filter, and between 0.4 and 15 fluctuations a second. 4 of 8 members fall inside it, and once rates within a factor of 2 are counted as one modulation channel there are 3. The members that fail do so for two different reasons: the low ones sit under the roughness ceiling but their coincidences are buried in a filter that holds three partials, and the high ones are resolved and far too fast.
Fig. 5 The earlier picture of the same family: four members inside the shaded box, of which two beat at 1.26 and 0.89 a second and fall in one modulation channel, leaving three. The fourth member sits at 1.03 filter widths — three per cent inside the line.

Of the eight members at A3, three carry more than half their own filter’s fluctuation: the 2:1 at 1.26 beats a second, the 4:2 at 0.89 and the 6:3 at 2.71. The 8:4, at 0.399, does not. Then the modulation-channel rule the tenth rung applied — two rates are one fluctuation unless they differ by a factor of two — takes 1.26 against 0.89, a ratio of 1.42, and folds them together. What survives is a slow beat at nine tenths a second and a faster one at 2.71, and the answer is two.

At A1 it is worse than the tenth rung reported, and worse in the direction a low chord stops being a chord found from the other side of the same wall. The only member it admitted, the 4:2, carries 0.44 of its filter’s fluctuation, so on this criterion a mistuned octave at the bottom of a piano delivers no countable beat at all. At A5 nothing changes: the 2:1 carries 0.996 and is the only member slow enough, so the count stays at one.

That leaves an obvious objection, which is that a result reached by tightening a criterion is a result about the criterion. It is, and the honest way to quote it is to sweep.

The count at A3 survives a criterion 3 per cent tighter and no more. How many separable beats a octave mistuned by 6.0 cents delivers at three registers, against the number the resolution criterion is: a coincidence counts as resolved when the gap to its neighbours exceeds this many filter widths. The shaded band is the published range for that convention, one bandwidth to one and a quarter. A3 holds 3 beats only as far as 1.03 and then falls to 2, which is 3 per cent into a band twenty-five per cent wide — so the larger count is not a robust reading of the criterion, it is the value at one end of it. The dashed marks are what the same families give when the criterion is asked in the depth's own units instead, and each of them is a count the bandwidth criterion also reaches inside its own published range.
Fig. 6 How many separable beats a 6-cent octave delivers at three registers, against the number the resolution criterion is. The shaded band is the published range for that convention. A3 holds three beats only as far as 1.03 and then falls to two, which is a tenth of the way into a band a quarter wide.

The number three survives a criterion less than three per cent tighter than the one it was quoted at, which is a tenth of the way into a published range a quarter wide. The tenth rung stated that range — from about one bandwidth to about one and a quarter — swept it at the far end, and reported that the shape of the register dependence held. It does hold. What that sweep did not report is that the headline count crosses at 1.03, so a criterion a tenth of the way toward the conservative end already gives two. A5’s one survives the whole band; A1’s one is gone by 1.15. The count of three was never a robust reading.

That is the correction, and it is a correction to a number rather than to an argument. Everything the tenth rung says about why the count is register-dependent stands untouched, and so does its central observation that the bass and the treble lose their families for opposite reasons.

Every interval, every register, again

The grid the tenth rung drew over seven intervals and seven octaves can be recomputed cell by cell.

The map keeps its shape and loses 8 of its beats. How many separable beats each interval a tuner sets delivers at seven points up the compass, counted with a member admitted only when more than 50 per cent of the fluctuation in its auditory filter is its own. The small figure beside a cell is what the same cell held under the looser criterion of a gap wider than one filter width; 8 of the 49 cells fall. The column totals run 0, 2, 3, 7, 5, 0, 0 against 0, 5, 7, 8, 5, 0, 0 before, so the bass loses most of what it had and the treble loses nothing it did not already lack. What does not move is the shape: nothing at either end of the keyboard, and the maximum still at C4 — the octave every bearing plan in the craft is laid in.
Fig. 7 The same seven intervals at the same seven registers, counted with a member admitted only when more than half the fluctuation in its filter is its own. The bracketed figures are what each cell held before. The totals fall from 25 to 17 and eight cells change, and the maximum is still at C4.

The column totals run 0, 2, 3, 7, 5, 0, 0 against 0, 5, 7, 8, 5, 0, 0. C2 loses three of its five and C3 loses four of its seven; C4 loses one and C5 loses none. The bass is where the tightening bites, and it bites there for the reason the tenth rung already identified — below about 500 hertz the auditory filter is nearly constant in width while the family’s coincidences march up in equal steps of f0f_0, so the whole family is pressed into its own analysis bandwidth and every member is sharing a channel with two others.

What does not move is the shape, and the shape is what the practical claim rests on. Nothing at C1, nothing at C6 or above, a band in the middle, and the maximum at C4. A bearing plan is laid in the octave from the F below middle C to the F above, and that is where the evidence for it is under either criterion. Outside it a tuner works by octaves, which is where the stretch a piano is tuned to comes from and where counting has already stopped.

The repertoire that claim is about should be named, because it is a claim about practice and not about physics. The tuning sequences with published beat rates for each test interval at each pitch are a nineteenth- and twentieth-century apparatus, written for the equal-tempered piano and taught from printed tables to be counted against a watch — the sequences the essay on the fraction of a comma traces back to Werckmeister’s and Neidhardt’s directions, which give narrowings rather than rates. The grid says that a method built on counting could only ever have been built in the middle four octaves, and the tightening leaves that conclusion where it was while making the middle look less generous than it did.

Which computation produced the numbers

Each family is the set of coincidences a mistuned m:nm{:}n makes among sixteen partials, with the partials of a stiff string at nf01+Bn2nf_0\sqrt{1+Bn^2} and BB from this collection’s own string design, exactly as the ninth rung and the tenth computed them. The coincidence frequency is the mean of the two partials that make it.

The amplitudes are a power law, 1/ns1/n^s with s=1s = 1 throughout unless the figure says otherwise. This is a change from the two lists of eight the earlier rungs read from, and it is made deliberately: sixteen partials need sixteen amplitudes, and this collection has already found that its own rounded string list — 0.17 where a pure 1/n1/n wants 0.167 — reverses the major and minor sixths in the seventh rung’s own table. A slope has no rounding in it and the closed form above is exact against it.

The modulation index is the collection’s own, 2min(a,b)/(a+b)2\min(a,b)/(a+b), unchanged since the rung that introduced it.

The auditory filter is the roex(pp) of Patterson and Moore, (1+pg)epg(1+pg)e^{-pg} in power, with p=4fc/ERB(fc)p = 4f_c/\mathrm{ERB}(f_c) so that its equivalent rectangular bandwidth is Glasberg and Moore’s 24.7(0.00437f+1)24.7(0.00437f + 1) hertz by construction — the same function the resolvability machinery here has always used, now used as a shape rather than as a number.

The effective index is the same expression with the roex-weighted sum of every other component added to the level. The share is exact rather than approximated: the intensity of a filtered sum is iwi2\sum_i w_i^2 plus 2wiwjcos(Δωijt)2w_iw_j\cos(\Delta\omega_{ij}t) over every pair, those terms sum to (iwi)2(\sum_i w_i)^2 identically, and the share is the pair’s own term over all the fluctuating ones. That identity is checked in the drawing rather than trusted.

The threshold is the temporal modulation transfer function the eighth rung brought in, mthr(f)=0.031+(f/50)2m_{\mathrm{thr}}(f) = 0.03\sqrt{1 + (f/50)^2}, flat below about fifty fluctuations a second and rising above.

A member counts when its share exceeds a half, its rate lies between 0.4 and 15 fluctuations a second, and its effective index exceeds the threshold at that rate. Rates within a factor of two are one modulation channel. The interval widths in the grid are equal temperament’s own departures from just, with the octave at its 2:1 stretch.

One of those numbers is not the one the tenth rung’s account of itself gives. The ceiling above which a fluctuation is roughness rather than a beat is written there as twenty a second; the constant this collection actually computes with, and which produced every figure in both essays, is fifteen. Fifteen is used here, so the two grids are comparable cell by cell. The difference is not quite nothing: exactly one cell in the whole compass holds a member between the two ceilings that passes every other test — the minor third at C4, whose 6:5 coincidence sits at 1.33 filter widths and beats 18.8 times a second — so under the ceiling the tenth rung states, its C4 column totals nine rather than the eight it quotes. Nothing else in either grid moves, and neither does the location of the maximum. It is worth writing down because it is the kind of gap a sweep of the other criterion cannot find.

Where the model stops

The pedestal is treated as level rather than as sound. The neighbours are added as amplitudes into the denominator, which is the worst case — it assumes they never help the fluctuation — and their own modulations are removed from the numerator, which is right for the share and approximate for the index. The share is the trustworthy half of the pair, because it is an identity rather than an approximation.

The filter shape has no level in it. The auditory filter broadens as the sound gets louder, and a tuner does not strike quietly. Every number here is for a filter at a moderate level, and a wider filter would lower every share and every count. Nothing in this arithmetic is quoted at a sound pressure at all.

The equal-amplitude case is untouched and is still the important one. A unison pairs partial one against partial one on two nominally identical sources, so its index is exactly one at every member, and that is the case three strings on one bridge is about. Nothing above changes it; the invariant simply says the same thing about it that it says about everything else.

A half is a criterion. So was one bandwidth. The whole content of the comparison is that the second is stricter than the first, and the sweep is what the number should be quoted with — three at A3 is a criterion of 1.00 to 1.03, and two is everything from there to 1.33.

One slope for both notes and for the whole compass. A real piano’s partials fall away faster in the treble than in the bass and faster still under a soft hammer, and the closed form says exactly what that does: it moves the index down the curve, never sideways along a family.

And there is still no listener. Every quantity here is a property of a signal delivered to a bank of filters. What anybody does with it is a different question and this ladder has not asked it once in eleven rungs.

What the picture cannot show

It cannot show the modulation filterbank. The factor of two between rates stands in for a set of channels tuned to modulation frequency, and the share says how much energy arrives at one rate without saying whether a channel tuned to it would win against the others. That is the difference between a spectrum and a detector.

It cannot show the attack. A piano note’s partials decay at different rates, so the amplitudes that fix the index are only the ones at the moment they are read, and the ratio of two partials is not constant through a note. The index is a constant of the spectrum, and a spectrum that changes moves it.

Nor can it show the phase. Whether the neighbours reinforce or oppose the pair at any instant depends on phases nothing here tracks, and a real envelope wanders where this arithmetic gives a mean.

And it cannot show what two beats at once are like. The claim is that two arrive separately, not that anybody hears them as two. A tuner listening to a mistuned octave at A3 is listening to a texture whose composition this says, and whose sound it does not.

Where this ladder goes next

Eleven rungs. Beats are arithmetic anybody can hear; a tuner counts them; a cellist’s wolf is the same arithmetic coupled; a piano’s unison sits below a bifurcation and a chorus above it; every partial beats at its own rate; a beat has a depth as well as a rate; the rate decides which intervals are usable; a single mistuning makes a whole family; how many of that family arrive separately; and now that the depth of every member of a family is identical, that the criterion which decides separateness was a depth criterion all along, and that the count in the middle of the keyboard is two.

What is owed next is the detector, and the shape of the debt is exact. The share computed here is a modulation spectrum: for each member’s filter it gives the amplitude at every rate present, from the pair’s own to the ones its neighbours make with each other. What decides audibility is not that spectrum but the output of a channel tuned to the member’s rate, and the modulation filterbank this ladder has been standing in for with a factor of two has a published shape — channels of quality factor near one, spaced logarithmically. Running the within-filter modulation spectrum through such a bank would replace the factor of two with a computed ratio and would answer the question the ninth rung asked and none of the rungs above it has: whether two fluctuations at 0.89 and 1.26 a second are one thing or two, at that depth. It is arithmetic on quantities this collection already holds, and it needs no listener.

A second debt sits beside it and is smaller. The auditory filter’s sharpness falls with level, so every share here is quoted for a moderately loud note and every count would fall for a loud one. Since a tuner strikes hard, that is not a refinement: it is a claim that the number of beats a mistuned octave delivers depends on how the note was struck. Running it needs one published coefficient this collection does not yet carry, and after that it is arithmetic too.

Part 11 of 16

One essay in the series on beating. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

AmplitudeBeatingCritical bandwidthDetectionModulationPartialPianoResolvabilitySpectrumTuning by ear