Every member of a beat family is the same depth
Assumes: Beats are arithmetic that anybody can hear · Three beats at most, and only in the middle of the keyboard
Three beats at most counted the members of a mistuned octave’s beat family twice — once by whether each fluctuation is separable in the cochlea from the partials beside it, and once by whether it is slow enough to wait for — and ended by naming the quantity it had counted them without. A fluctuation is audible only if it is deep enough as well — beats are arithmetic anybody can hear only while there is something to hear — and the depth of the third and fourth members at A3 was expected to be one or two orders of magnitude below the first, because the amplitudes that make them fall away up the spectrum. Applying a threshold to that fall would take the answer from four members to two.
The depth is now computed, and the first half of that expectation is wrong by about the largest margin an expectation can be wrong by. It does not fall.
The depth does not know which member it is
The arithmetic is two lines and it was available the whole time.
A beat has a depth fixed what the depth of a beat is: two components of amplitude and make an envelope swinging between and , so the modulation index — the quantity a demodulator reports and the quantity every published detection threshold is stated on — is . It depends on the two amplitudes only through their ratio. Scale both by any factor and the index is unchanged, because a fluctuation of a tenth of a quiet thing is the same fluctuation as a tenth of a loud one.
Now take the family. The -th member of a mistuned pairs the lower note’s partial with the upper note’s partial . On a spectrum whose partials fall as those two amplitudes are and , and their ratio is
in which has cancelled. Every member of every family therefore has the same modulation index, and it is
a number set by the interval and the spectral slope and by nothing else at all. For an octave on a one-over- string spectrum it is two thirds, which is a dip of 9.5 decibels — the exact figure in the octave row of the table the seventh rung computed from a single pair of partials, arrived at here from eight pairs at once.
That is a stronger claim than the seventh rung was in a position to make. Its table of seven intervals was computed from the coincidence a tuner is said to listen to — the lower note’s -th partial against the upper note’s -th — and treated as one row per interval. It is one row per family. A minor third’s beat is 20.8 decibels deep at its 6:5 coincidence and 20.8 decibels deep at its 12:10 and its 18:15, and the ordering of the seven intervals is the ordering of every member of all seven. That matters to the tuner counting seconds against a watch, who is choosing a member as well as an interval and has been told only about the first.
What falls as a power of the member number, and what does not
Something does fall away up a family, and it is worth being exact about what.
The absolute size of the fluctuation falls. Member ’s beat swings by , which for an octave on a spectrum is — the eighth member’s fluctuation is an eighth the size of the first’s in the pressure waveform. The energy in the beating term falls faster still, as the product of the two amplitudes, which is . This is the direction every partial beats at its own rate is looking when it climbs a detuned pair’s spectrum: the rates rise and the components get quieter, and it is easy to read the second as the beats fading out.
That product is a real quantity and this collection computes it, and the trouble is that it is a product of two amplitudes, which is not a depth of anything. It was read as one. The prediction that the third and fourth members would come in one to two orders of magnitude below the first is a reading of that product with a steeper spectrum in mind than a string has: with partials falling as the product falls as , and the fourth member is then 256 times smaller than the first. The number is right about the product. The product is the wrong object.
A quieter beat that is fully modulated is still fully modulated. This is the whole of it, and it is worth stating as a principle rather than as a correction, because the same confusion is available in every ladder that puts a threshold on a fluctuation. The ear’s amplitude-modulation threshold is a ratio — about three per cent of the carrier, and the carrier here is the member’s own two partials, which have gone quiet together. The beat a tuner can actually use already found the same thing from the other end: it swept the threshold and reported that on a string spectrum the first interval is not lost until the threshold is 26.5 times the published one. A quantity with that much room in it does not become binding because the components got smaller.
The figure above says how much room there is. An octave’s index reaches 0.03 only at a spectral slope of 6.0 — a second partial 36 decibels down, which is not a string, not a plucked string, not a struck one, and not a bowed one. The instrument that does approach it is the one the seventh rung already named: a clarinet’s second partial is about 28 decibels down, its octave’s index is 0.077, and that is only 2.6 times the threshold. The depth criterion has teeth on a stopped pipe and none at all on wire.
The depth that does change is the depth inside one filter
There is a second depth, and it is the one the ladder should have been asking about.
A modulation index of two thirds is what the pair delivers on its own. A listener does not receive the pair on its own; a listener receives the output of an auditory filter centred near the coincidence, and that filter passes whatever else is nearby. Those neighbours are not modulated at this member’s rate. They arrive as level, and level in the denominator of a ratio makes the ratio smaller.
The filter is Patterson’s rounded exponential, with its sharpness tied to the same equivalent rectangular bandwidth this collection’s resolvability machinery has used since it first asked which harmonics of a note are individually available: , gain at a fractional offset . At the 8:4 coincidence, 1777 hertz, that bandwidth is 217 hertz and the neighbours are 220 apart, so they sit right at its skirt and a little of each gets through.
The dilution across the whole family at A3 runs 0.662, 0.630, 0.575, 0.513, 0.454, and so on down — a fall of not quite a quarter over the four members the tenth rung admitted, against the factor of a hundred it expected. Every one of those numbers is between thirteen and twenty-two times the threshold at its own rate. The depth threshold removes nothing anywhere on a piano.
A buried beat is not a shallow one
So the first result of running the depth is that the recorded expectation was wrong. The second is more interesting, and it arrives from noticing what that pedestal actually is.
The neighbouring partials do not sit in the filter doing nothing. They beat — with each other, and with the pair, at rates of their own. The filter’s output is not one fluctuation diluted by a steady background; it is several fluctuations at several rates, and the question a listener faces is not how deep this member’s beat is but how much of what arrives at that place is this member’s beat.
That quantity is exact and needs no approximation. The squared envelope of a sum of components has a constant term plus one cosine per pair, at the difference of their frequencies and with amplitude twice the product of theirs, and the terms add up to the square of the summed amplitudes exactly. So the share of the filter’s fluctuation belonging to the member’s own rate can be read straight off.
And here is the thing worth the essay. The criterion the tenth rung used to decide whether a member is separately available is a depth criterion in disguise. “The gap to the nearest neighbour exceeds one filter width” is a binary line drawn across this continuous quantity, and writing the same question in the depth’s own units gives a stricter answer at every register: at A3 the share falls through a half at member 3.4 while the gap falls through one bandwidth at 4.1; at A1 the two are 1.7 and 2.6; at A5, 3.8 and 4.1.
The two thresholds this ladder has been treating as separate — a place and a depth — are one threshold. A member the tenth rung called buried is not shallow at all: its own fluctuation is untouched, at two thirds. What has happened is that its filter is delivering somebody else’s fluctuation as well, which is the definition of roughness rather than of quiet.
So the count at A3 is two
Applying the share criterion where the bandwidth criterion sat changes the answer, and it changes it to the number the recorded expectation named.
Of the eight members at A3, three carry more than half their own filter’s fluctuation: the 2:1 at 1.26 beats a second, the 4:2 at 0.89 and the 6:3 at 2.71. The 8:4, at 0.399, does not. Then the modulation-channel rule the tenth rung applied — two rates are one fluctuation unless they differ by a factor of two — takes 1.26 against 0.89, a ratio of 1.42, and folds them together. What survives is a slow beat at nine tenths a second and a faster one at 2.71, and the answer is two.
At A1 it is worse than the tenth rung reported, and worse in the direction a low chord stops being a chord found from the other side of the same wall. The only member it admitted, the 4:2, carries 0.44 of its filter’s fluctuation, so on this criterion a mistuned octave at the bottom of a piano delivers no countable beat at all. At A5 nothing changes: the 2:1 carries 0.996 and is the only member slow enough, so the count stays at one.
That leaves an obvious objection, which is that a result reached by tightening a criterion is a result about the criterion. It is, and the honest way to quote it is to sweep.
The number three survives a criterion less than three per cent tighter than the one it was quoted at, which is a tenth of the way into a published range a quarter wide. The tenth rung stated that range — from about one bandwidth to about one and a quarter — swept it at the far end, and reported that the shape of the register dependence held. It does hold. What that sweep did not report is that the headline count crosses at 1.03, so a criterion a tenth of the way toward the conservative end already gives two. A5’s one survives the whole band; A1’s one is gone by 1.15. The count of three was never a robust reading.
That is the correction, and it is a correction to a number rather than to an argument. Everything the tenth rung says about why the count is register-dependent stands untouched, and so does its central observation that the bass and the treble lose their families for opposite reasons.
Every interval, every register, again
The grid the tenth rung drew over seven intervals and seven octaves can be recomputed cell by cell.
The column totals run 0, 2, 3, 7, 5, 0, 0 against 0, 5, 7, 8, 5, 0, 0. C2 loses three of its five and C3 loses four of its seven; C4 loses one and C5 loses none. The bass is where the tightening bites, and it bites there for the reason the tenth rung already identified — below about 500 hertz the auditory filter is nearly constant in width while the family’s coincidences march up in equal steps of , so the whole family is pressed into its own analysis bandwidth and every member is sharing a channel with two others.
What does not move is the shape, and the shape is what the practical claim rests on. Nothing at C1, nothing at C6 or above, a band in the middle, and the maximum at C4. A bearing plan is laid in the octave from the F below middle C to the F above, and that is where the evidence for it is under either criterion. Outside it a tuner works by octaves, which is where the stretch a piano is tuned to comes from and where counting has already stopped.
The repertoire that claim is about should be named, because it is a claim about practice and not about physics. The tuning sequences with published beat rates for each test interval at each pitch are a nineteenth- and twentieth-century apparatus, written for the equal-tempered piano and taught from printed tables to be counted against a watch — the sequences the essay on the fraction of a comma traces back to Werckmeister’s and Neidhardt’s directions, which give narrowings rather than rates. The grid says that a method built on counting could only ever have been built in the middle four octaves, and the tightening leaves that conclusion where it was while making the middle look less generous than it did.
Which computation produced the numbers
Each family is the set of coincidences a mistuned makes among sixteen partials, with the partials of a stiff string at and from this collection’s own string design, exactly as the ninth rung and the tenth computed them. The coincidence frequency is the mean of the two partials that make it.
The amplitudes are a power law, with throughout unless the figure says otherwise. This is a change from the two lists of eight the earlier rungs read from, and it is made deliberately: sixteen partials need sixteen amplitudes, and this collection has already found that its own rounded string list — 0.17 where a pure wants 0.167 — reverses the major and minor sixths in the seventh rung’s own table. A slope has no rounding in it and the closed form above is exact against it.
The modulation index is the collection’s own, , unchanged since the rung that introduced it.
The auditory filter is the roex() of Patterson and Moore, in power, with so that its equivalent rectangular bandwidth is Glasberg and Moore’s hertz by construction — the same function the resolvability machinery here has always used, now used as a shape rather than as a number.
The effective index is the same expression with the roex-weighted sum of every other component added to the level. The share is exact rather than approximated: the intensity of a filtered sum is plus over every pair, those terms sum to identically, and the share is the pair’s own term over all the fluctuating ones. That identity is checked in the drawing rather than trusted.
The threshold is the temporal modulation transfer function the eighth rung brought in, , flat below about fifty fluctuations a second and rising above.
A member counts when its share exceeds a half, its rate lies between 0.4 and 15 fluctuations a second, and its effective index exceeds the threshold at that rate. Rates within a factor of two are one modulation channel. The interval widths in the grid are equal temperament’s own departures from just, with the octave at its 2:1 stretch.
One of those numbers is not the one the tenth rung’s account of itself gives. The ceiling above which a fluctuation is roughness rather than a beat is written there as twenty a second; the constant this collection actually computes with, and which produced every figure in both essays, is fifteen. Fifteen is used here, so the two grids are comparable cell by cell. The difference is not quite nothing: exactly one cell in the whole compass holds a member between the two ceilings that passes every other test — the minor third at C4, whose 6:5 coincidence sits at 1.33 filter widths and beats 18.8 times a second — so under the ceiling the tenth rung states, its C4 column totals nine rather than the eight it quotes. Nothing else in either grid moves, and neither does the location of the maximum. It is worth writing down because it is the kind of gap a sweep of the other criterion cannot find.
Where the model stops
The pedestal is treated as level rather than as sound. The neighbours are added as amplitudes into the denominator, which is the worst case — it assumes they never help the fluctuation — and their own modulations are removed from the numerator, which is right for the share and approximate for the index. The share is the trustworthy half of the pair, because it is an identity rather than an approximation.
The filter shape has no level in it. The auditory filter broadens as the sound gets louder, and a tuner does not strike quietly. Every number here is for a filter at a moderate level, and a wider filter would lower every share and every count. Nothing in this arithmetic is quoted at a sound pressure at all.
The equal-amplitude case is untouched and is still the important one. A unison pairs partial one against partial one on two nominally identical sources, so its index is exactly one at every member, and that is the case three strings on one bridge is about. Nothing above changes it; the invariant simply says the same thing about it that it says about everything else.
A half is a criterion. So was one bandwidth. The whole content of the comparison is that the second is stricter than the first, and the sweep is what the number should be quoted with — three at A3 is a criterion of 1.00 to 1.03, and two is everything from there to 1.33.
One slope for both notes and for the whole compass. A real piano’s partials fall away faster in the treble than in the bass and faster still under a soft hammer, and the closed form says exactly what that does: it moves the index down the curve, never sideways along a family.
And there is still no listener. Every quantity here is a property of a signal delivered to a bank of filters. What anybody does with it is a different question and this ladder has not asked it once in eleven rungs.
What the picture cannot show
It cannot show the modulation filterbank. The factor of two between rates stands in for a set of channels tuned to modulation frequency, and the share says how much energy arrives at one rate without saying whether a channel tuned to it would win against the others. That is the difference between a spectrum and a detector.
It cannot show the attack. A piano note’s partials decay at different rates, so the amplitudes that fix the index are only the ones at the moment they are read, and the ratio of two partials is not constant through a note. The index is a constant of the spectrum, and a spectrum that changes moves it.
Nor can it show the phase. Whether the neighbours reinforce or oppose the pair at any instant depends on phases nothing here tracks, and a real envelope wanders where this arithmetic gives a mean.
And it cannot show what two beats at once are like. The claim is that two arrive separately, not that anybody hears them as two. A tuner listening to a mistuned octave at A3 is listening to a texture whose composition this says, and whose sound it does not.
Where this ladder goes next
Eleven rungs. Beats are arithmetic anybody can hear; a tuner counts them; a cellist’s wolf is the same arithmetic coupled; a piano’s unison sits below a bifurcation and a chorus above it; every partial beats at its own rate; a beat has a depth as well as a rate; the rate decides which intervals are usable; a single mistuning makes a whole family; how many of that family arrive separately; and now that the depth of every member of a family is identical, that the criterion which decides separateness was a depth criterion all along, and that the count in the middle of the keyboard is two.
What is owed next is the detector, and the shape of the debt is exact. The share computed here is a modulation spectrum: for each member’s filter it gives the amplitude at every rate present, from the pair’s own to the ones its neighbours make with each other. What decides audibility is not that spectrum but the output of a channel tuned to the member’s rate, and the modulation filterbank this ladder has been standing in for with a factor of two has a published shape — channels of quality factor near one, spaced logarithmically. Running the within-filter modulation spectrum through such a bank would replace the factor of two with a computed ratio and would answer the question the ninth rung asked and none of the rungs above it has: whether two fluctuations at 0.89 and 1.26 a second are one thing or two, at that depth. It is arithmetic on quantities this collection already holds, and it needs no listener.
A second debt sits beside it and is smaller. The auditory filter’s sharpness falls with level, so every share here is quoted for a moderately loud note and every count would fall for a loud one. Since a tuner strikes hard, that is not a refinement: it is a claim that the number of beats a mistuned octave delivers depends on how the note was struck. Running it needs one published coefficient this collection does not yet carry, and after that it is arithmetic too.
Part 11 of 16
One essay in the series on beating. The essays either side of this one:
What links here
Essays that reach for this one mid-argument — the half of a link its own author cannot write down.
What this makes readable
Essays that declare this one a prerequisite.
The objects named here
The third way in, after the field and the series: the things themselves, and every essay that touches each one.
AmplitudeBeatingCritical bandwidthDetectionModulationPartialPianoResolvabilitySpectrumTuning by ear
- A clarinet keeps what a string loses critical bandwidth, partial, spectrum
- A firm touch buys beats until the aftersound sinks with it beating, piano, tuning by ear
- A roughness with a rate of its own beating, critical bandwidth, partial
- A section against another section beating, critical bandwidth, partial
- A string that decays twice is counted early beating, partial, tuning by ear
- Eleven partials is one too many critical bandwidth, partial, resolvability