Perception and the listener

Three answers to how finely a pitch can be heard

Two notes one after the other are told apart at about four cents at A440. Whether a melodic interval is in tune is a judgement an order of magnitude coarser. And two notes held together are heard to beat at a third of a cent, because the question is answered by counting rather than by hearing pitch at all. Every equal division ever built sits between the coarsest and the finest.

Assumes: How small a difference is audible · A scale is not a set of pitches

“How small a pitch difference can be heard” is asked as though it had one answer. It has at least three, they are an order of magnitude apart, and which one applies depends entirely on how the two pitches are presented.

Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.
Fig. 1 Three resolutions across five octaves, on a logarithmic scale of cents, with the step sizes of four equal divisions drawn across them. The three are not close together and no notation is anywhere near the finest.

One after the other: about four cents

The difference limen for frequency is the classical measurement: two pure tones in succession, one slightly higher, and the listener says which. This site has drawn it before and the number is around 4 cents at A440, rising to 8.6 cents three octaves down at A110 and settling near 3.4 in the treble.

Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.
Fig. 2 The three across seven octaves rather than five, which is where the ratio between them is worth reading. Successive tones are told apart at 4.0 cents at A440 and 8.6 three octaves down; beat counting is 2.6 at A440 and 20.9 at A0. The gap between hearing and counting runs from 1.2 in the bass to twenty-five times that in the treble — so the advantage a tuner has over a listener is not a constant, it is a function of register, and it is smallest exactly where a bass line lives.

That is a threshold in the strict sense — the point at which a forced choice becomes better than guessing — and it is measured in laboratory conditions on pure tones with no other sound present. It is the finest of the three for successive tones, and it is not the finest overall.

In a melody: twenty-five to fifty cents

Ask a different question. Play an interval — two notes, one after the other, separated by roughly a fifth — and ask whether it is in tune. That is not a discrimination task; it is a categorisation task, and categorisation is coarse by design.

Drawn as identification curves, a melodic interval’s categories are 100 cents apart with a crossing about 24 cents wide, which is where the middle answer comes from: a melodic judgement is not a comparison of two frequencies but an assignment to a box.

The published range for how far a melodic interval can be mistuned before trained listeners reliably call it wrong is roughly 25 to 50 cents, depending on the interval, the training and the context. That is between six and twelve times the discrimination threshold at the same frequency.

The gap is not a contradiction. Discrimination asks are these two things different; interval judgement asks is this the expected interval, and the second is answered against a learned category with a width. The category is the theory, and a category wide enough to survive singers, vibrato and instrumental error has to be far wider than a threshold.

Held together: a third of a cent

Now play the two notes at the same time and the question changes completely.

Two tones a fifth apart, slightly mistuned, produce beats between the third partial of the lower and the second of the upper. The beat rate is the difference between those two partial frequencies, and a beat is audible as a beat down to roughly one a second — slower than that it becomes a swell to be waited out rather than a rhythm to be counted, and a tuner’s printed instructions are given in beats per second for exactly that reason.

Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.
Fig. 3 The divisions somebody has actually proposed, against the three thresholds. Twelve equal’s step is 100 cents — coarser than every one of them; fifty-three’s is 22.6 and seventy-two’s 16.7, both above the discrimination threshold and far above the beat-counting one. So a finer division does not let a musician play more accurately in any sense the ear can check: a performer with continuous pitch is already placing notes more finely than any of these specify, and one with fixed pitches is placing them exactly where the frets are.

Set the beat rate to one a second and solve for the mistuning. At A440 the answer is 1.3 cents; at A880 it is 0.66; at A1760 it is 0.33. And the limit is not the ear at all — it is patience. Wait two seconds for a beat instead of one and every number halves.

This is how instruments are actually tuned. A tuner counts beats precisely because the beat-rate method is two orders of magnitude finer than pitch discrimination and requires no memory for pitch whatever. The person tuning a piano is not comparing two pitches; they are counting a rhythm.

Why the three are not the same measurement made three ways

It would be tidy if the three numbers were one threshold seen through three amounts of noise. They are not, and the clearest evidence is that they scale differently with frequency.

Discrimination improves with frequency and then flattens: 8.6 cents at A110, 4.0 at A440, 3.4 at A880 and 3.6 at A1760, so in cents it varies by a factor of two and a half across the range and in hertz it varies by a factor of twenty.

Beat counting improves without limit: the mistuning that gives one beat every two seconds halves with every octave, because the beat rate is a difference in hertz and the same number of cents is twice as many hertz an octave up. At A110 it is 5.2 cents and at A1760 it is 0.33.

And melodic-interval judgement is flat in cents, near enough, because it is a judgement about a ratio and a category, and neither has a frequency in it.

Three quantities with three different dependences on frequency are three quantities. A single underlying resolution seen through different amounts of noise would scale the same way in all three cases.

And the subjective octave stretch is a fourth quantity of the same order — 8 cents at 125 hertz and 30 at 4,000 — which is not a threshold at all but a systematic departure, and it is larger than the discrimination limit everywhere.

Where the notations sit

Now put the equal divisions on the same axis.

Three answers to how finely a pitch can be heard. Three resolutions across five octaves, on a logarithmic scale of cents. Two notes one after the other are told apart at 4.0 cents at A440 and 8.6 cents three octaves down. Whether a melodic interval is in tune is a judgement an order of magnitude coarser, 25 to 50 cents. And two notes held a fifth apart are heard to beat once every 2 seconds at 1.31 cents, which is finer than either. The horizontal lines are the step sizes of the equal divisions that have been built: 12 at 100.0 cents, 24 at 50.0 cents, 53 at 22.6 cents, 72 at 16.7 cents. Every one of them is coarser than discrimination and finer than melodic judgement.
Fig. 4 The middle answer with the note lengths that produce it, which is the variable the other two do not have. A melodic judgement is made on notes of some duration, and a note too short to specify its own frequency cannot be judged more finely than its length allows — so the 25-to-50-cent band is partly a fact about how long musical notes are rather than only about categories. Two seconds of note would give a much finer band than any melody supplies, and no melody supplies it.

Twelve equal is 100 cents a step. Twenty-four is 50. Fifty-three is 22.6. Seventy-two — the finest in any real use, in some Byzantine chant theory and in a handful of twentieth-century scores — is 16.7.

Every one of them is coarser than discrimination and finer than or equal to melodic judgement. Not one of them is anywhere near the beat-counting resolution.

To notate at the discrimination limit would take about 150 equal steps — 1200 divided by twice the four-cent limen at A440 — so that any target pitch is within half a limen of some step. Nobody has proposed such a system, and the reason is obvious once the three numbers are separated: a notation is read by a person playing a melody, and melodic judgement is the resolution that matters for that.

Every equal division from five to eighty, with its best fifth and third, says which divisions are worth building — and the answer is a tuning answer rather than a perceptual one: nineteen and thirty-one have better thirds, measured against just ratios and not against any threshold.

A fourth answer, briefly, and why it is not on the axis

There is a resolution finer than all three and it belongs to a different sense of the word.

Two notes held together at a simple ratio do not merely beat when mistuned; they lose a quality — the fused, single-object character that a pure octave or a pure fifth has and a mistuned one does not. The auditory-scene ladder measures the same effect from the other side: a partial mistuned by about one per cent stops belonging to its own tone and is heard out as a separate whistle.

One per cent is seventeen cents, which is far coarser than beat counting. But the two are measuring different things — one is about noticing a modulation and the other about a partial ceasing to be part of an object — and neither is a pitch judgement. Both are on the plot only in the sense that they can be converted to cents.

That conversion is the trap this whole essay is about. Anything can be converted to cents, and having a number in cents does not make two quantities comparable. The three resolutions here are plotted together to show how far apart they are, not to suggest that they measure one thing.

Which threshold a comma has to beat

The three resolutions decide different arguments, and the tuning ladder has already used two of them without naming the distinction.

The schisma is 1.95 cents — below the discrimination limen everywhere, so as a pitch difference it is inaudible; but two notes a schisma apart sounded together at A440 beat once every 1.5 seconds, which is comfortably countable. So the schisma is inaudible and measurable at the same time, and which of those is true depends on whether the two notes are played one after the other or together.

The Pythagorean comma is 23.5 cents: above the limen at every frequency, and about the width of one step of 53-equal. The syntonic comma is 21.5. Both are audible as pitch differences and both are far below the melodic-interval category width — which is exactly why temperament works at all: the intervals are wrong by more than the ear can detect in isolation and by less than the ear’s categories care about.

Nine commas, and the line under which none of them matters. Each named comma at its true size in cents, against the difference limen of 5.10 cents at 261.626 Hz — the smallest change of frequency a listener can detect, computed from the same formula the perception essays use. One comma falls below it, and it is the only gap in the subject that no tuning system has to do anything about.
Fig. 5 The commas of tuning theory ranked by size, with what each one costs. Read against the three resolutions: everything here is above the discrimination limen except the schisma, everything is below the melodic-category width, and every one of them is enormous by the beat-counting standard.

The consequence for a division nobody has built

Put the three numbers together and a question that gets argued about becomes arithmetic.

Would a finer equal division let a musician play more accurately? Not in any sense the ear can check. Twelve equal already asks for placement well inside a melodic category, and the categories are 25 to 50 cents wide. A performer on a continuously variable instrument is already placing pitches more finely than any division specifies; a performer on a fixed-pitch instrument is placing them exactly where the frets or keys are, and adding more of them changes which pitches are available rather than how accurately they are hit.

Would a finer division sound better in tune? That is a different question and the answer is yes, up to a point, but the point is a tuning one rather than a perceptual one — nineteen and thirty-one have better thirds and the improvement is measured against just ratios, not against a threshold.

Is there a division fine enough that no interval is audibly mistuned? By the discrimination standard, about 150 steps. By the beat-counting standard, there is no such division at all, because the requirement has no floor. The second answer is the operative one for a sustained chord and is the reason just intonation keeps being reinvented: the error a temperament leaves is always audible if two notes are held long enough.

A notation is a compromise between three thresholds and it has never been in the middle of them. It sits at the coarse end, because that is where the person reading it is working.

How much counting buys over hearing, register by register

The three resolutions are quoted at A440 and all three move with frequency, so the interesting quantity is not any one of them but the ratio — how much better off a listener is counting beats than simply hearing a difference. It varies by a factor of twenty-five across the compass.

discrimination beat counting ratio
A0, 27.5 Hz 25.1 cents 20.9 1.20
A1, 55 14.3 10.5 1.37
A2, 110 8.6 5.2 1.65
A3, 220 5.6 2.6 2.13
A4, 440 4.0 1.3 3.08
A5, 880 3.4 0.66 5.19
A6, 1760 3.6 0.33 10.83
A7, 3520 5.0 0.16 30.63

Counting is finer than hearing at every frequency, so the two never cross and there is no register in which the beat method is the worse one. But the margin is not remotely constant. In the top octave of a piano counting is thirty times finer; in the bottom octave it is finer by twenty per cent, which is to say by nothing a person could rely on.

That is a computable account of something every tuner reports. Down at the bottom A, a fifth mistuned enough to beat once a second is mistuned by 20.9 cents — and 25.1 cents is what the ear can simply hear down there, so the countable target is barely finer than the thing it was meant to replace. The bass is also where a beat takes longest to count, because the same beat rate is a smaller fraction of the note’s own decay. The two difficulties compound, and the practice they produce is the standard one: set the temperament in the middle octaves, where the ratio is three or five, and tune the bass by octaves from what is already set.

The same table corrects a number from earlier in this essay. Notating at the discrimination limit was put at about 150 equal steps, from the four-cent limen at A440 — but the limen is finest in the treble, at 3.4 cents, so a uniform division that stays inside half a limen everywhere needs 177. Matched to the bass instead it would need 70, and no equal division can be both. A notation fine enough for the ear at one end of the compass is two and a half times finer than it needs to be at the other, which is one more reason the finest division anybody built stopped at seventy-two.

What the picture cannot show

The middle number is the weakest. The difference limen and the beat-rate limit are both computed here from stated models — one an empirical fit that this site plots, the other a line of arithmetic. The 25-to-50-cent range for melodic interval judgement is quoted from the literature as a range, and it varies with the interval, with training, with whether the context is tonal and with what the listener expects. It is drawn as a band because it is a band.

All three assume steady tones. A real note has vibrato of anything from 20 to 100 cents peak to peak, a pitch that moves during the attack, and inharmonic partials that give it more than one pitch to be. Every threshold here is measured on something no instrument produces.

And the beat-rate limit is not a limit of the ear. It is where this essay decided to stop waiting. The arithmetic has no floor: a beat every ten seconds is 0.26 cents at A440 and a beat every minute is 0.04, and both are detectable by a patient person with a stopwatch. Calling it a perceptual resolution is a convenience, and the honest statement is that simultaneous mistuning has no threshold — only a time cost.

The numbers, in one place

For reference, at A440 and in cents:

  • Discrimination, two successive pure tones: about 4.
  • Melodic interval judgement, whether an interval is in tune: 25 to 50.
  • Beat counting, two tones held a fifth apart at one beat a second: 1.3. At one beat every two seconds: 0.66.
  • One step of twelve equal: 100. Of twenty-four: 50. Of fifty-three: 22.6. Of seventy-two: 16.7.
  • The Pythagorean comma: 23.5. The syntonic comma: 21.5. The schisma: 2.0.
  • The stretch a listener prefers on an octave: about 15 at the extremes of the range.

Reading down that list is the argument. There is no single line at which “audible” begins, four of the six quantities in the middle are within a factor of three of each other, and the two ends are separated by a factor of forty.

Whose music this is a claim about

The three resolutions map onto three activities and the mapping is clean.

Tuning uses the finest. Every historical temperament is specified, in the sources, as a set of beat rates — so many beats per second on this fifth, none on that third — and the reason is that the alternative would have been unusable. A tuner told to make a fifth “two cents narrow” has been given an instruction two limens fine and no way to check it; told to make it beat once a second, they have a countable target. The whole method is a rhythm, and it is the only place in music where a pitch is set by a clock.

Performing uses the middle one. A singer or a violinist places a pitch within the category, and the evidence that this is what they are doing is that measured performance intonation is scattered over tens of cents while being unhesitatingly heard as correct.

And listening for a difference uses the discrimination threshold for successive notes, which is why the octave-stretch of a piano is audible as a preference and not as a mistuning: seventeen cents is four limens and about half a category — large enough to be heard as a difference and small enough not to be heard as a wrong note.

The traditions this ladder is about complicate the middle number and not the other two. A maqam performer’s control of a neutral third is finer than a category-width account would suggest, and the reason is that the category is a different category — a learned one, in a system with more of them. The resolution of a judgement is a property of the vocabulary as much as of the ear, and only the two ends of this essay’s range are properties of the ear alone.

The ladder from here

If every notation is coarser than the ear and finer than the category, then the choice of division is not settled by perception, and the previous rungs have shown it is not settled by consonance either where a tradition does not harmonise. So what does settle it?

The strongest available answer is that the spectrum does — that a scale is what a family of instruments’ partials ask for. That account designed a working scale for a synthetic spectrum two rungs ago. The next rung runs it on the tradition it is most often invoked to explain, and it does not survive the encounter.

Part 4 of 14

One essay in the series on beyond twelve. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

What this makes readable

Essays that declare this one a prerequisite.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

BeatingCategorical perceptionCentsDifference limenEqual divisionMicrotonalityPitch discrimination