Intervals and chords

The chord that is not played at once

Every chord until now starts its notes at the same instant, and no figure ever set the asynchrony to anything else. A spread chord is not a defective simultaneity: at forty milliseconds a triad of half-second notes is still eighty-seven per cent simultaneous, so it carries almost all its roughness, and it stops being a chord at all only when the last note arrives after the first has finished. What none of this can explain is the one thing every keyboard player knows — that a chord is rolled upward. The masking asymmetry that ought to explain it is 2.4 decibels at a close voicing, which is not enough.

Assumes: Three notes at once, and why these three · What makes two partials one note

A chord is defined by its notes sounding together. Every figure in this ladder has taken that literally: the four ways of stacking two thirds, the roughness of every chord of every size, the root an ear supplies and the register a chord is in are all computed for notes beginning at the same instant.

Almost no chord is played that way. A keyboard player rolls; a harp arpeggiates; a guitar’s strum takes twenty or thirty milliseconds from the low string to the high; a string quartet is together to within a few tens of milliseconds at best. The asynchrony is not a failure to be simultaneous — a spread chord is a specific and much-used sound — and this ladder has never had a number for it.

Where the chord actually is

Start with the arithmetic, which is simpler than it looks. Three notes of length d, onsets Δ apart: all three are sounding from the last onset until the first note ends, which is d − 2Δ.

a major triad, spread by 40 milliseconds. Three notes of a major triad, each lasting 600 ms, with their onsets 40 ms apart. The shaded band is the window in which all three are sounding — 520 ms, which is 87 per cent of a note's length. The chord exists as a simultaneity only inside that band; before it and after it the passage is a melody. At 40 ms the notes are further apart than the asynchrony at which a mistimed partial leaves its note.
Fig. 1 A close-position major triad on C3, notes lasting 600 ms, rolled at 40 ms — an ordinary keyboard spread. The shaded band is the window in which all three are sounding: 520 ms of a 600 ms note, 87 per cent. Before it and after it, the passage is a melody. A chord is a thing that happens in the middle of a spread rather than a thing a spread approximates.

That framing gives the quantity the ladder was missing. Roughness is a property of two partials sounding at the same time, so a chord that is 87 per cent simultaneous carries 87 per cent of its roughness — and the fall is linear in the spread, with a slope set by the note length.

What a spread costs a chord. a major triad of 3 notes lasting 600 ms each, with the onsets spread by up to 320 ms. The upper line is the share of each note's length during which every note is sounding; the lower is the chord's roughness weighted by that share, since roughness is a property of two partials sounding at the same time. At a spread of 30 ms — the asynchrony at which a mistimed partial stops belonging to its note — the chord is still 89 per cent simultaneous. It stops being simultaneous at all at 300 ms, which is where the last note arrives after the first has finished.
Fig. 2 The whole sweep, for 600 ms notes. The chord is fully simultaneous at zero and not simultaneous at all at 300 ms, where the last onset coincides with the first note’s end. The two dashed lines are thresholds this site measured for other reasons: 30 ms, the asynchrony at which a mistimed partial stops belonging to its note, and 200 ms, where forward masking is spent.

The important thing about that picture is where the thresholds fall. Every asynchrony that matters perceptually is in the first ten per cent of the range, because the note is long compared with every one of them. At 30 ms — enough to make a partial leave its note entirely — the chord is still 90 per cent itself.

Two different fusions, and only one of them is at 30 ms

That gap is worth naming carefully, because it is easy to read the 30 ms figure as though it applied here and it does not.

Mistune or mistime one partial of a complex tone and it leaves the note: 30 milliseconds of head start is enough. That is a statement about partials becoming a note — a grouping decision the auditory system makes below the level of anything musical, and it has to be fast because the alternative is hearing an orchestra as several hundred separate sine tones.

Notes becoming a chord is a different grouping at a different level, and its threshold is not 30 ms. A three-note spread at 100 ms is unambiguously three attacks and unambiguously one harmony; nobody hears a rolled chord as three unrelated notes. The window that bounds that is the one the phrase ladder measured — the psychological present, two to eight seconds — with the note’s own length setting the tighter bound in practice.

So a chord has two fusion thresholds three orders of magnitude apart, and they belong to different mechanisms. The 30 ms one decides what counts as a note; the seconds-long one decides what counts as a simultaneity for the purposes of harmony.

Partial 3, 60 ms earlyA 10-partial tone on 220 Hz with one partial treated differently from the rest. Mistuning it moves it off the harmonic grid by 0.0 per cent, which is 0.0 Hz — slow enough to be heard as a beat rather than as a separate pitch, and enough for the partial to be heard out of the note as a whistle of its own. An onset difference does the same to a partial that is exactly in tune.19.8 Hz off12345678910partial numberamplitude1% is enoughto hear the partialout of the note;3% and it stopscontributing to thepitch of the wholeMoore, Peters &Glasberg, 1985
Fig. 3 The fast threshold in its own terms: one partial of a complex tone given a 60-millisecond head start, which is twice what it takes to hear it out as a separate sound. This is the mechanism the number belongs to, and it is not the mechanism that decides whether a rolled chord is a chord.

There is another window in this range, drawn for a different purpose and worth putting beside the spread: two arrivals inside a few tens of milliseconds are heard as one event with one location. The chord spreads this essay is about sit across that boundary rather than inside it, which is why a spread chord is heard as a chord arriving late rather than as several events.

The explanation that does not survive

Now the interesting failure. Chords are rolled upward — overwhelmingly, in every keyboard tradition that rolls them, and in the notation, where the wavy line is read from the bottom.

There is an acoustic explanation available and this site has the machinery for it. Masking spreads upward and barely downward: a low tone raises the threshold at higher frequencies far more than a high tone does at lower ones. So a chord rolled from the bottom has each note arriving into the previous note’s masking skirt, and one rolled from the top does not — which would predict that an upward roll blends and a downward roll does not.

The prediction is right in direction and wrong in size.

Spreading a chord upward is not the same as spreading it down. The masked threshold each note of a chord arrives under, at a spread of 40 ms, for three voicings of the same triad. Masking spreads upward far more than downward, so a chord rolled from the bottom puts every note into the previous note's skirt and one rolled from the top does not. The difference grows with the spacing — 2.4 dB close, 7.4 dB open, 19.1 dB wide — and at the wide voicing both thresholds are below audibility anyway, so the asymmetry is real, computable, and too small at ordinary spreads to be why arpeggios go up.
Fig. 4 The masked threshold each note arrives under at a 40 ms spread, rolled up against rolled down, for three voicings. The asymmetry is real and grows with the spacing — 2.4 dB at a close voicing, 7.4 at an open one, 19.1 at a wide one. But at the wide voicing both thresholds are below the audible floor, and at the close voicing, where the effect would be audible, it is two decibels.

Two decibels is about the smallest level difference anybody notices, and it is the difference in a masked threshold rather than in anything heard. Where the asymmetry gets large the notes are so far apart that neither masks the other at all. Between fundamentals there is no voicing at which the masking explanation is both true and large, and the reason looks structural: masking asymmetry needs a wide frequency separation and masking needs a narrow one.

Sweeping every voicing of the triad inside four octaves gives the trade-off exactly. The asymmetry climbs from 2.4 decibels at the close voicing to 41.9 at the widest one tried, and the masked threshold itself falls from 24.4 decibels to −65.4 over the same span — from comfortably audible to forty decibels below anything. The product of the two is maximised at the open voicing, [0, 7, 16], where the asymmetry is 7.4 decibels and the threshold is 12.2. That is the best case the explanation has anywhere, and it is a seven-decibel effect in a threshold, which is still not an account of a four-century convention.

Except that the notes have partials, and that changes the answer

The caveat below says that including partials would raise the upward figure and would not change the finding. Running it changes the finding.

A note is not a sine tone. A low note’s harmonic series lies above it, blanketing exactly the frequencies the upper notes of a chord occupy, while a high note’s series lies above it too and therefore reaches nothing below. So the asymmetry the explanation needs is not really about the skirt of a masking pattern at all — it is about the fact that harmonic sounds mask upward by having partials there.

Computing the masked threshold from every partial of the earlier note, rather than from its fundamental alone:

voicing up, fundamentals up, with partials down
close [0, 4, 7] 24.4 dB 24.4 dB 22.0 dB
open [0, 7, 16] 12.2 dB 14.6 dB 4.8 dB
wide [0, 12, 28] −14.8 dB +13.9 dB −33.9 dB

At the wide voicing the upward figure moves by 28.7 decibels and crosses from inaudible to audible, while the downward figure does not move at all. The asymmetry goes from 18.9 decibels to 47.8, and it is now an asymmetry between a threshold that is well above the floor and one that is thirty-four decibels below it.

The mechanism is a coincidence that is not a coincidence. At [0, 12, 28] the second note is an octave above the first, which is to say it is on the first note’s second partial; the third note then sits between the second and third partials of the note before it. A wide voicing rolled upward has each note arriving where the previous note already has energy, and that is a description of what a spread triad in open position is. Rolled downward, every note arrives in a region the previous note has left empty.

So the honest verdict is narrower than the one this essay set out to give. The masking explanation fails for the close voicings and works for the wide ones, and it fails precisely where the fundamentals are close enough for the skirt argument to be the whole story. Whether that rescues the convention is a separate question — keyboard players roll close chords too, and there the effect really is two decibels — but the claim that no voicing makes the explanation both true and large does not survive giving the notes their partials.

What one tone hides, and in which direction. The masked threshold beside a tone masker: any probe below one of these curves is inaudible while the masker sounds. The frequency axis is in Bark, the scale on which the ear's filters are evenly spaced, so the pattern is a pair of straight lines. The upper slope is much shallower than the lower one and gets shallower still as the masker gets louder — masking spreads upward, not downward.
Fig. 5 The mechanism the explanation rests on, in its own right: the masked threshold beside a tone, for the bottom and the top note of the same triad at the same level. The low masker’s skirt reaches further up than the high masker’s reaches down, which is the asymmetry. What the calculation above adds is the other factor — by the time the second note arrives, forward masking has taken the first note’s effective level from 80 decibels to 30, and a skirt is proportional to what is casting it.

So the practice is only partly explained here, and it is worth saying what would explain the rest. The partial calculation above accounts for wide voicings and leaves the close ones — which are the commonest case and the one the convention is most obviously about — with an asymmetry of two decibels and no explanation at all.

Three candidates for that remainder, none of them measurable with what this site has. The bass note establishing the harmonic root before the upper notes arrive, which is a claim about the root an ear supplies being computed incrementally rather than from a finished set. The mechanics of a hand rolling from a stable thumb, which is a fact about anatomy and would predict the convention to be weaker on the harp, where both hands are free, than on a keyboard. And the notational convention being self-reinforcing once established, which is the explanation that requires no acoustics and which nothing here can rule out.

The second of those is the one with a testable consequence, and it is the sort of thing a corpus rather than a model would settle. It is worth naming as a shortfall rather than leaving as a list: the essay has an acoustic mechanism that works for the voicings the convention is least about, and the case it most needs to explain is the one where every quantity it can compute comes out too small.

What a spread is for

Two uses of the spread follow from the overlap arithmetic, and both are ordinary practice.

It buys time for a bass note to be heard as a bass note. The lowest note of a chord is the one the ear’s root-finding machinery leans on hardest, and it is also the one whose pitch takes longest to establish: a low tone needs more cycles to have a pitch than a high one, which is a fact about periodicity rather than about attention. Rolling from the bottom gives the bass a head start of exactly the kind that would help — which is the practice’s best defence, and is not the masking argument.

It weakens a dissonance without changing it. A chord’s roughness is carried only during the simultaneous window, so spreading a harsh chord makes it measurably less harsh while leaving every pitch in place. At a 150 ms roll a triad of 600 ms notes is half simultaneous and carries half its roughness.

the same triad, spread by 150 milliseconds. Three notes of the same triad, each lasting 600 ms, with their onsets 150 ms apart. The shaded band is the window in which all three are sounding — 300 ms, which is 50 per cent of a note's length. The chord exists as a simultaneity only inside that band; before it and after it the passage is a melody. At 150 ms the notes are further apart than the asynchrony at which a mistimed partial leaves its note.
Fig. 6 The same triad rolled at 150 ms. The simultaneous window has fallen to 300 of the 600 milliseconds, and the passage now reads as three notes with a chord inside them. This is what a harpist does to a chord that would be ugly struck.

And it makes a chord audible in a reverberant room. Beyond the critical distance the room is louder than the instrument, and a struck chord in a stone building arrives as one smear. Rolling it separates the attacks by more than the first eighty milliseconds in which a room is not yet a room, so each note gets its own early arrival. That is a claim this site could check and has not.

And it separates a voicing that would otherwise be a smear. A close triad low on a keyboard is five times rougher than the same three pitch classes in the middle, because the notes fall inside one critical band. Rolling it does not move the notes and does not fix the roughness — but it removes most of the simultaneity the roughness needs.

What a spread costs a chord. a major triad of 3 notes lasting 600 ms each, with the onsets spread by up to 320 ms. The upper line is the share of each note's length during which every note is sounding; the lower is the chord's roughness weighted by that share, since roughness is a property of two partials sounding at the same time. At a spread of 30 ms — the asynchrony at which a mistimed partial stops belonging to its note — the chord is still 89 per cent simultaneous. It stops being simultaneous at all at 300 ms, which is where the last note arrives after the first has finished.
Fig. 7 A major triad of three notes lasting 600 milliseconds each, with the onsets spread by up to 320. The upper line is the share of each note’s length during which every note is sounding; the lower is the chord’s roughness weighted by that share.

Roughness is a property of two partials sounding together, so a spread chord delivers less of it — not because the notes are less rough but because they are less often simultaneous. Every roughness number this ladder has computed assumes an overlap of one, and this is the curve that says what the assumption is worth.

Which computation produced the numbers

The overlap is arithmetic: with n notes of length d and onsets Δ apart, all of them sound for d − (n−1)Δ, and the fraction is that over d.

The roughness is chordRoughness — the same function every figure in this ladder uses, over a string timbre — multiplied by the overlap fraction. That multiplication is a model rather than a measurement, and it is a crude one: it assumes roughness accumulates in proportion to simultaneous time, which is reasonable given that a dissonance has to last and its own rung found roughness needs a few hundred milliseconds to build.

The masked thresholds combine two functions the site already has. Forward masking decays the earlier note’s effective level over the spread — 80 dB falls to 30.5 dB at 40 ms — and the spreading function then gives the masked threshold at the later note’s frequency from that residual level. Both are the published shapes this site’s masking figures draw.

The note length of 600 ms is a choice and every number scales with it. A 200 ms note rolled at 40 ms is only 60 per cent simultaneous, which is why a fast passage cannot be rolled and a slow one can.

Whose music, and when

The rolled chord as a notated device belongs to keyboard and harp writing from the seventeenth century onward, and the convention that the wavy line means upward is European and not universal. A guitar strum is a spread with a direction that alternates by design; a rasgueado is a spread long enough to be a rhythm.

What travels is the arithmetic and one claim about the model. This ladder’s roughness numbers are upper bounds. Every one of them was computed for a simultaneity that almost no performance produces, and the correction is a multiplication by an overlap fraction that a performer chooses. A chord scored as rough here is as rough as it can be played, not as rough as it is played.

What the picture cannot show

Note length is uniform and the notes do not decay. A real rolled chord on a piano has its first note already decaying when the last arrives, so the “simultaneous” window is not one of equal levels — the bass is quieter by then, which shifts the balance in exactly the direction the roll was supposed to fix.

Roughness is taken as accruing linearly in time, which no measurement here supports. The dissonance rung found that roughness needs time to establish; whether it accrues linearly, or with a threshold, or with an onset weighting, is not settled by anything on this site.

The figures the direction figure draws are between fundamentals. Real notes have partials, and a low note’s upper partials sit exactly where the higher notes of the chord are — which is the mechanism a third is rougher in the bass rests on. That was previously recorded here as a correction that would not change the finding; the section above computes it, and at wide voicings it does. The partial sum used there is a power sum over the earlier note’s partials with a string timbre’s amplitudes, which is a crude way to combine maskers and almost certainly overstates the total where several partials contribute at once.

And nothing here is heard. The claim that a rolled chord is still one harmony is a report about listening rather than a measurement, and the site’s own machinery has no way to make it. What it can say is that the roughness falls, by a computable amount, and that is a smaller claim.

The one case where the spread is the harmony

There is a limiting case worth naming because it inverts the whole argument. Spread a chord until the overlap is zero and it is not a weakened chord — it is an arpeggio, and an arpeggio is heard as a harmony anyway.

That cannot be an overlap effect, because there is no overlap. It is memory: the notes are held in the same window that holds a phrase, and the harmony is assembled from them after the fact. Which means a chord has two routes to being heard as one thing — simultaneity, which is acoustic and which everything in this ladder computes, and succession within the present, which is cognitive and which none of it computes.

the triad as an arpeggio, spread by 300 milliseconds. Three notes of the triad as an arpeggio, each lasting 600 ms, with their onsets 300 ms apart. The shaded band is the window in which all three are sounding — 0 ms, which is 0 per cent of a note's length. The chord exists as a simultaneity only inside that band; before it and after it the passage is a melody. At 300 ms the notes are further apart than the asynchrony at which a mistimed partial leaves its note.
Fig. 8 The same three notes at a 300 ms spread, where the simultaneous window has closed entirely. Nothing about this picture is a chord and every listener hears one. Everything computed here about triads applies to the shaded band, and here there is no band — so the fact that the harmony survives is evidence that the acoustic account was never the whole of it.

The boundary between the two routes is where the overlap reaches zero, and that is not a perceptual threshold at all: it is the note length. A harpsichord, whose notes are short, crosses into the second regime at a spread a piano would still be in the first regime at. The same roll is a chord on one instrument and an arpeggio on another, and the deciding number is the decay time rather than anything about hearing.

Where this ladder goes next

Seven rungs of this ladder have asked what three notes at once are and why these three. This one asks what happens when at once is relaxed, and finds that the answer is mostly arithmetic, that the perceptual thresholds are all in the first tenth of the range, and that the practice with the most obvious acoustic explanation available does not have it.

The rung after this one is the doubling. A triad has three notes and a texture that plays it has four voices or six, so one note sounds twice — and which one is doubled is a decision the four-part machinery leaves free and every treatise has opinions about. Roughness, virtual pitch and the masking figures above all have something to say about it, and none of them has been asked.

Part 8 of 9

One essay in the series on the triad. The essays either side of this one:

What links here

Essays that reach for this one mid-argument — the half of a link its own author cannot write down.

The objects named here

The third way in, after the field and the series: the things themselves, and every essay that touches each one.

Attack transientFusionMaskingOnsetRoughnessSimultaneityTriadVoicing