Concept

Scale-space — where it appears

Reading a signal at several degrees of smoothing at once, and taking the features that survive the most smoothing as the strongest. It turns a list of phrase boundaries into a hierarchy, and it replaces a threshold with a scale — which then has to be chosen, or supplied from outside.

Named by 4 essays across 2 fields — each of them below, with the objects they name alongside it.

The boundaries that survive each amount of smoothing. The local boundary strengths of Twinkle, twinkle read at every scale: the curve is smoothed with a Gaussian of the width on the horizontal axis and the peaks that survive are counted. Small scales give 11 boundaries and large ones give one, and the notation marks 5. The level with that many falls at a width of 2, where the model finds 100 per cent of the notated boundaries and 100 per cent of what it finds is notated — a comparison with no threshold in it, which is what the scale parameter buys.

A boundary at a stated level

A boundary detector run over three tunes agreed with the notation on one and barely at all on another, and left two things owing: a version with a scale parameter, and a version run on performance timings. Both are paid here, and they pay differently — the scale removes a free parameter from the comparison and does not rescue the hard case, while two per cent of rubato does.

form · Phrase
Twinkle, twinkle, phrased at the level each tempo selects. The number of boundaries the model finds when its smoothing scale is set by the psychological present rather than chosen, against the tempo the tune is taken at. The scale in notes is the present's 3.5 seconds divided by the mean note length, so a fast tempo puts more notes inside the present and smooths harder. The page's own phrasing has 5 boundaries, drawn as the flat line; the model matches it best at 160 beats per minute, where the present holds 8.2 notes. The same tune at two tempos is read at two levels, which is the prediction and is not a free parameter.

The level the tempo chooses

The boundary detector has a scale parameter and an earlier essay left it free, ending with the sentence that names this one: the scale is in notes and the psychological present is in seconds. The psychological present is two to eight seconds, a tempo converts one to the other, and the level a listener reads then stops being a parameter at all — which is a prediction with teeth, because the same tune at two tempos should be phrased differently at levels the arithmetic names in advance.

form · Phrase
How far the detector looks, note by note. The number of notes that fit inside a 3.5-second present at each point of the tune, once the performance has lengthened its phrase-final notes by 30 per cent. It runs from 5 to 10 notes against a constant 7 for the unperformed version, and it dips exactly where a boundary is, because a boundary is where the performance slows. Reading the boundary-strength curve with that width at every point instead of one width everywhere gives an agreement of 0.55 with the notated phrasing, against 0.36 for the fixed width the present dictates and 0.71 for a fixed width fitted to this tune. The dips are marked, and the notated boundaries are the vertical lines: the detector narrows itself at the places it is supposed to find, which is the circularity this figure has to be honest about — the lengthening was put there by the notation.

A detector whose resolution the performance sets

The boundary detector lost its free parameter when the psychological present became a number of notes at a stated tempo, and what that held still was named at the time: a performance slows into a phrase end, so the number of notes inside the present is not the same everywhere in a tune — it falls exactly where a boundary is. Making the width follow the performance recovers half of what removing the parameter cost, and honestly leaves the other half.

form · Phrase
The same tune read at six widths of the psychological present. A later essay made the detector's smoothing width a function of position, which removed its last free parameter but one — and the one it cannot remove is the width of the psychological present, because that is a fact about listeners rather than a choice. So the honest object is not a reading but a family of them, one per width. A short present finds 6 boundaries and a long one finds 2, and the family agrees on 0 of them. The fixed-width control, at its own best width, scores 0.67 against the adaptive readings' 0.67, 0.75, 0.33, 0.33, 0.40, 0.40 — so the adaptation does not win, which is what that essay reported too. What the family adds is the ordering: a boundary in every row is a different claim from one in a single row, and a single reading has no way to say so.

A family of readings

Removing the detector's free parameter, and then its constant tempo, cost persistence both times — the property that made its boundaries ordered rather than merely found. Recovering it means a family of adaptive readings rather than one, indexed by the width of the psychological present, which is the one parameter that cannot be removed, because it is a fact about listeners.

rhythm · Phrase

Named alongside it

The objects these essays reach for when they reach for this one.

PhraseSegmentationRubatoGroupingHierarchyLbdmPerceptual presentTempoFinal lengtheningNotationPersistencePsychological present

All concepts