Orrery — Copenhagen
2026-08-23
Why GSM calls sound underwater
A 1990s mobile call does not sound bad at random. It sounds bad in a specific, repeatable way: a faint gurgle under vowels, sibilants that thin into a soft hiss, a texture closer to a voice heard through a wall of water than through wire. That quality has a name and a cause. It is RPE-LTP — Regular Pulse Excitation, Long-Term Prediction — the speech codec GSM adopted to carry a call at 13 kilobits a second, and the underwater grain is not a defect in some particular handset. It is the sound of the compression working as specified.
GSM’s full-rate codec, standardized as GSM 06.10 by the European GSM standardization group (later folded into ETSI), takes voice sampled 8,000 times a second and fits each 20-millisecond frame — 160 samples — into 260 bits. That is 13,000 bits a second — a fraction of what a landline call carries, let alone a CD. Nearly everything about the underwater sound follows from how that small budget gets spent.
RPE-LTP does not transmit the waveform. It predicts it. A short-term linear predictor models the shape of the vocal tract for that frame — the resonant shape of mouth and throat that gives a vowel its character. A second, long-term predictor models pitch: the repeating pulse of the vocal cords in voiced speech, spaced by the pitch period. Together the two predictors reconstruct most of what a vowel sounds like from very little data, because a vowel is repetitive and a handful of numbers describes it well.
What is left after both predictions — the residual, the part neither predictor accounted for — is where an unpredictable sound actually lives: consonants, sibilants, breath, background noise. RPE-LTP does not send that residual whole. It sends a regularly spaced subset of it — every third sample, thirteen out of forty, on one of four candidate grids chosen for whichever carries the most energy. The decoder puts those thirteen pulses back on the grid and zeros between them, and leaves the synthesis filters to make a signal out of that. For a clean vowel the approximation holds and the loss is barely audible. For a fricative or a hard consonant, where nearly all the information sits in that discarded residual, the approximation runs out — and what returns is a wavering, band-limited, faintly liquid texture. It is the codec’s honest best guess at a signal type its architecture was never built to carry well.
Thirteen kilobits a second was not a stingy choice; it was the ceiling. Channel coding turned it into 22.8 kilobits on the air, and eight of those shared a single 200-kilohertz carrier — decoded in real time on mobile silicon with a fraction of a modern phone’s processing budget. RPE-LTP was the trade that fit inside that ceiling: spend the bits on what speech does predictably — pitch, vocal-tract shape — and let everything else arrive as a coarser approximation. The alternative was not a clearer call. The alternative was a network that could not carry the call at all.
Salo Oy built the network-side test hardware that lived with exactly this trade in the mid-1990s — codec boards and comfort-noise generators that told a Nordic carrier’s engineers what a call would sound like before it reached a customer’s ear. Orrery’s restoration is modeled from one of those units:
Salo 3310 · Full-Rate voice codec · restored from unit no. CU-9426-0173 · Salo Oy, Salo, Finland, 1994
The watery quality is the RPE-LTP algorithm working exactly as designed. We have not clarified it; clarity was never the assignment.
The instrument: https://orrery.dk/salo/3310/
Listening
Demonstration recordings, prepared at the bench.
Eight generations down
Degrade is walked from nought to full — one encode generation stacking toward eight — and at twenty-four seconds Rate drops to Half. Distance accumulates.
0:00 / 0:30