Glottal Stops in British English TTS

This article examines the essential acoustic cues that text-to-speech (TTS) engines must synthesize to accurately reproduce glottal stops in colloquial British accents. Accurately modeling glottal replacement—most commonly associated with T-glottalization in dialects such as Cockney, Estuary English, and Multicultural London English—requires more than simply dropping phonemes. To achieve natural, authentic British speech synthesis, an engine must precisely control silent intervals, laryngealization, formant dynamics, and the absence of oral release bursts.

Complete Silence and Duration Control

A canonical glottal stop ([ʔ]) involves the total adduction of the vocal folds, blocking the airstream at the glottis rather than the alveolar ridge.

Laryngealization and Creaky Voice

In natural colloquial British speech, full acoustic closure does not always occur; instead, speakers frequently produce glottalization or creaky voice. A high-quality neural vocoder or parametric synthesizer must generate these specific non-modal phonation patterns:

Suppression of Oral Release Bursts

When replacing voiceless alveolar plosives (/t/) with a glottal stop, the engine must actively suppress the standard acoustic cues of oral stops:

Preserved Vowel Formant Trajectories

A common failure in synthetic glottal stops is treating the glottal stop as an alveolar consonant with missing audio.

Contextual and Phonotactic Modulation

To ensure realistic regional inflection, the TTS architecture must apply glottal cues conditionally based on phonological context: