Glottal Stops in British English TTS
This article examines the essential acoustic cues that text-to-speech (TTS) engines must synthesize to accurately reproduce glottal stops in colloquial British accents. Accurately modeling glottal replacement—most commonly associated with T-glottalization in dialects such as Cockney, Estuary English, and Multicultural London English—requires more than simply dropping phonemes. To achieve natural, authentic British speech synthesis, an engine must precisely control silent intervals, laryngealization, formant dynamics, and the absence of oral release bursts.
Complete Silence and Duration Control
A canonical glottal stop ([ʔ]) involves the total adduction of the vocal folds, blocking the airstream at the glottis rather than the alveolar ridge.
- Silence Interval: The TTS model must insert a brief silent gap into the acoustic waveform. In natural speech, this closure duration typically ranges between 30 to 80 milliseconds, depending on speaking rate and stress.
- Abrupt Amplitude Drop: Unlike the gradual decay seen in lenis consonants, the transition into a glottal stop requires an immediate drop in sound pressure level across all frequency bands.
Laryngealization and Creaky Voice
In natural colloquial British speech, full acoustic closure does not always occur; instead, speakers frequently produce glottalization or creaky voice. A high-quality neural vocoder or parametric synthesizer must generate these specific non-modal phonation patterns:
- Fundamental Frequency (\(F_0\)) Perturbation: As the vocal folds compress, the pitch often drops sharply immediately before and after the glottal event.
- Periodicity and Jitter: The engine must reproduce irregular pitch periods (jitter) and cycle-to-cycle amplitude variations (shimmer) in the surrounding vowel margins.
- Increased Spectral Tilt: Creaky voice features longer closed phases within each glottal cycle, altering the harmonic structure by boosting lower harmonics relative to higher frequencies.
Suppression of Oral Release Bursts
When replacing voiceless alveolar plosives (/t/) with a glottal stop, the engine must actively suppress the standard acoustic cues of oral stops:
- Elimination of Aspiration: Standard British Received Pronunciation (RP) produces a characteristic burst of high-frequency noise followed by aspiration upon the release of /t/. In colloquial accents using glottal replacement (e.g., in words like water, bottle, or butter), all transient burst noise and frication between 3 kHz and 8 kHz must be entirely absent.
- Unvoiced Release: If a release phase is present, it must manifest as an unobtrusive resumption of voicing or silence, lacking the sharp alveolar click.
Preserved Vowel Formant Trajectories
A common failure in synthetic glottal stops is treating the glottal stop as an alveolar consonant with missing audio.
- Supraglottal Neutrality: Because the constriction happens at the vocal folds rather than with the tongue tip, the vocal tract above the glottis does not shift toward alveolar target values.
- Formant Transitions: The second (\(F_2\)) and third (\(F_3\)) formants should not bend toward the typical 1800 Hz alveolar locus. Instead, the formants of the preceding vowel must remain relatively flat or reflect only the anticipatory positioning of the subsequent sound.
Contextual and Phonotactic Modulation
To ensure realistic regional inflection, the TTS architecture must apply glottal cues conditionally based on phonological context:
- Intervocalic Placement: Foot-internal, intervocalic positions (city, better) require pronounced creakiness or full closure depending on the targeted regional dialect.
- Pre-Consonantal and Syllabic Environments: Before syllabic nasals (button) or liquids (little), the glottal stop transitions directly into the consonant without intermediate oral gestures.
- Absolute Word-Final Boundaries: At utterance boundaries, glottal stops often resolve into prolonged creak rather than a distinct acoustic break.