Glottal Pulse Profile and TTS Vocal Brightness
In Text-to-Speech (TTS) synthesis, the glottal pulse profile of the fundamental frequency (\(F_0\)) acts as the primary acoustic driver of perceived vocal brightness. Vocal brightness refers to the prominence of high-frequency harmonic energy relative to the lower frequencies in a voice. By dictating the abruptness of vocal fold closure and the shape of the excitation waveform within the classic source-filter framework, the glottal pulse profile directly establishes the spectral tilt of the raw voice source before it ever reaches the vocal tract filters.
The Source-Filter Foundation and Spectral Tilt
Speech synthesis fundamentally relies on an excitation source passing through a linear vocal tract filter. While the vocal tract shapes vowel formants, the glottal source waveform determines the baseline distribution of harmonic energy across the frequency spectrum.
The profile of this pulse—traditionally parameterized by models such as the Liljencrants-Fant (LF) model—defines how the glottis opens, reaches peak airflow, and abruptly closes. The rate of change in airflow during the closing phase, known as the maximum flow declination rate (MFDR), dictates the spectral roll-off. A standard voice typically rolls off at roughly -12 dB per octave. When a glottal pulse exhibits an extremely sharp, discontinuous closure, the spectral slope flattens toward -6 dB per octave, pushing significant acoustic energy into the 2 kHz to 5 kHz region. This excess high-frequency energy produces a perceptually "bright," "pressed," or "cutting" vocal timbre.
Glottal Parameters Governing Brightness
Several specific timing properties within the glottal cycle determine whether a synthesized voice sounds dark, balanced, or bright:
- Open Quotient (\(O_q\)): The ratio of time the glottis is open relative to the entire pitch period (\(1/F_0\)). A lower open quotient corresponds to a longer closed phase, which generally concentrates energy into well-defined higher harmonics, enhancing brightness.
- Speed Quotient (\(S_q\)): The ratio of the opening phase duration to the closing phase duration. Asymmetry favoring a rapid closing phase sharpens the excitation impulse, increasing high-frequency harmonic generation.
- Return Phase (\(T_a\)): The effective duration of the residual airflow deceleration after initial closure. A shorter, more abrupt return phase prevents high-frequency damping, directly contributing to a brighter timbre, whereas a prolonged return phase introduces low-pass filtering, resulting in a darker or breathier tone.
Interaction with the Fundamental Frequency
While the glottal pulse profile sets the envelope of source energy, the fundamental frequency (\(F_0\)) sets harmonic spacing. At higher \(F_0\) values, harmonics are spaced further apart. If the glottal pulse lacks high-frequency energy, high-pitched synthetic voices quickly sound thin or muffled due to sparse harmonic content. Conversely, if an aggressively sharp glottal pulse profile is applied at high frequencies, the synthesis can easily become harsh or unnaturally metallic. TTS engines must dynamically scale the pulse profile with \(F_0\) variations to maintain consistent perceived brightness across a speaker's pitch range.
Implications for Modern Neural and Parametric TTS
In parametric and source-filter-based neural TTS architectures (such as DDSP or neural vocoders with explicit excitation modules), failing to accurately model the glottal pulse profile leads to common synthesis artifacts. An overly smoothed synthetic excitation results in dull, muffled speech, whereas an oversimplified impulse train introduces an unnatural, buzzy brightness. By explicitly controlling or accurately predicting the glottal flow derivative, modern synthesis engines can realistically manipulate vocal effort, expressiveness, and perceived brightness without distorting speaker identity or linguistic clarity.