Formant Preservation: Singing Synthesis vs Speech TTS

Formant preservation during pitch shifting is a fundamental requirement in acoustic signal processing, but its precision is far more critical in singing voice synthesis (SVS) than in standard text-to-speech (TTS). While speech synthesis operates within narrow fundamental frequency ranges where minor acoustic distortions often pass unnoticed, singing voice synthesis demands wide-range, dynamic pitch shifts over sustained periods. Without strict formant preservation, pitch-shifted singing immediately exhibits the unnatural "munchkin" or "giant" effects, compromises vowel intelligibility, and alters the perceived anatomical identity of the singer.

Source-Filter Independence and Vocal Anatomy

Human vocal production relies on the source-filter model: the vocal folds generate a fundamental frequency (\(F_0\)) along with harmonics (the source), and the vocal tract acts as a resonant acoustic cavity that filters this sound (the filter). The resonant frequency peaks created by this cavity are formants (\(F_1, F_2, F_3\), etc.).

When a human sings higher or lower, they stretch or relax their vocal folds to change \(F_0\); their vocal tract length remains relatively static. Therefore, their formants stay relatively fixed regardless of the pitch. If a pitch-shifting algorithm scales the entire spectral envelope uniformly alongside \(F_0\), it alters the formants, effectively modeling an impossible physical transformation where the singer's throat expands or contracts with every note.

Extreme Dynamic Pitch Ranges

Sustained Notes vs. Transient Phonemes

Speech is characterized by rapid transitions between consonants and vowels. The transient nature of conversational speech masks minor acoustic artifacts, phase errors, and slight spectral envelope shifts because the listener's brain prioritizes linguistic decoding over acoustic purity.

Singing, conversely, emphasizes sustained vowels held over beats or measures. During these prolonged states, listeners have ample time to evaluate the timbre, warmth, and naturalness of the voice. Any unnatural drift in formant positions during sustained notes stands out immediately as artificial, robotic, or metallic.

Vowel Intelligibility and Formant Positioning

Formant locations—specifically the relationship between the first two formants (\(F_1\) and \(F_2\))—define vowel identity in human language.

In SVS, if pitch shifting forces formants to shift proportionally with the musical score:

  1. High-pitched notes shift lower formants into frequency bands designated for completely different vowels (e.g., an /ɑ/ turning into an /o/ or /u/).
  2. High fundamental frequencies may exceed the natural frequency of the first formant, causing severe acoustic dropouts if the filter is not carefully decoupled and preserved.

Trained singers often perform subtle "formant tuning" to align tract resonances with individual harmonics for vocal projection. To synthesize realistic singing, an engine must maintain precise control over these resonant boundaries, ensuring that changes in pitch do not degrade linguistic comprehension.

Timbral Consistency and Singer Identity

A vocal tract's fixed formant profile is the primary acoustic fingerprint defining a specific vocalist. In speech TTS, subtle shifts in timbre still sound like an acceptable human speaker, even if slightly varied. In music, a singer’s unique timbre must remain completely coherent across the entire range of a song. Formant preservation ensures that a synthetic vocalist sounds like the exact same person whether they are singing a low bass note or a high belt.