Anti-Resonances and Nasal Zeros in Articulatory TTS
Physical articulatory Text-to-Speech (TTS) models speech by simulating the anatomical mechanics and wave propagation of the human vocal tract. While this approach offers high naturalness and physical realism for simple non-nasal vowels, the introduction of branched acoustic pathways—specifically during the production of nasalized vowels and consonants—creates significant modeling challenges. This article examines why anti-resonances (spectral zeros) caused by the nasal and sinus cavities complicate mathematical formulation, parameter estimation, and digital filter stability in articulatory synthesis.
Violation of All-Pole Source-Filter Approximations
Traditional articulatory synthesis and linear predictive coding (LPC) rely on an all-pole assumption, where the vocal tract is treated as a single, unbranched acoustic tube. This representation produces a transfer function characterized entirely by resonance peaks (formants). When the velum lowers, the acoustic path bifurcates into the pharyngeal-oral and pharyngeal-nasal tracts.
This parallel branching produces anti-resonances, or transmission zeros, where specific frequency components cancel each other out due to destructive interference. Because an all-pole filter cannot reproduce these sharp spectral nulls without an infinite or impractical number of poles, articulatory models must adopt pole-zero (ARMA) architectures. These architectures are far more difficult to design, parameterize, and compute in real time.
The Shunting Effect of Closed Cavities
During nasal consonants such as /m/, /n/, and /ŋ/, the oral cavity is completely occluded at the lips, alveolar ridge, or velum, while the velopharyngeal port remains open. The occluded oral cavity acts as an acoustic side-branch, or shunt, absorbing acoustic energy at its own natural resonant frequencies.
Because this trapped energy cannot radiate past the closure, it manifests in the radiated acoustic output as deep spectral troughs (nasal zeros). Modeling these zeros requires dynamic tracking of the exact physical point of oral closure, the precise volume of the occluded oral cavity, and the dynamic acoustic coupling between the pharynx, oral tract, and nasal passage.
Non-Linearities of Sinus Cavities and Soft Tissue
Nasal acoustics are not governed solely by the primary nasal passage; the paranasal sinuses act as Helmholtz resonators attached along the nasal walls. These small, rigid side-cavities introduce additional fixed and highly complex pairs of poles and zeros into the frequency spectrum, particularly between 300 Hz and 2500 Hz.
Furthermore, unlike the oral cavity, the nasal tract possesses a high surface-area-to-volume ratio lined with vascularized, mucous-covered tissue. This soft-tissue environment introduces disproportionately high thermal and viscous boundary-layer losses. Accurately modeling the shifting bandwidths and depths of nasal anti-resonances requires precise specification of complex wall impedances, which are difficult to measure in vivo and computationally expensive to simulate numerically.
Complications in Acoustic-to-Articulatory Inversion
Training articulatory TTS systems often requires acoustic-to-articulatory inversion, where articulatory trajectories are deduced from recorded speech audio. Anti-resonances severely obscure this mapping.
Because nasal zeros can cancel out or heavily damp nearby vocal tract formants (such as the first formant, \(F_1\)), standard peak-picking and spectral tracking algorithms fail. The presence of zeros also increases the non-uniqueness problem of articulatory inversion: multiple different physical tract configurations can generate nearly identical spectral envelopes when zeros obscure dominant resonant frequencies.
Dynamic Topology Changes and Filter Transients
In running speech, transitions into and out of nasality involve rapid topological changes. As the velum opens or closes, the acoustic network switches dynamically between a single transmission line and a complex branched manifold.
In discrete-time physical models, such as digital waveguide meshes or transmission line analogs, switching boundary states dynamically can introduce severe numerical artifacts, energy imbalances, and audible discontinuities. Smoothly transitioning anti-resonance locations across moving velopharyngeal boundaries requires precise interpolation of cross-sectional area functions to avoid synthesis instability and maintain high acoustic fidelity.