Filtering Inconsistent Vocal Fry in TTS Datasets

Building a high-quality, multi-day Text-to-Speech (TTS) dataset requires strict acoustic consistency, yet voice talent naturally experiences physiological changes that introduce unintended vocal fry—also known as creaky voice—across different recording sessions. This article details the computational and procedural workflows dataset curators use to detect, measure, and filter out inconsistent vocal fry. By combining acoustic feature extraction, machine learning classification, session baseline tracking, and targeted pruning strategies, audio engineers ensure the synthetic voice maintains stable timbre and pitch across thousands of generated utterances.

Acoustic Indicators of Vocal Fry

Vocal fry occurs when the vocal folds close tightly and vibrate irregularly at a very low frequency, creating a distinct rattling or popping sound. To identify this algorithmically, curators extract several key acoustic features from the raw audio:

Automated Detection Pipelines

Curators rarely inspect tens of hours of speech manually. Instead, they deploy automated pipelines combining signal processing tools with machine learning models:

  1. Feature Extraction Engines: Open-source speech processing tools (such as Praat/Parselmouth and openSMILE) calculate frame-by-frame metrics including Cepstral Peak Prominence (CPP), local jitter, and subharmonic ratios.
  2. Frame-Level Classifiers: Lightweight binary classifiers (such as Random Forests or Support Vector Machines trained on acoustic features) or deep learning architectures (like 1D-CNNs or fine-tuned wav2vec 2.0 checkpoints) classify speech frames as modal or creaky.
  3. Boundary Aggregation: Because vocal fry frequently manifests at the ends of phrases (terminal fry) due to falling subglottal pressure, detection scripts specifically calculate the percentage of fry present in the final 200–500 milliseconds of each phrase.

Session Baseline Normalization and Drift Tracking

Vocal fry fluctuates based on hydration, fatigue, time of day, and recording environment. Curators mitigate multi-day inconsistency through baseline calibration:

Remediation and Filtering Workflows

Once inconsistent vocal fry is pinpointed, curators take distinct actions depending on the severity and location of the creak: