Studio Acoustic Standards for High-Fidelity TTS Datasets
Training high-fidelity Text-to-Speech (TTS) models requires pristine source audio, as modern neural vocoders and acoustic models inherently amplify background noise, room reflections, and spectral inconsistencies. To build an optimal reference speech dataset, recording environments must adhere to strict acoustic standards regarding ambient noise, reverberation, sound isolation, and hardware linearity. Failing to meet these baseline criteria results in synthetic voices characterized by phasing artifacts, high-frequency hiss, and robotic unnaturalness.
Ambient Noise Floor (Noise Criterion)
The ambient noise floor of the recording booth must be exceptionally low to prevent machine learning algorithms from modeling background noise as part of the speaker's vocal characteristics.
- Target Rating: The studio must achieve a Noise Criterion rating of NC-15 or lower (equivalent to roughly NR-15 in Europe).
- SPL Measurement: The equivalent continuous sound level (\(L_{Aeq}\)) should not exceed 18 to 20 dBA across the audible spectrum.
- HVAC Mitigation: Air handling units must be externally located, heavily baffled, and run at ultra-low velocities to eliminate sub-bass rumble (<80 Hz) and broadband airflow noise.
Reverberation Time (\(RT_{60}\))
Neural TTS architectures require dry, direct vocal sound. Any room coloration or early reflections will blur phoneme boundaries and induce comb filtering in generated speech.
- \(RT_{60}\) Target: Decay time across the vocal range (125 Hz to 8 kHz) must remain between 0.1 and 0.2 seconds.
- Spectral Flatness: The decay curve must be linear across the frequency spectrum. Excess low-end build-up ("boxiness") caused by inadequate bass trapping alters the natural spectral tilt of the voice.
- Treatment Type: Studios must use broadband absorption combined with targeted diffusion rather than thin acoustic foam, which selectively absorbs only high frequencies and skews tonal balance.
Sound Isolation (Sound Transmission Class)
To ensure long recording campaigns proceed without acoustic interruptions from external urban environments, structural isolation is mandatory.
- Isolation Rating: Partitions, doors, and viewing windows must provide a minimum Sound Transmission Class (STC) of 55 to 60.
- Mechanical Decoupling: The studio requires a "room-within-a-room" construction with floating floors and resilient wall mounts to reject low-frequency ground-borne vibrations (e.g., traffic, HVAC vibration, footsteps).
Electroacoustic Signal Chain Purity
Acoustic control extends to the interaction between the physical room and the capture chain. Equipment selection must emphasize neutrality over coloration.
- Microphone Self-Noise: The primary microphone (typically a large-diaphragm condenser with a flat response) must have an equivalent noise level of \(\le\) 7 to 9 dBA.
- Preamplification: Preamps must operate transparently with Total Harmonic Distortion plus Noise (THD+N) below -100 dB at operating levels, ensuring no analog warmth or saturation enters the dataset.
- Format Standards: Audio must be digitized at uncompressed linear PCM, minimum 24-bit depth and 48 kHz or 96 kHz sampling rates, preserving high-frequency air bands required for modern vocoder upsampling.
Consistency and Spatial Control
Because dataset creation typically spans weeks, spatial relationships within the acoustic field must remain static to avoid variations that degrade model convergence.
- Mic-to-Mouth Distance: Fixed mechanically (e.g., using physical spacers or calibrated pop filters) at 15 to 20 cm, slightly off-axis to avoid plosive energy.
- Thermal and Humidity Stability: Environmental controls must maintain steady temperature (20–22°C) and relative humidity (40–50%) to stabilize both the voice talent’s vocal folds and the velocity of sound in the acoustic chamber across sessions.