Causes of Phase Smearing in Audacity Time Stretch
Extreme tempo changes in Audacity often introduce phase smearing, watery textures, and flanging artifacts due to the mathematical limitations of time-stretching algorithms. When audio is drastically slowed down without changing pitch, the software must duplicate and redistribute audio information across time. This process destabilizes the phase relationships between individual frequencies and overlapping audio grains, creating comb filtering and transient blurring that manifest as metallic, hollow, or flanging sounds.
How Time-Stretching Works in Audacity
Audacity typically relies on two primary algorithms for time-shifting: SoundTouch (used in the standard "Change Tempo" effect) and SBSMS (Subband Sinusoidal Modeling Synthesis, used for higher-quality rendering).
Both approaches attempt to decouple duration from pitch, but they operate through different mechanics:
- Time-Domain Approaches (WSOLA/SoundTouch): Waveform Similarity Overlap-Add chops audio into small segments (grains), spaces them further apart to stretch the duration, and crossfades the overlaps where waveforms match most closely.
- Frequency-Domain Approaches (Phase Vocoder/SBSMS): These divide the audio into overlapping Short-Time Fourier Transform (STFT) analysis windows, measuring the frequency, amplitude, and phase of sinusoidal components over time before resynthesizing them across a stretched timeline.
Phase Incoherence and Drift
The primary driver of flanging and smearing is phase incoherence. Natural acoustic sounds consist of complex harmonically related frequencies that maintain strict, coherent phase relationships with one another.
When a phase vocoder shifts these components across an extended timeline, it must recalculate the phase angle of every frequency bin for each synthesized frame. Even minor rounding errors or miscalculations cause individual frequencies to drift out of alignment with their adjacent harmonics. This internal desynchronization—known as horizontal phase incoherence—turns focused, punchy instruments into diffuse, smeared textures.
Comb Filtering and Flanging
When time-domain algorithms like WSOLA overlap audio grains to fill the stretched time gap, identical or near-identical sound fragments are summed together with slight time offsets.
Mixing an audio signal with a slightly delayed version of itself causes constructive and destructive interference, creating a series of regularly spaced frequency notches known as comb filtering. When these time offsets dynamically shift from frame to frame, the frequency notches sweep across the spectrum, producing a distinct flanging or jet-engine chorus effect.
Transient Smearing and Dispersion
Transients, such as drum hits and consonant speech sounds, require precise vertical phase alignment across all frequencies at an exact point in time.
Drastic time-stretching breaks this vertical alignment. Instead of arriving simultaneously, the high and low frequencies of a transient arrive slightly out of sync. This disperses the energy of a sharp strike over tens or hundreds of milliseconds, transforming crisp attacks into dull, "watery" thuds or generating unnatural pre-echo artifacts.
Compounding at Extreme Stretch Ratios
At modest tempo reductions (such as 5% to 15%), phase-locking heuristics can reasonably estimate phase continuity, keeping artifacts below the threshold of human perception. However, when tempo is stretched drastically (such as 50% slower or more), the algorithm is forced to synthesize vast amounts of nonexistent time data. The predictive models collapse under this workload, multiplying phase estimation errors and rendering phase smearing unmistakably audible.