PESQ and POLQA in Text-to-Speech Evaluation

Perceptual Evaluation of Speech Quality (PESQ) and Perceptual Objective Listening Quality Analysis (POLQA) are standardized objective metrics originally engineered to predict human Mean Opinion Scores (MOS) in telecommunications. While researchers frequently attempt to use them to benchmark Text-to-Speech (TTS) systems, their utility in this domain is highly constrained. PESQ and POLQA apply effectively only to isolated subcomponents of a synthesis pipeline—such as vocoder reconstruction against known ground-truth audio—but fundamentally fail as end-to-end TTS metrics due to their reliance on time-aligned reference signals, inability to score prosody or expressiveness, and insensitivity to generative synthesis artifacts.

The Design Purpose of PESQ and POLQA

PESQ (ITU-T P.862) and its successor POLQA (ITU-T P.863) are full-reference perceptual quality algorithms. They compare a degraded audio signal against an original, pristine reference signal to quantify quality degradation introduced by telecommunication networks, speech codecs (such as AMR or Opus), packet loss, and jitter. Both models simulate human psychoacoustics—mapping signals into loudness and pitch representations to calculate perceptual distance scores that correlate with subjective listening tests.

Where PESQ and POLQA Apply to TTS

Because these algorithms require a reference signal, their application to TTS is limited strictly to scenarios featuring a direct acoustic ground truth:

Where PESQ and POLQA Fail to Apply

Despite their success in network telephony, PESQ and POLQA break down when applied to general or end-to-end TTS evaluation:

Preferred Alternatives for TTS Quality Evaluation

Due to the architectural limitations of PESQ and POLQA, TTS assessment relies on other methods: