Statistical Tests for TTS Listening Tests

Subjective listening tests, such as Mean Opinion Score (MOS), MUSHRA, and AB preference tests, represent the benchmark for evaluating Text-to-Speech (TTS) audio quality and naturalness. However, reporting raw score increases without proper statistical testing undermines the validity of research findings. This article outlines the essential statistical significance tests required to rigorously substantiate perceived improvements in TTS systems, categorizing methods by experimental design, data distribution, and error correction.

Mean Opinion Score (MOS) Tests

In an absolute category rating (ACR) MOS test, listeners evaluate audio samples on an ordinal scale, typically ranging from 1 (Bad) to 5 (Excellent). Because these ratings represent discrete, ordinal data that rarely follow a normal distribution, standard parametric tests can produce misleading results.

MUSHRA Tests

MUlti Stimulus test with Hidden Reference and Anchor (MUSHRA) evaluations present listeners with multiple systems simultaneously on a continuous scale from 0 to 100. This within-subject design yields paired, continuous data.

AB and ABX Preference Tests

Preference tests ask listeners to choose which of two samples sounds better, more natural, or more intelligible. The output produces categorical or binomial data.

Multiple Comparisons Corrections

When comparing three or more TTS architectures against one another or against a shared baseline, conducting multiple pairwise tests increases the probability of Type I errors (false positives). Applying a correction method is mandatory:

Effect Sizes and Confidence Intervals

Statistical significance denotes only that an observed difference is unlikely due to random chance; it does not measure practical relevance. Comprehensive TTS reporting requires: