Limitations of Crowdsourced MOS in Text-to-Speech

Crowdsourcing Mean Opinion Score (MOS) evaluations has become the standard approach for benchmarking Text-to-Speech (TTS) models due to its scalability, low cost, and fast turnaround. However, relying on distributed, non-expert workers introduces significant statistical noise, behavioral artifacts, and systematic biases that compromise the validity of the results. This article explores the primary limitations and biases inherent in crowdsourced TTS MOS evaluations, detailing how environmental variability, worker behavior, cognitive heuristics, and demographic factors skew subjective audio quality assessments.

Hardware and Acoustic Environment Variability

In traditional laboratory listening tests, participants use calibrated, high-fidelity headphones within sound-isolated listening booths. In contrast, crowdsourced workers participate using an uncontrollable range of hardware, including low-quality built-in smartphone speakers, standard consumer earbuds, and high-end studio monitors. Each audio transducer has a distinct frequency response that can either mask or artificially emphasize synthesis artifacts like high-frequency phase distortion, robotic metallic buzzing, or low-frequency sub-band noise. Furthermore, varying ambient noise levels in remote worker environments prevent consistent perception of subtle prosodic flaws and background artifacts.

Worker Incentives and Low-Effort Participation

Micro-task crowdsourcing platforms typically compensate workers based on the number of completed tasks rather than the quality of evaluation. This economic dynamic incentivizes high speed over careful listening, leading to:

While trap questions, attention checks, and synthetic audio anchors can mitigate fraudulent submissions, they add complexity to test design and do not completely eliminate low-effort responses from real workers.

Cognitive and Contextual Biases

Subjective human judgment is inherently non-linear and context-dependent. Crowdsourced evaluations regularly suffer from several systemic perceptual biases:

Demographic and Linguistic Misalignment

TTS naturalness is deeply intertwined with linguistic familiarity, regional accents, and cultural expectations. Crowdsourced labor pools often fail to align with the target listener demographic:

Collapse of Multi-Dimensional Quality into a Single Metric

Standard absolute category rating (ACR) MOS tests compress complex perceptual dimensions—intelligibility, prosodic naturalness, speaker similarity, and audio fidelity—into a single scalar score from 1 to 5. Crowdsourced raters frequently conflate these separate axes. For example, a model that produces pristine 48kHz audio with flat, robotic prosody may receive higher scores than a model with highly expressive human-like prosody that contains faint background artifacts. This obscures the exact failure modes of the underlying neural architecture.

Lack of Cross-Study Comparability

Because crowdsourced MOS relies on transient, non-standardized participant pools and varying anchor sets, MOS values are not absolute measurements. A MOS of 4.2 reported in one research paper cannot be directly compared to a 4.2 in another paper unless both models were evaluated side-by-side within the exact same listening test by the identical pool of workers. The misuse of MOS as an absolute benchmarking metric across different publications remains a widespread methodological vulnerability in modern speech synthesis research.