Accented TTS Scoring: Native vs Non-Native Listeners
Subjective evaluation of Text-to-Speech (TTS) systems relies heavily on human perception, but the linguistic background of the evaluator introduces significant variance into the results. This article explores how a listener's native-speaker status influences the scoring of accented synthetic speech. It details the mechanisms behind perceptual differences, including phonological sensitivity, shared-accent intelligibility benefits, and varying criteria for naturalness versus comprehensibility, providing key insights for building more accurate TTS evaluation benchmarks.
The Role of Listener Background in Speech Evaluation
Subjective evaluation methods—such as Mean Opinion Score (MOS), Word Error Rate (WER) via transcription tasks, and degradation category ratings—serve as standard benchmarks for measuring TTS quality. These tests assess characteristics like naturalness, intelligibility, and overall user preference.
However, speech perception is inherently subjective and shaped by a listener's native language (L1). When synthetic voices exhibit regional, foreign, or non-standard accents, a listener's native-speaker status directly alters how speech errors, prosodic irregularities, and acoustic artifacts are perceived and penalized.
Perceptual Sensitivity and Scoring Strictness
Native listeners typically demonstrate higher sensitivity to subtle phonological and prosodic deviations. Because they possess deeply ingrained internal models of their language's standard phonetics, rhythm, and intonation, native speakers readily detect unnatural stress placement, improper vowel elongation, or inaccurate pitch contours produced by accented TTS models.
Consequently, native listeners tend to assign lower naturalness and acceptability scores to accented TTS output. While they can often understand the message, the cognitive friction caused by unexpected acoustic patterns leads to harsher subjective penalties. Non-native (L2) listeners, in contrast, frequently rely on broader contextual cues and may be less critical of subtle prosodic flaws, provided the speech remains functionally comprehensible.
The Interlanguage Speech Intelligibility Benefit
A significant factor in subjective scoring is the "matched-accent" or interlanguage speech intelligibility benefit. When an L2 listener evaluates a TTS system that features an accent matching their own native language, their comprehension and subjective ratings often increase.
This occurs because the listener and the accented synthetic voice share similar underlying phonetic categories and phonological rules. For example:
- A Spanish-accented English TTS system may be rated as more intelligible and pleasant by native Spanish speakers than by native English speakers.
- Native English listeners, lacking familiarity with these specific phonetic substitutions, may struggle more with word recognition and assign lower scores.
Conversely, if the synthetic accent does not match the non-native listener's L1, the cognitive load doubles. The listener must process both non-native speech patterns and an unfamiliar accent, which can drive subjective intelligibility scores even lower than those reported by native listeners.
Differing Criteria: Naturalness vs. Intelligibility
Listener status also impacts how distinct evaluation dimensions are weighted:
- Intelligibility (Can the speech be understood?): Non-native listeners often prioritize intelligibility over acoustic realism. If the synthesized voice is clear enough to convey meaning without ambiguity, L2 evaluators are more likely to award favorable scores, even if the accent sounds synthetic or mismatched.
- Naturalness (Does the speech sound human and authentic?): Native listeners weigh naturalness heavily. Even if an accented TTS system achieves near-perfect intelligibility, native raters will significantly downgrade the voice if the accent feels artificial, inconsistent, or caricatured.
Implications for TTS Testing and System Design
The divergence between native and non-native listener scores highlights the danger of using uncharacterized crowdsourced listener pools. If a TTS benchmark fails to balance or record the native-speaker status of its evaluators, the resulting scores can be misleadingly high or unduly critical.
For speech synthesis developers, evaluating accented TTS requires targeting the listener pool to the intended use case. If an accented synthetic voice is deployed for local language learners or regional services, evaluations must incorporate representative L1 and L2 groups. Establishing distinct scoring cohorts ensures that naturalness optimizations do not mask underlying intelligibility barriers across diverse user populations.