Annotating Emotion Tags in TTS Speech Databases

This article provides an overview of the primary methodologies used to annotate subjective emotional tags in expressive speech databases for Text-to-Speech (TTS) synthesis. Developing natural, expressive synthetic speech requires reliable affective labeling, which presents challenges due to the inherently subjective nature of human emotion. The sections below outline the dominant annotation frameworks, measurement dimensions, annotator workflows, and quality-control techniques used to transform subjective human perception into structured data suitable for training modern neural TTS models.

Categorical Emotion Annotation

The categorical approach assigns discrete emotion labels to speech utterances based on predefined taxonomies.

Dimensional Emotion Models

Dimensional modeling avoids discrete boundaries by positioning emotional states along continuous axes. This method provides the fine-grained, continuous control needed for latent emotion embeddings in neural TTS.

Utterance-Level vs. Continuous Time-Series Annotation

Annotation methodologies vary based on temporal granularity:

Annotator Sourcing and Bias Mitigation

Because emotional perception varies based on culture, gender, and individual context, database curators use distinct sourcing methodologies:

Machine-Assisted and Semi-Supervised Annotation

Given the cost and fatigue associated with human listening tests, hybrid methodologies are common: