Addressing data scarcity in speech emotion recognition: A comprehensive review

Citations

WEB OF SCIENCE

9
Citations

SCOPUS

20

초록

Speech emotion recognition (SER) is a critical field within affective computing, aiming to detect and classify emotional states from speech signals, which vary dynamically over time. These signals encode complex relationships between features at multiple time scales, effectively reflecting a speaker’s emotional state. Despite significant progress, SER faces the persistent challenge of labeled data scarcity, a major obstacle given the data-intensive requirements of deep learning models. This scarcity often results in small, imbalanced datasets that hinder model generalization. Various strategies, including feature selection, data augmentation, domain adaptation, and fusion techniques, have been employed to mitigate these issues. However, comprehensive reviews that critically analyze these methods remain limited. In this paper, we provide an extensive review of these data scarcity strategies in SER, assessing their merits and limitations in terms of efficiency and robustness. Special attention is given to how these strategies enhance the performance of both acoustic and multimodal SER systems when operating on limited datasets. Additionally, we highlight the potential of fusion strategies combined with attention mechanisms as promising solutions to improve convergence and reduce model complexity.

키워드

Emotion recognition; Data scarcity; Limited datasets; Attention mechanisms; SENTIMENT ANALYSIS; FUSION; TEXT
제목
Addressing data scarcity in speech emotion recognition: A comprehensive review
저자
Kakuba, Samuel; Han, Dong Seog
DOI
10.1016/j.icte.2024.11.003
발행일
2025-02
유형
Article
저널명
ICT Express
권
11
호
1
페이지
110 ~ 123