상세 보기
Addressing data scarcity in speech emotion recognition: A comprehensive review
- Kakuba, Samuel;
- Han, Dong Seog
WEB OF SCIENCE
9SCOPUS
20초록
Speech emotion recognition (SER) is a critical field within affective computing, aiming to detect and classify emotional states from speech signals, which vary dynamically over time. These signals encode complex relationships between features at multiple time scales, effectively reflecting a speaker’s emotional state. Despite significant progress, SER faces the persistent challenge of labeled data scarcity, a major obstacle given the data-intensive requirements of deep learning models. This scarcity often results in small, imbalanced datasets that hinder model generalization. Various strategies, including feature selection, data augmentation, domain adaptation, and fusion techniques, have been employed to mitigate these issues. However, comprehensive reviews that critically analyze these methods remain limited. In this paper, we provide an extensive review of these data scarcity strategies in SER, assessing their merits and limitations in terms of efficiency and robustness. Special attention is given to how these strategies enhance the performance of both acoustic and multimodal SER systems when operating on limited datasets. Additionally, we highlight the potential of fusion strategies combined with attention mechanisms as promising solutions to improve convergence and reduce model complexity.
키워드
- 제목
- Addressing data scarcity in speech emotion recognition: A comprehensive review
- 저자
- Kakuba, Samuel; Han, Dong Seog
- 발행일
- 2025-02
- 유형
- Article
- 저널명
- ICT Express
- 권
- 11
- 호
- 1
- 페이지
- 110 ~ 123
- 언어
- ENG
- 출판사
- ELSEVIER
- 발행국가
- 네덜란드
- 분량
- 14 페이지
- ISSN
- E 2405-9595
P 2405-9595