상세 보기
Speech Emotion Recognition using Context-Aware Dilated Convolution Network
- Kakuba, Samuel;
- Han, Dong Seog
WEB OF SCIENCE
9SCOPUS
15초록
Deep learning-based speech emotion recognition has been applied for social living assistance, health monitoring, authentication, and other human-to-machine interaction applications. Because of the ubiquitous nature of the applications, computationally efficient and robust speech emotion recognition models are required. The nature of the speech signal requires tracking of time steps, analyzing long-term dependencies and the contexts of the utterances as well as the spatial cues. Recurrent neural networks like long short-term memory and gated recurrent units coupled with attention mechanisms are often used to consider long-term dependencies and context in the speech signal. However, they do not take care of the spatial cues that may exist in the speech signal. Moreover, the operation of most of these systems is sequential which causes slow convergence, and sluggish training. Therefore, we propose a model that employs dilated convolutions layers in combination with hybrid attention mechanisms. The model uses multi-head attention to extract the global context in the feature representations which are fed into the bidirectional long short-term memory configured with self-attention to further handle the context and long-term dependencies. The model uses spectral and voice quality features extracted from the raw speech signals as input. The proposed model achieves comparable performance in terms of F1 score and accuracy. The proposed model's performance is also presented in terms of confusion matrices.
키워드
- 제목
- Speech Emotion Recognition using Context-Aware Dilated Convolution Network
- 저자
- Kakuba, Samuel; Han, Dong Seog
- 발행일
- 2022
- 유형
- Proceedings Paper
- 저널명
- 2022 27TH ASIA PACIFIC CONFERENCE ON COMMUNICATIONS (APCC 2022): CREATING INNOVATIVE COMMUNICATION TECHNOLOGIES FOR POST-PANDEMIC ERA
- 페이지
- 601 ~ 604
- 언어
- ENG
- 출판사
- IEEE
- 발행국가
- 미국
- 분량
- 4 페이지