BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification

  • Kim, June-Woo; 
  • Toikkanen, Miika; 
  • Choi, Yera; 
  • Moon, Seoung-Eun; 
  • June, Ho-Young
Citations

WEB OF SCIENCE

12
Citations

SCOPUS

18

초록

Respiratory sound classification (RSC) is challenging due to varied acoustic signatures, primarily influenced by patient demographics and recording environments. To address this issue, we introduce a text-audio multimodal model that utilizes metadata of respiratory sounds, which provides useful complementary information for RSC. Specifically, we fine-tune a pretrained text-audio multimodal model using free-text descriptions derived from the sound samples' metadata which includes the gender and age of patients, type of recording devices, and recording location on the patient's body. Our method achieves state-of-the-art performance on the ICBHI dataset, surpassing the previous best result by a notable margin of 1.17%. This result validates the effectiveness of leveraging metadata and respiratory sound samples in enhancing RSC performance. Additionally, we investigate the model performance in the case where metadata is partially unavailable, which may occur in real-world clinical setting.

키워드

Respiratory Sound Classification; Pretrained Language-Audio Model; ICBHI; Metadata
제목
BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification
저자
Kim, June-Woo; Toikkanen, Miika; Choi, Yera; Moon, Seoung-Eun; June, Ho-Young
DOI
10.21437/Interspeech.2024-492
발행일
2024
유형
Proceedings Paper
저널명
INTERSPEECH 2024
페이지
1690 ~ 1694