상세 보기
RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification
- Kim, June-woo;
- Toikkanen, Miika;
- Bae, Sangmin;
- Kim, Minseok;
- Jung, Ho-young
SCOPUS
12초록
Recent advancements in AI have democratized its deployment as a healthcare assistant. While pretrained models from large-scale visual and audio datasets have demonstrably generalized to this task, surprisingly, no studies have explored pretrained speech models, which, as human-originated sounds, intuitively would share closer resemblance to lung sounds. This paper explores the efficacy of pretrained speech models for respiratory sound classification. We find that there is a characterization gap between speech and lung sound samples, and to bridge this gap, data augmentation is essential. However, the most widely used augmentation technique for audio and speech, SpecAugment, requires 2-dimensional spectrogram format and cannot be applied to models pretrained on speech waveforms. To address this, we propose RepAugment, an input-agnostic representation-level augmentation technique that outperforms SpecAugment, but is also suitable for respiratory sound classification with waveform pretrained models. Experimental results show that our approach outperforms the SpecAugment, demonstrating a substantial improvement in the accuracy of minority disease classes, reaching up to 7.14%. © 2024 IEEE.
- 제목
- RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification
- 저자
- Kim, June-woo; Toikkanen, Miika; Bae, Sangmin; Kim, Minseok; Jung, Ho-young
- 발행일
- 2024
- 유형
- Conference paper
- 언어
- ENG
- 출판사
- Institute of Electrical and Electronics Engineers Inc.
- ISSN
- E 0589-1019
P 1557-170X