Two-Stage Speech Enhancement Framework: Balancing Spectral Suppression and Compensation

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Speech enhancement (SE) aims to preserve speech components while removing noise that degrades signal quality. However, existing methods often distort speech components by over-suppressing noise. This paper proposes a two-stage SE framework called the suppression and compensation network (SCNet) that estimates high-quality target speech by balancing noise reduction and speech preservation. In the first stage, the noise suppression module (NS-Module) removes major noise components by modeling a complex mask function that estimates both magnitude and phase information. The NS-Module distinguishes speech from noise based on the regular energy distribution of the speech spectrum, utilizing a time-frequency convolutional module (TFCM) that models long-range dependencies across both time and frequency dimensions. In the second stage, the spectral detail compensation module (SDC-Module) restores missing spectral details by focusing on the difference between the original spectrum and the enhanced spectrum from the previous stage. Additionally, the loss function includes an asymmetric component that penalizes over-suppressed spectra, guiding the model toward an optimal balance. The experimental results on the DNS-Challenge and VoiceBank + DEMAND datasets demonstrate that the proposed method more effectively captures and restores the inherent spectral structure of speech compared to existing SE models. © 2013 IEEE.

키워드

compensation; Deep learning; masking; speech enhancement; time-frequency convolution; two-stage network
제목
Two-Stage Speech Enhancement Framework: Balancing Spectral Suppression and Compensation
저자
Il Koh, Hyeong; Nam Kim, Myoung; Na, Sungdae
DOI
10.1109/ACCESS.2025.3635221
발행일
2025-11
유형
Article
저널명
IEEE Access
권
13
페이지
198347 ~ 198357