상세 보기
Deflationary Extraction Transformer for Speech Separation with Unknown Number of Talkers
- Lee, Sangwon;
- Kim, Han-Gyu;
- Jang, Gil-Jin
WEB OF SCIENCE
0SCOPUS
0초록
Most speech separation techniques require knowing the number of talkers mixed in an input, which is not always available in real situations. To address this problem, we present a novel speech separation method that automatically finds the number of talkers in input mixture recordings. The proposed method extracts the voices of individual talkers one by one in a deflationary manner and stops the extraction sequence when a predefined termination criterion is satisfied. The backbone separation model is built based on the transformer architecture with permutation-invariant training to avoid ambiguity in identifying talkers at the output. The experimental results on the Libri5Mix and Libri10Mix datasets show that the proposed method without the number of talkers as input significantly outperforms state-of-the-art models that are provided with the number of talkers.
키워드
- 제목
- Deflationary Extraction Transformer for Speech Separation with Unknown Number of Talkers
- 저자
- Lee, Sangwon; Kim, Han-Gyu; Jang, Gil-Jin
- 발행일
- 2025-08-08
- 유형
- Article
- 저널명
- Sensors
- 권
- 25
- 호
- 16
- 언어
- ENG
- 출판사
- MDPI
- 발행국가
- 스위스
- ISSN
- E 1424-8220
P 1424-8220