Deflationary Extraction Transformer for Speech Separation with Unknown Number of Talkers

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Most speech separation techniques require knowing the number of talkers mixed in an input, which is not always available in real situations. To address this problem, we present a novel speech separation method that automatically finds the number of talkers in input mixture recordings. The proposed method extracts the voices of individual talkers one by one in a deflationary manner and stops the extraction sequence when a predefined termination criterion is satisfied. The backbone separation model is built based on the transformer architecture with permutation-invariant training to avoid ambiguity in identifying talkers at the output. The experimental results on the Libri5Mix and Libri10Mix datasets show that the proposed method without the number of talkers as input significantly outperforms state-of-the-art models that are provided with the number of talkers.

키워드

speech separation; speaker diarization; SepFormer; permutation-invariant training; Conv-TasNet
제목
Deflationary Extraction Transformer for Speech Separation with Unknown Number of Talkers
저자
Lee, Sangwon; Kim, Han-Gyu; Jang, Gil-Jin
DOI
10.3390/s25164905
발행일
2025-08-08
유형
Article
저널명
Sensors
권
25
호
16