상세 보기
A study on the application of residual vector quantization for vector quantized-variational autoencoder-based foley sound generation model
WEB OF SCIENCE
0SCOPUS
0초록
Among the Foley sound generation models that have recently begun to be studied, a sound generation technique using the Vector Quantized-Variational AutoEncoder (VQ-VAE) structure and generation model such as Pixelsnail are one of the important research subjects. On the other hand, in the field of deep learning-based acoustic signal compression, residual vector quantization technology is reported to be more suitable than the conventional VQ-VAE structure. Therefore, in this paper, we aim to study whether residual vector quantization technology can be effectively applied to the Foley sound generation. In order to tackle the problem, this paper applies the residual vector quantization technique to the conventional VQ-VAE-based Foley sound generation model, and in particular, derives a model that is compatible with the existing models such as Pixelsnail and does not increase computational resource consumption. In order to evaluate the model, an experiment was conducted using DCASE2023 Task7 data. The results show that the proposed model enhances about 0.3 of the Frechet audio distance. Unfortunately, the performance enhancement was limited, which is believed to be due to the decrease in the resolution of time-frequency domains in order to do not increase of the resources.
키워드
- 제목
- A study on the application of residual vector quantization for vector quantized-variational autoencoder-based foley sound generation model
- 저자
- Lee, Seokjin
- 발행일
- 2024-03
- 유형
- Article
- 저널명
- 한국음향학회지
- 권
- 43
- 호
- 2
- 페이지
- 243 ~ 252
- 언어
- ENG
- 출판사
- ACOUSTICAL SOC KOREA
- 발행국가
- 대한민국
- 분량
- 10 페이지
- ISSN
- E 2287-3775
P 1225-4428