A study on the application of residual vector quantization for vector quantized-variational autoencoder-based foley sound generation model

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Among the Foley sound generation models that have recently begun to be studied, a sound generation technique using the Vector Quantized-Variational AutoEncoder (VQ-VAE) structure and generation model such as Pixelsnail are one of the important research subjects. On the other hand, in the field of deep learning-based acoustic signal compression, residual vector quantization technology is reported to be more suitable than the conventional VQ-VAE structure. Therefore, in this paper, we aim to study whether residual vector quantization technology can be effectively applied to the Foley sound generation. In order to tackle the problem, this paper applies the residual vector quantization technique to the conventional VQ-VAE-based Foley sound generation model, and in particular, derives a model that is compatible with the existing models such as Pixelsnail and does not increase computational resource consumption. In order to evaluate the model, an experiment was conducted using DCASE2023 Task7 data. The results show that the proposed model enhances about 0.3 of the Frechet audio distance. Unfortunately, the performance enhancement was limited, which is believed to be due to the decrease in the resolution of time-frequency domains in order to do not increase of the resources.

키워드

Foley sound generation model; Vector Quantized-Variational AutoEncoder (VQ-VAE); Residual vector quantization; Generative Model
제목
A study on the application of residual vector quantization for vector quantized-variational autoencoder-based foley sound generation model
저자
Lee, Seokjin
DOI
10.7776/ASK.2024.43.2.243
발행일
2024-03
유형
Article
저널명
한국음향학회지
권
43
호
2
페이지
243 ~ 252