텍스트-비디오 검색 모델에서의 캡션을 활용한 비디오 특성 대체 방안 연구

A Study on the Alternative Method of Video Characteristics Using Captioning in Text-Video Retrieval Model

초록

In this paper, we propose a method that performs a text-video retrieval model by replacing video properties using captions. In general, the exisiting embedding-based models consist of both joint embedding space construction and the CNN-based video encoding process, which requires a lot of computation in the training as well as the inference process. To overcome this problem, we introduce a video-captioning module to replace the visual property of video with captions generated by the video-captioning module. To be specific, we adopt the caption generator that converts candidate videos into captions in the inference process, thereby enabling direct comparison between the text given as a query and candidate videos without joint embedding space. Through the experiment, the proposed model successfully reduces the amount of computation and inference time by skipping the visual processing process and joint embedding space construction on two benchmark dataset, MSR-VTT and VATEX.

키워드

Multimodal Deep Learning; Video-Captioning; Text-Video Retrieval
제목
텍스트-비디오 검색 모델에서의 캡션을 활용한 비디오 특성 대체 방안 연구
제목 (타언어)
A Study on the Alternative Method of Video Characteristics Using Captioning in Text-Video Retrieval Model
저자
이동훈; 허찬; 박혜영; 박상효
DOI
10.14372/IEMEK.2022.17.6.347
발행일
2022-12
유형
Y
저널명
대한임베디드공학회논문지
권
17
호
6
페이지
347 ~ 353