Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions

Citations

WEB OF SCIENCE

4
Citations

SCOPUS

12

초록

In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional captions often have struggled to capture the rich semantics, including temporal changes, inherent in the video. In addition, incorrect information caused by generative models can lead to inaccurate retrieval. To address these issues, we propose a new framework, Narrating the Video (NarVid), which strategically leverages the comprehensive information available from frame-level captions, the narration. The proposed NarVid exploits narration in multiple ways: 1) feature enhancement through cross-modal interactions between narration and video, 2) query-aware adaptive filtering to suppress irrelevant or incorrect information, 3) dual-modal matching score by adding query-video similarity and query-narration similarity, and 4) hard-negative loss to learn discriminative features from multiple perspectives using the two similarities from different views. Experimental results demonstrate that NarVid achieves state-of-the-art performance on various benchmark datasets. © 2025 IEEE.

키워드

multimodal retrieval; text-video retrieval
제목
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
저자
Hur, Chan; Hong, Jeong-hun; Lee, Donghun; Kang, Dabin; Myeong, Semin; Park, S.; Park, Hyeyoung
DOI
10.1109/CVPR52734.2025.02242
발행일
2025
유형
Proceedings Paper
저널명
Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
페이지
24077 ~ 24086