Deep learning-guided video compression for machine vision tasks

  • Kim, Aro; 
  • Woo, Seung-taek; 
  • Park, Minho; 
  • Kim, Dong-hwi; 
  • Lim, Hanshin; 
  • ... Park, Sang-hyo; 
  • 외 2명
Citations

WEB OF SCIENCE

2
Citations

SCOPUS

4

초록

In the video compression industry, video compression tailored to machine vision tasks has recently emerged as a critical area of focus. Given the unique characteristics of machine vision, the current practice of directly employing conventional codecs reveals inefficiency, which requires compressing unnecessary regions. In this paper, we propose a framework that more aptly encodes video regions distinguished by machine vision to enhance coding efficiency. For that, the proposed framework consists of deep learning-based adaptive switch networks that guide the efficient coding tool for video encoding. Through the experiments, it is demonstrated that the proposed framework has superiority over the latest standardization project, video coding for machine benchmark, which achieves a Bjontegaard delta (BD)-rate gain of 5.91% on average and reaches up to a 19.51% BD-rate gain.

키워드

Video compression; Video coding for machines; Deep learning
제목
Deep learning-guided video compression for machine vision tasks
저자
Kim, Aro; Woo, Seung-taek; Park, Minho; Kim, Dong-hwi; Lim, Hanshin; Jung, Soon-heung; Kwak, Sangwoon; Park, Sang-hyo
DOI
10.1186/s13640-024-00649-w
발행일
2024-09-20
유형
Article
저널명
Eurasip Journal on Image and Video Processing
권
2024
호
1