A Region Descriptive Pre-training Approach with Self-attention Towards Visual Question Answering

  • Kolawole, Bisi Bode; 
  • Lee, Minho
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Concatenation of text (question-answer) and image has been the bedrock of most visual language systems. Existing models concatenate the text (question-answer) and image inputs in a forced manner. In this paper, we introduce a region descriptive pre-training approach with self-attention towards VQA. The model is a new learning method that uses the image region descriptions combined with object labels to create a proper alignment between the text(question-answer) and the image inputs. We study the text associated with each image and discover that extracting the region descriptions from the image and using it during training greatly improves the model's performance. In this research work, we use the region description extracted from the images as a bridge to map the text and image inputs. The addition of region description makes our model perform better against some recent state-of-the-art models. Experiments demonstrated in this paper show that our model significantly outperforms most of these models.

키워드

Visual question answering; Region descriptions; Object label; Pre-training
제목
A Region Descriptive Pre-training Approach with Self-attention Towards Visual Question Answering
저자
Kolawole, Bisi Bode; Lee, Minho
DOI
10.1007/978-3-030-92310-5_9
발행일
2022
유형
Proceedings Paper
저널명
Communications in Computer and Information Science
권
1517
페이지
73 ~ 80