Implementation of a Virtual Assistant System Based on Deep Multi-modal Data Integration

  • Baek, Sungdae; 
  • Kim, Jonghong; 
  • Lee, Junwon; 
  • Lee, Minho
Citations

WEB OF SCIENCE

2
Citations

SCOPUS

5

초록

In this study, we propose a virtual assistant system that is applied to real life using signal processing and deep learning. First, the overall structure of the proposed system that integrates and controls various modules is introduced, after which we present a multi-modal fusion module that provides services to users. It integrates a natural language processing module for interpreting Korean chatbots and a behavior recognition module for understanding user behavior using an RGB camera. In addition, a hand gesture recognition module was utilized to understand the user's intentions using depth and RGB images. We explain the implementation of a customized service system with several parts: i) a user interface module that interacts with the user, ii) a face recognition module that distinguishes different users, and iii) a voice processing module that can replace the input and output methods through a keyboard and monitor. To check the performance of each module, a testbed was configured in an office environment. Through test results, we successfully demonstrate the realization of the proposed system in real life Finally, we list the challenges discovered during the operation of this system and suggest directions for further research.

키워드

Integrated system; Transfer learning; Natural language processing; Image processing; Action recognition; Gesture recognition; CONVOLUTIONAL NEURAL-NETWORK; RECOGNITION; ATTENTION
제목
Implementation of a Virtual Assistant System Based on Deep Multi-modal Data Integration
저자
Baek, Sungdae; Kim, Jonghong; Lee, Junwon; Lee, Minho
DOI
10.1007/s11265-022-01829-5
발행일
2024-03
유형
Article
저널명
Journal of Signal Processing Systems
권
96
호
3
페이지
179 ~ 189