Multimodal feature fusion for human activity recognition using human centric temporal transformer

Citations

WEB OF SCIENCE

6
Citations

SCOPUS

16

초록

In recent years, human activity recognition (HAR) has focused considerable interest due to its manifold monitoring applications. Mainstream HAR approaches often face challenges with the reliability of results when relying on a single data modality, especially when integrating heterogeneous data sources. A notable limitation of the implemented artificial intelligence (AI) models is their limited capability to handle dynamic scenarios, as they lack the necessary contextual information from multiple sources, which impedes the models' adaptability and accuracy. This paper proposes a multi-modality framework for HAR that fuses human concern patterns using various spatiotemporal model flavors. In addition, to get the spatial features, a swin transformer with a dual attention concept is applied to process visual sensor data, while one-dimensional convolutional neural network leverages human skeleton information obtained from the detection model with numerous key points. Later, these multi-modality features are fused to improve the robust analysis and comprehension of activities. Next, these resulting features are passed to the human centric temporal transformer (HCTT), that has the capabilities to process multimodal sequence data for temporal learning. Moreover, the attention block of HCTT enables human-related attentive patterns followed by a dual fusion mechanism. The proposed model was evaluated on four open-access large-scale HAR datasets, where comprehensive ablation studies and comparative analyses demonstrated that our developed multimodal approach outperforms recent baseline HAR models. This underscores its potential for advancing AI applications and human activity analysis.

키워드

Vision transformer; Multi modality; Human activity recognition; Surveillance data; Attention module; Human centric transformer; Implemented artificial intelligence; Application of artificial intelligence; Urban safety; LSTM; INTERNET; FLOW; CNN
제목
Multimodal feature fusion for human activity recognition using human centric temporal transformer
저자
Khan, Samee Ullah; Sultana, Maryam; Danish, Sufyan; Gupta, Deepak; Alghamdi, Norah Saleh; Woo, Suchang; Lee, Dong-Gyu; Ahn, Sangtae
DOI
10.1016/j.engappai.2025.111844
발행일
2025-11-23
유형
Article
저널명
Engineering Applications of Artificial Intelligence
권
160