Multi-Modal Integration of 2D and 3D Attributes for Multi-Vehicles Tracking

Citations

SCOPUS

2

초록

Tracking multiple objects is crucial in autonomous vehicles, but relying on one sensor is unreliable due to potential failures in challenging scenarios. 2D cameras provide texture information, whereas LiDAR offers 3D structural data, each excelling under different conditions. Therefore, combining the features of these two sensors is essential for learning distinct characteristics. Effective fusion is challenging because the modal-ities contain fundamentally different data. In this study, we introduce multi-modal integration of point-level and pixel-level features to enhance feature distinctiveness. We utilize VoxelNet for obtaining multi-scale point cloud representations, and ResNet-50 for 2D image-based feature extraction. Additionally, we assess the benefits of pre-training individual modalities followed by fine-tuning the multi-modal. Our technique achieves MOTA 91.28% and 73.53% HOTA on the KITTI dataset, surpassing many methods without multi-modal integration. © 2024 IEEE.

키워드

autonomous vehicles; deep learning; merge features; Multiple object tracking; neural networks
제목
Multi-Modal Integration of 2D and 3D Attributes for Multi-Vehicles Tracking
저자
Altaf, Muhammad Adeel; Kim, Min Young
DOI
10.1109/ICTC62082.2024.10827704
발행일
2024
유형
Conference paper
저널명
International Conference on ICT Convergence
페이지
364 ~ 369