Hierarchical Action Understanding : Fine-to-Coarse Reasoning Framework for Video Interpretation

  • Moon, JunBeom; 
  • Won, Jiye; 
  • Joo, YeEun; 
  • Heo, Sehwan; 
  • Jung, Soon Ki
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Existing video action recognition methods often fail to bridge the semantic gap between fine-grained, frame-level labels and coarse, high-level actions. To address this, we propose a hierarchical reasoning framework that predicts coarse action labels from sequences of fine-grained labels. Using the Breakfast dataset, we apply rule-based preprocessing to align and pair fine and coarse labels, resolving temporal misalignment and background segments. Our model-agnostic framework supports integration with outputs from Temporal Action Segmentation (TAS) models. We evaluate sequence models-LSTM, TCN, Transformer, and Mamba-under causal and non-causal settings using multiple loss functions, including cross-entropy and cosine similarity. Results show that causal models, particularly LSTM and Mamba, outperform others in accuracy, edit distance, and F1-score, especially with hybrid losses. Our method is robust to noisy fine labels and preserves interpretability through explicit fine-to-coarse mapping. This work offers a scalable and modular solution for multi-level action understanding across diverse video domains.

제목
Hierarchical Action Understanding : Fine-to-Coarse Reasoning Framework for Video Interpretation
저자
Moon, JunBeom; Won, Jiye; Joo, YeEun; Heo, Sehwan; Jung, Soon Ki
DOI
10.1109/AVSS65446.2025.11149940
발행일
2025
유형
Proceedings Paper
저널명
2025 IEEE INTERNATIONAL CONFERENCE ON ADVANCED VISUAL AND SIGNAL-BASED SYSTEMS, AVSS
호
2025