상세 보기
Activation-driven discrepancy as evidence for generalized neural interpretability
- Shin, Ho Kyung;
- Nam, Woo-Jeoung
WEB OF SCIENCE
0SCOPUS
0초록
Architectural and operational variances among Deep Neural Networks (DNNs) limit the universal applicability of existing activation-based explainability methods. In this paper, we propose a model-agnostic framework for interpreting network decisions by clarifying the gap between the transferred activations and those of the original model. Notably, even when given an empty input, a model can produce a plausible prediction based solely on the activation information retained from the original prediction. Our motivation starts from this phenomenon, assuming that an incomplete prediction contains a rough representation of neural information that is approximate to the original model's prediction, yet missing key details. Conversely, clarifying the distributional gap between these activation differences could help identify crucial factors influencing model predictions. To address this, we introduce Activation Presence Transfer (APT), a method that transfers neuron activation information to a cloned model, preserving the nature of the original model's prediction when given a blank image. Next, we propose the Residual Attribution Map (RAM), which quantifies distributional discrepancies across layers and generates integrated attribution maps using both weights and gradients. In a verified experimental setup, we conduct comparative analyses against existing methods on various CNN and Transformer models to validate its model-agnostic nature. The experimental results demonstrate superior performance compared to existing approaches that are specifically tailored to particular model types.
키워드
- 제목
- Activation-driven discrepancy as evidence for generalized neural interpretability
- 저자
- Shin, Ho Kyung; Nam, Woo-Jeoung
- 발행일
- 2026-03-01
- 유형
- Article
- 권
- 299
- 언어
- ENG
- 출판사
- PERGAMON-ELSEVIER SCIENCE LTD
- 발행국가
- 영국
- ISSN
- E 1873-6793
P 0957-4174