Activation-driven discrepancy as evidence for generalized neural interpretability

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Architectural and operational variances among Deep Neural Networks (DNNs) limit the universal applicability of existing activation-based explainability methods. In this paper, we propose a model-agnostic framework for interpreting network decisions by clarifying the gap between the transferred activations and those of the original model. Notably, even when given an empty input, a model can produce a plausible prediction based solely on the activation information retained from the original prediction. Our motivation starts from this phenomenon, assuming that an incomplete prediction contains a rough representation of neural information that is approximate to the original model's prediction, yet missing key details. Conversely, clarifying the distributional gap between these activation differences could help identify crucial factors influencing model predictions. To address this, we introduce Activation Presence Transfer (APT), a method that transfers neuron activation information to a cloned model, preserving the nature of the original model's prediction when given a blank image. Next, we propose the Residual Attribution Map (RAM), which quantifies distributional discrepancies across layers and generates integrated attribution maps using both weights and gradients. In a verified experimental setup, we conduct comparative analyses against existing methods on various CNN and Transformer models to validate its model-agnostic nature. The experimental results demonstrate superior performance compared to existing approaches that are specifically tailored to particular model types.

키워드

Explainable AI; Explanation mechanism; Model agnostic; Activation; CAM VISUAL EXPLANATIONS; NETWORKS; ATTENTION
제목
Activation-driven discrepancy as evidence for generalized neural interpretability
저자
Shin, Ho Kyung; Nam, Woo-Jeoung
DOI
10.1016/j.eswa.2025.130001
발행일
2026-03-01
유형
Article
저널명
Expert Systems with Applications
권
299