A Self-Attention Classifier Head for Improved Image Classification and Interpretability of ViT

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

The transformer architecture, particularly the vision transformer (ViT), has demonstrated remarkable success in various computer vision tasks. While numerous studies have explored architectural modifications within the transformer backbone, the final classifier head has received less attention. Conventional ViT classifier heads typically employ a simple linear layer acting on a class token or globally pooled features. This letter introduces a self-attention classifier head (SACH) that replaces the standard linear classifier with a self-attention layer. SACH enables the model to adaptively weigh the importance of different token representations from the transformer backbone, facilitating a more nuanced and powerful feature aggregation. This mechanism is particularly effective when training models from scratch, where learning robust representations is most critical. We demonstrate that SACH consistently and significantly improves performance in this setting, achieving a relative error rate reduction of up to 13.02% on SVHN with ViT-Tiny and 25.52% on CIFAR-10 with ViT-Base. These improvements are achieved with only a marginal increase in computational overhead, adding approximately 2% to both parameters and FLOPs across various model scales. Furthermore, the attention mechanism in SACH offers enhanced interpretability; its attention maps can be combined with those from the ViT backbone to provide clearer visualisations of the model's focus. In transfer learning scenarios, SACH demonstrates robust generalisation by consistently delivering performance gains, in contrast to methods like global average pooling, which can exhibit unstable results and sometimes degrade performance.

키워드

computer vision; image classification; learning (artificial intelligence); neural net architecture
제목
A Self-Attention Classifier Head for Improved Image Classification and Interpretability of ViT
저자
Ahn, Jaesin; Jung, Heechul
DOI
10.1049/ell2.70478
발행일
2025-11-13
유형
Article
저널명
Electronics Letters
권
61
호
1