Vision Transformer Compression and Architecture Exploration with Efficient Embedding Space Search

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

This paper addresses theoretical and practical problems in the compression of vision transformers for resource-constrained environments. We found that deep feature collapse and gradient collapse can occur during the search process for the vision transformer compression. Deep feature collapse diminishes feature diversity rapidly as the layer depth deepens, and gradient collapse causes gradient explosion in training. Against these issues, we propose a novel framework, called VTCA, for accomplishing vision transformer compression and architecture exploration jointly with embedding space search using Bayesian optimization. In this framework, we formulate block-wise removal, shrinkage, cross-block skip augmentation to prevent deep feature collapse, and Res-Post layer normalization to prevent gradient collapse under a knowledge distillation loss. In the search phase, we adopt a training speed estimation for a large-scale dataset and propose a novel elastic reward function that can represent a generalized manifold of rewards. Experiments were conducted with DeiT-Tiny/Small/Base backbones on the ImageNet, and our approach achieved competitive accuracy to recent patch reduction and pruning methods. The code is available at https://github.com/kdaeho27/VTCA.

제목
Vision Transformer Compression and Architecture Exploration with Efficient Embedding Space Search
저자
Kim, Daeho; Kim, Jaeil
DOI
10.1007/978-3-031-26313-2_32
발행일
2023
유형
Proceedings Paper
저널명
Lecture Notes in Computer Science
권
13843
페이지
524 ~ 540