Adaptive Attention Link-based Regularization for Vision Transformers

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Transformer networks are recently employed in various vision tasks with outperforming performance, but they require extensive training data and a lengthy training time to train a model to disregard an inductive bias. In this paper, we present an adaptive regularization technique to improve the training efficiency of ViT in terms of training time and data quantity. Our technique utilizes trainable links between the channel-wise spatial attention of a pre-trained Convolution Neural Network (CNN) and the attention head of Vision Transformers (ViT), which is called the attention augmentation module. The attention augmentation module is trained simultaneously with ViT, so automatically discovers the complicated attention similarity between levels of CNN and ViT models. By considering the trained attention similarity, even with a small amount of data, we validate that the suggested method considerably improves the performance of ViT while achieving faster convergence during training. Additionally, we can also obtain a meaningful analysis of the relevant relationship between each CNN activation map and each ViT attention head from the trained links of the attention augmentation module.

제목
Adaptive Attention Link-based Regularization for Vision Transformers
저자
Jin, Heegon; Choi, Jongwon
DOI
10.1109/AVSS65446.2025.11149958
발행일
2025-08
유형
Proceedings Paper
저널명
Proceedings - IEEE International Conference on Advanced Video and Signal-Based Surveillance, AVSS