상세 보기
초록
Despite its widespread success on vision tasks, standard attention employs a shared dot-product mechanism that uniformly scores all query–key interactions before applying softmax. In this paper, we hypothesize that explicitly controlling the amplification or suppression of individual query-key token-pair interactions can lead to more expressive and discriminative representations. To this end, we propose Pairwise Logit Gating (PLuG) attention, a simple yet effective plug-in approach that introduces a learnable gating mechanism operating on each token-pair to modulate attention logits prior to softmax. This gating enables the model to selectively amplify informative interactions and suppress spurious ones through gating coefficient matrix, improving its ability to capture spatial and semantic relationships critical for vision tasks. Experimental results demonstrate that PLuG can be seamlessly integrated into standard attention mechanisms across various vision tasks by modifying only the attention module. In particular, PLuG consistently improves strong baselines, boosting DeiT-Ti top-1 accuracy from 72.2% to 73.2%, increasing Mask2Former mIoU from 47.7 to 48.7, and reducing DiT-S/2 FID from 67.40 to 66.15, highlighting the effectiveness of PLuG as a practical plug-in enhancement for attention-based vision models.
키워드
- 제목
- PLuG: Pairwise Logit Gating for Expressive Attention Modulation in Vision Transformers
- 저자
- Lee, Dongheon; Yoon, Sangwoo; Kim, Seongsu; Paik, Joonki
- 발행일
- 2026
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 87364 ~ 87376