PLuG: Pairwise Logit Gating for Expressive Attention Modulation in Vision Transformers

  • Lee, Dongheon
  • Yoon, Sangwoo
  • Kim, Seongsu
  • Paik, Joonki
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Despite its widespread success on vision tasks, standard attention employs a shared dot-product mechanism that uniformly scores all query–key interactions before applying softmax. In this paper, we hypothesize that explicitly controlling the amplification or suppression of individual query-key token-pair interactions can lead to more expressive and discriminative representations. To this end, we propose Pairwise Logit Gating (PLuG) attention, a simple yet effective plug-in approach that introduces a learnable gating mechanism operating on each token-pair to modulate attention logits prior to softmax. This gating enables the model to selectively amplify informative interactions and suppress spurious ones through gating coefficient matrix, improving its ability to capture spatial and semantic relationships critical for vision tasks. Experimental results demonstrate that PLuG can be seamlessly integrated into standard attention mechanisms across various vision tasks by modifying only the attention module. In particular, PLuG consistently improves strong baselines, boosting DeiT-Ti top-1 accuracy from 72.2% to 73.2%, increasing Mask2Former mIoU from 47.7 to 48.7, and reducing DiT-S/2 FID from 67.40 to 66.15, highlighting the effectiveness of PLuG as a practical plug-in enhancement for attention-based vision models.

키워드

attention logit modulationAttention mechanismimage classificationimage generationsemantic segmentation Ivision transformer
제목
PLuG: Pairwise Logit Gating for Expressive Attention Modulation in Vision Transformers
저자
Lee, DongheonYoon, SangwooKim, SeongsuPaik, Joonki
DOI
10.1109/ACCESS.2026.3701015
발행일
2026
유형
Article
저널명
IEEE Access
14
페이지
87364 ~ 87376

파일 다운로드