DIAL: Dense Image-Text ALignment for Weakly Supervised Semantic Segmentation

  • Jang, Soojin; 
  • Yun, Jungmin; 
  • Kwon, Junehyoung; 
  • Lee, Eunju; 
  • Kim, Youngbin
Citations

WEB OF SCIENCE

6
Citations

SCOPUS

26

초록

Weakly supervised semantic segmentation (WSSS) approaches typically rely on class activation maps (CAMs) for initial seed generation, which often fail to capture global context due to limited supervision from image-level labels. To address this issue, we introduce DALNet, Dense Alignment Learning Network that leverages text embeddings to enhance the comprehensive understanding and precise localization of objects across different levels of granularity. Our key insight is to employ a dual-level alignment strategy: (1) Global Implicit Alignment (GIA) to capture global semantics by maximizing the similarity between the class token and the corresponding text embeddings while minimizing the similarity with background embeddings, and (2) Local Explicit Alignment (LEA) to improve object localization by utilizing spatial information from patch tokens. Moreover, we propose a cross-contrastive learning approach that aligns foreground features between image and text modalities while separating them from the background, encouraging activation in missing regions and suppressing distractions. Through extensive experiments on the PASCAL VOC and MS COCO datasets, we demonstrate that DALNet significantly outperforms state-of-the-art WSSS methods. Our approach, in particular, allows for more efficient end-to-end process as a single-stage method. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.

키워드

image-level labels supervision; single-stage framework; weakly supervised semantic segmentation
제목
DIAL: Dense Image-Text ALignment for Weakly Supervised Semantic Segmentation
저자
Jang, Soojin; Yun, Jungmin; Kwon, Junehyoung; Lee, Eunju; Kim, Youngbin
DOI
10.1007/978-3-031-72890-7_15
발행일
2025
유형
Proceedings Paper
저널명
Lecture Notes in Computer Science
권
15127
페이지
248 ~ 266