SAM: cross-modal semantic alignments module for image-text retrieval

Citations

WEB OF SCIENCE

11
Citations

SCOPUS

11

초록

Cross-modal image-text retrieval has gained increasing attention due to its ability to combine computer vision with natural language processing. Previously, image and text features were extracted and concatenated to feed the transformer-based retrieval network. However, these approaches implicitly aligned the image and text modalities since the self-attention mechanism computes attention coefficients for all input features. In this paper, we propose cross-modal Semantic Alignments Module (SAM) to establish an explicit alignment through enhancing an inter-modal relationship. Firstly, visual and textual representations were extracted from an image and text pair. Secondly, we constructed a bipartite graph by representing the image regions and words in the sentence as nodes, and the relationship between them as edges. Then our proposed SAM allows the model to compute attention coefficients based on the edges in the graph. This process helps explicitly align the two modalities. Finally, a binary classifier was used to determine whether the given image-text pair is aligned. We reported extensive experiments on MS-COCO and Flickr30K test sets, showing that SAM could capture the joint representation between the two modalities and could be applied to the existing retrieval networks. © 2023, The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature.

키워드

Cross-modalGraph neural networksImage-text retrievalVision-language
제목
SAM: cross-modal semantic alignments module for image-text retrieval
저자
Park, PilseoJang, SoojinCho, YunsungKim, Youngbin
DOI
10.1007/s11042-023-15798-9
발행일
2024-01
유형
Article
저널명
Multimedia Tools and Applications
83
4
페이지
12363 ~ 12377