Multi-modal CrossViT using 3D spatial information for visual localization

Citations

SCOPUS

1

초록

Visual Localization entails estimating the position and orientation of a camera from input images. Since it is pivotal to robotics, autonomous vehicles, and augmented reality fields, improvements to its accuracy and computational efficiency are vital. Although several approaches to hierarchical visual localization have been proposed, the convolutional operations in their global localization stage inflate their computational requirements. This study proposes a hierarchical framework comprised of a multi-modal CrossViT (Vision Transformer) that leverages both image features and 3D spatial information to generate more robust global descriptors. In the contrastive learning approach employed, the positive and negative image sets for each anchor image are designated based on the presence of shared 3D points. The intersection-over-union between 3D bounding boxes generated from a pair of images is used as a quantitative measure of similarity for positive sets in the loss computation. The embedding capacity of the proposed multi-modal CrossViT is transferred onto an architecture that takes a single image as input using a knowledge distillation approach. Local matching models are used to establish correspondences between the anchor image and each of the retrieved reference images. The final camera pose is determined using the random sample consensus and perspective-n-point algorithm. The large-scale Aachen Day-Night datasets were used to evaluate the efficiency and accuracy of the proposed approach. Experimental results show that the proposed approach achieves performance comparable to that of previous state-of-the-art approaches with significantly less processing and memory requirements (58.9 times fewer per-second floating-point operations and 21.6 times fewer parameters than the NetVLAD model). © The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2024.

키워드

Deep metric learningKnowledge distillationMulti-modalVisual localization
제목
Multi-modal CrossViT using 3D spatial information for visual localization
저자
Kang, JunekooMpabulungi, MarkHong, Hyunki
DOI
10.1007/s11042-024-20382-w
발행일
2025-02
유형
Article
저널명
Multimedia Tools and Applications
84
5
페이지
2059 ~ 2083