상세 보기
APT : A Two-stage Image-Video Transfer Learning for Video Retrieval
- Kang, Seong-Min;
- Cho, Yoon-Sik
SCOPUS
0초록
Image-text pre-trained models, such as CLIP, have gained significant traction in the field of video-text learning. Recent approaches have extended these models to video tasks, achieving unprecedented performance on the foundational task of video understanding: text-video retrieval. However, unlike conventional transfer learning within the same domain, cross-modal transfer learning from images to videos often requires fine-tuning all pre-trained weights instead of keeping them frozen. This may result in overfitting and distorting the pre-trained weights, leading to a degradation in performance. To address this challenge, we introduce a learning strategy, termed Adaptive Parameter-wise Two-Stage Training with Integrated Scene Transition Mask Adapters (APT). Our two-stage learning process mitigates the distortion of the pre-trained weights. In the first stage, we use a novel method to identify an optimal subset of parameters to update by monitoring their fluctuations. Once the remaining parameters are frozen, the second stage is dedicated to learning temporal dynamics in videos with an adapter module. APT can be applied to any existing models in a plug-and-play manner and consistently achieves performance improvements over the base models. We report state-of-the-art performance across key text-video benchmark datasets, including MSRVTT and LSMDC. Our code is available at Our code is available at https://github.com/kang7734/APT.
키워드
- 제목
- APT : A Two-stage Image-Video Transfer Learning for Video Retrieval
- 저자
- Kang, Seong-Min; Cho, Yoon-Sik
- 발행일
- 2026
- 유형
- Article