Detection of Videos with Audio-Visual Inconsistency for Video Representation Learning

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Audio-visual alignment using video data is a conventional approach for the self-supervision of multi-modal representation learning. Nevertheless, the presence of background music, external noise, and human conversational audio might lead to misalignment between the audio and visual elements in videos. In this paper, we introduce a method that concurrently identifies erroneous videos and trains the multi-modal representation model. Throughout the training process of the multi-modal representation model, we enhance a module responsible for detecting misaligned audio-visual videos, establishing accurate audiovisual pairs by eliminating erroneous videos. The misaligned audio-visual video detection module features an architecture based on VQ-VAE, extending to consider label information from input videos, effectively reconstructing video-based features along with labels. We evaluate our method on the tasks of video recognition and video retrieval on UCF-51 and UCF-101 datasets, achieving competitive performance with existing representation learning methods for audio-visual knowledge transfer.

제목
Detection of Videos with Audio-Visual Inconsistency for Video Representation Learning
저자
Park, Soohyun; Lim, Hyoungjun; Choi, Jongwon
DOI
10.1109/AVSS65446.2025.11149931
발행일
2025
유형
Proceedings Paper
저널명
Proceedings - IEEE International Conference on Advanced Video and Signal-Based Surveillance, AVSS