상세 보기
Detection of Videos with Audio-Visual Inconsistency for Video Representation Learning
- Park, Soohyun;
- Lim, Hyoungjun;
- Choi, Jongwon
WEB OF SCIENCE
0SCOPUS
0초록
Audio-visual alignment using video data is a conventional approach for the self-supervision of multi-modal representation learning. Nevertheless, the presence of background music, external noise, and human conversational audio might lead to misalignment between the audio and visual elements in videos. In this paper, we introduce a method that concurrently identifies erroneous videos and trains the multi-modal representation model. Throughout the training process of the multi-modal representation model, we enhance a module responsible for detecting misaligned audio-visual videos, establishing accurate audiovisual pairs by eliminating erroneous videos. The misaligned audio-visual video detection module features an architecture based on VQ-VAE, extending to consider label information from input videos, effectively reconstructing video-based features along with labels. We evaluate our method on the tasks of video recognition and video retrieval on UCF-51 and UCF-101 datasets, achieving competitive performance with existing representation learning methods for audio-visual knowledge transfer.
- 제목
- Detection of Videos with Audio-Visual Inconsistency for Video Representation Learning
- 저자
- Park, Soohyun; Lim, Hyoungjun; Choi, Jongwon
- 발행일
- 2025
- 유형
- Proceedings Paper
- 저널명
- Proceedings - IEEE International Conference on Advanced Video and Signal-Based Surveillance, AVSS
- 언어
- ENG
- 출판사
- Institute of Electrical and Electronics Engineers Inc.
- ISSN
- E 2643-6213
P 2643-6213