Apparatus and method for video representation learning

초록

An apparatus for video representation learning according to an embodiment may extract video features from video data to generate a video embedding, extract image features from image data extracted from the video data to generate an image embedding, and extract audio features from audio data extracted from the video data to generate an audio embedding. Further, contrastive learning may be performed by generating a first compositional embedding based on the video embedding and the audio embedding, generating a second compositional embedding based on the video embedding and the audio embedding, generating a positive sample and a negative sample based on a correlation between the image embedding and the audio embedding, and then using the data.

제목
Apparatus and method for video representation learning
저자
Jong Won Choi; Soo Hyun Park; Jong Su Youn
발행일
2024-07-02
출원번호
18/396142
등록번호
12026936
출원일
2023-12-26
등록일
2024-07-02