Lexical Distractor Mining Network for Causal Video Question Answering

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Causal Video Question Answering (Causal VideoQA) aims to answer questions about a video by reasoning over causal relationships among multiple events within the video. Unlike conventional VideoQA, which involves selecting a single correct answer from multiple choices, Causal VideoQA requires selecting both the answer and the causal reasoning that supports it. Despite significant progress, recent Causal VideoQA systems often struggle to identify accurate causal reasoning. This is largely due to the presence of lexical distractors in the reasoning candidates, which confuse models and lead them to incorrect inferences. A lexical distractor refers to an incorrect reasoning choice that mimics the correct reasoning sentence by exploiting the lexical and structural similarity of words, without reflecting the actual intended meaning. To this end, we propose the Lexical Distractor Mining Network (LDMNet), which performs sensible causal reasoning by progressively mining these distractors using our defined lexical similarity score and utilizing them as hard negatives for contrastive learning. Extensive experiments on two recent Causal VideoQA benchmarks (i.e., Causal-VidQA, NExT-QA) demonstrate that LDMNet achieves state-of-the-art performance, surpassing existing methods across multiple question types by reducing the model’s unnecessary reliance on lexical cues. The code will be made publicly available.

키워드

Causal Video Question AnsweringContrastive LearningHard Negative Sampling
제목
Lexical Distractor Mining Network for Causal Video Question Answering
저자
Kim, JunyeongYoon, Sunjae
DOI
10.1109/ACCESS.2026.3666880
발행일
2026-02
유형
Article
저널명
IEEE Access
14
페이지
29984 ~ 29994