VolDoGER: LLM-Assisted Datasets for Domain Generalization in Vision-Language Tasks

  • Choi, Juhwan; 
  • Kwon, Junehyoung; 
  • Yun, Jungmin; 
  • Yu, Seunguk; 
  • Kim, YoungBin
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Domain generalizability is a crucial aspect of a deep learning model since it determines the capability of the model to perform well on data from unseen domains. How-ever, research on the domain generalizability of deep learning models for vision-language tasks remains limited, primarily because of the lack of required datasets. To address these challenges, we propose VolDoGER: Vision-Language Dataset for Domain Generalization, a dedicated dataset designed for domain generalization that addresses three vision-language tasks: image captioning, visual question answering, and visual entailment. We constructed VolDoGER by extending LLM-based data annotation techniques to vision-language tasks, thereby alleviating the burden of recruiting human annotators. We evaluated the domain generalizability of various models through VolDoGER.

키워드

Domain Generalization; LLM for Data Annotation; Synthetic Data
제목
VolDoGER: LLM-Assisted Datasets for Domain Generalization in Vision-Language Tasks
저자
Choi, Juhwan; Kwon, Junehyoung; Yun, Jungmin; Yu, Seunguk; Kim, YoungBin
DOI
10.1109/ICCVW69036.2025.00712
발행일
2025
유형
Proceedings Paper
저널명
Proceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
페이지
6892 ~ 6902