상세 보기
VolDoGER: LLM-Assisted Datasets for Domain Generalization in Vision-Language Tasks
- Choi, Juhwan;
- Kwon, Junehyoung;
- Yun, Jungmin;
- Yu, Seunguk;
- Kim, YoungBin
WEB OF SCIENCE
0SCOPUS
0초록
Domain generalizability is a crucial aspect of a deep learning model since it determines the capability of the model to perform well on data from unseen domains. How-ever, research on the domain generalizability of deep learning models for vision-language tasks remains limited, primarily because of the lack of required datasets. To address these challenges, we propose VolDoGER: Vision-Language Dataset for Domain Generalization, a dedicated dataset designed for domain generalization that addresses three vision-language tasks: image captioning, visual question answering, and visual entailment. We constructed VolDoGER by extending LLM-based data annotation techniques to vision-language tasks, thereby alleviating the burden of recruiting human annotators. We evaluated the domain generalizability of various models through VolDoGER.
키워드
- 제목
- VolDoGER: LLM-Assisted Datasets for Domain Generalization in Vision-Language Tasks
- 저자
- Choi, Juhwan; Kwon, Junehyoung; Yun, Jungmin; Yu, Seunguk; Kim, YoungBin
- 발행일
- 2025
- 유형
- Proceedings Paper
- 저널명
- Proceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
- 페이지
- 6892 ~ 6902
- 언어
- ENG
- 출판사
- Institute of Electrical and Electronics Engineers Inc.
- 분량
- 11 페이지
- ISSN
- P 2473-9936