Data augmentation for dense passage retrieval using corpus-passage frequency-based token deletion

Citations

WEB OF SCIENCE

5
Citations

SCOPUS

5

초록

This paper proposes a novel data augmentation method to address class imbalance in large-scale information retrieval systems. In particular, a corpus-passage frequency-based token deletion technique is introduced to improve the accuracy of Dense Passage Retrieval, which is a dense vector-based information retrieval model. Unlike traditional random token deletion methods that delete tokens with equal probability, the proposed method calculates token importance by considering both passage and corpus-level frequencies, leading to more effective token deletion. Experimental results demonstrate that the proposed approach significantly improves Top-k accuracy on smaller datasets compared to conventional augmentation techniques. While maintaining competitive performance on larger-scale datasets, its relative effectiveness is particularly notable in scenarios characterized by limited training data and severe class imbalance. This confirms its potential to improve the generalizability of information retrieval models. The source code is publicly available at https://github.com/asmoon002/DPR_TD.

키워드

Information retrieval; Data augmentation; Natural language processing; Class imbalance
제목
Data augmentation for dense passage retrieval using corpus-passage frequency-based token deletion
저자
Moon, A-Seong; Kim, Kyumin; Lee, Jaesung
DOI
10.1186/s40537-025-01257-9
발행일
2025-08
유형
Article
저널명
JOURNAL OF BIG DATA
권
12
호
1

파일 다운로드

Thumbnail