ETA: Enriching Typos Automatically from Real-World Corpora for Few-Shot Learning

Citations

SCOPUS

0

초록

Spell checking is the task of rectifying errors in a sentence resulting from various factors, and despite continuous research in this field, research often focused on widely known specific languages, such as English and Chinese. In this study, we focus on the Korean language and its linguistic features, particularly the propensity for a single character can be incorrect in diverse ways. Therefore, we categorize spelling errors from real-world corpora and automatically construct an error corpus based on their statistical patterns. When we employed them to leverage the impact of a pretrained large language model (LLM), we confirmed that utilizing the introduced spelling errors as samples for few-shot learning can be helpful in error correction tasks. We hope that this study contributes to the automatic construction of error corpora and prompt-based approaches for providing a better quality of consumer service, especially provoke research related to other low-resource languages in spell checking tasks. © 2024 IEEE.

키워드

Deep Learning; Natural Language Processing; Spell Checking Task
제목
ETA: Enriching Typos Automatically from Real-World Corpora for Few-Shot Learning
저자
Yu, Seunguk; Kim, Yeonghwa; Jin, Kyohoon; Kim, Youngbin
DOI
10.1109/ICCE-Asia63397.2024.10773956
발행일
2024
유형
Conference paper
저널명
2024 IEEE International Conference on Consumer Electronics-Asia, ICCE-Asia 2024