Generative Data Augmentation via Wasserstein Autoencoder for Text Classification

Citations

SCOPUS

0

초록

Generative latent variable models are commonly used in text generation and augmentation. However generative latent variable models such as the variational autoencoder(VAE) experience a posterior collapse problem ignoring learning for a subset of latent variables during training. In particular, this phenomenon frequently occurs when the VAE is applied to natural language processing, which may degrade the reconstruction performance. In this paper, we propose a data augmentation method based on the pre-trained language model (PLM) using the Wasserstein autoencoder (WAE) structure. The WAE was used to prevent a posterior collapse in the generative model, and the PLM was placed in the encoder and decoder to improve the augmentation performance. We evaluated the proposed method on seven benchmark datasets and proved the augmentation effect. © 2022 IEEE.

키워드

Generative model; Text augmentation; Text classification
제목
Generative Data Augmentation via Wasserstein Autoencoder for Text Classification
저자
Jin, K.; Lee, J.; Choi, J.; Jang, S.; Kim, Youngbin
DOI
10.1109/ICTC55196.2022.9952762
발행일
2022-10
유형
Conference Paper
저널명
International Conference on ICT Convergence
권
2022-October
페이지
603 ~ 607