Memetic feature selection for multilabel text categorization using label frequency difference

Citations

WEB OF SCIENCE

55
Citations

SCOPUS

72

초록

Multilabel text categorization is an important task in modern text mining applications. Text datasets comprise an excessive number of terms, and this can degrade the accuracy. Therefore, conventional studies applied a feature selection method before text categorization. Recently, memetic feature selection methods that hybridize an evolutionary feature wrapper and a filter have gained popularity and showed promising results. However, conventional memetic text feature selection methods suffer from limited performance because the used feature filter requires problem transformation that degrades the search capability, resulting in unrefined feature subsets with poor accuracy. In this study, we propose an effective memetic feature selection method based on a novel feature filter that is highly specialized to multilabel text categorization. Our experiments demonstrate that the proposed method significantly outperforms several conventional methods. © 2019 Elsevier Inc.

키워드

Multi-label text categorization; Feature selection; Memetic search; Population-based incremental learning; NAIVE BAYES; CLASSIFICATION; ALGORITHM; SCHEME
제목
Memetic feature selection for multilabel text categorization using label frequency difference
저자
Lee, Jaesung; Yu, Injun; Park, Jaegyun; Kim, Dae-Won
DOI
10.1016/j.ins.2019.02.021
발행일
2019-06
유형
Article
저널명
Information Sciences
권
485
페이지
263 ~ 280