Cost-constrained Group Feature Selection Using Information Theory

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

A problem of cost-constrained group feature selection in supervised classification is considered. In this setting, the features are grouped and each group is assigned a cost. The goal is to select a subset of features that does not exceed a user-specified budget and simultaneously allows accurate prediction of the class variable. We propose two sequential forward selection algorithms based on the information-theoretic framework. In the first method, a single feature is added in each step, whereas in the second one, we select the entire group of features in each step. The choice of the candidate feature or group of features is based on the novel score function that takes into account both the informativeness of the added features in the context of previously selected ones as well as the cost of the candidate group. The score is based on the lower bound of the mutual information and thus can be effectively computed even when the conditioning set is large. The experiments were performed on a large clinical database containing groups of features corresponding to various diagnostic tests and administrative data. The results indicate that the proposed method allows achieving higher accuracy than the traditional feature selection method, especially when the budget is low. © 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.

키워드

costly features; group feature selection; information theory; mutual information
제목
Cost-constrained Group Feature Selection Using Information Theory
저자
Klonecki, Tomasz; Teisseyre, Paweł; Lee, Jaesung
DOI
10.1007/978-3-031-33498-6_8
발행일
2023-06
유형
Proceedings Paper
저널명
Lecture Notes in Computer Science
권
13890 LNCS
페이지
121 ~ 132