Data synthesis method preserving correlation of features

Citations

WEB OF SCIENCE

11
Citations

SCOPUS

11

초록

Abundant data are essential for improving the performance of machine learning algorithms. Thus, if only limited data are available, data synthesis can be used to enlarge datasets. Data synthesis methods based on the covariance matrix are useful because of their fast data synthesis capabilities. However, artificial datasets generated via classical techniques show statistical discrepancies when compared to original datasets. To address this problem, we developed a new data synthesis method that preserves the correlation (between features) observed in the original dataset. This preservation was realized by considering not only the correlation but also the random noises used in data synthesis process. This method was applied to various biosignals (i.e., electrocortiography, electromyogram, and electrocardiogram), wherein data points are insufficient. Several classifiers (i.e., convolutional neural network, support vector machine, and k-nearest neighbor) were used to verify that the classification accuracy can be improved by the proposed data synthesis method. © 2021 Elsevier Ltd

키워드

Artificial dataset; Correlation; Data synthesis; Random noise; Classification (of information); Correlation methods; Learning algorithms; Nearest neighbor search; Neural networks; Support vector machines; Artificial datasets; Classical techniques; Correlation; Covariance matrices; Data synthesis; Limited data; Machine learning algorithms; Performance; Random noise; Synthesis method; Covariance matrix
제목
Data synthesis method preserving correlation of features
저자
Yang, W.; Nam, W.
DOI
10.1016/j.patcog.2021.108241
발행일
2022-02
유형
Article
저널명
Pattern Recognition
권
122