Product failure prediction with missing data

  • Kang, Seokho; 
  • Kim, Eunji; 
  • Shim, Jaewoong; 
  • Chang, Wonsang; 
  • Cho, Sungzoon
Citations

WEB OF SCIENCE

23
Citations

SCOPUS

26

초록

In production data, missing values commonly appear for several reasons including changes in measurement and inspection items, sampling inspections, and unexpected process events. When applied to product failure prediction, the incompleteness of data should be properly addressed to avoid performance degradation in prediction models. Well-known approaches for missing data treatment, such as elimination and imputation, would not perform well under usual scenarios in production data, including high missing rate, systematic missing and class imbalance. To address these limitations, here we present a method for predictive modelling with missing data by considering the characteristics of production data. It builds multiple prediction models on different complete data subsets derived from the original data-set, each of which has different coverage of instances and input variables. These models are selectively used to make predictions for new instances with missing values. We demonstrate the effectiveness of the proposed method through a case study using actual data-sets from a home appliance manufacturer.

키워드

data mining; predictive modelling; failure prediction; production data; missing value; NEURAL-NETWORKS; FAULT-DETECTION; DATA IMPUTATION; ROC CURVE; VALUES; QUALITY; AREA; MAP
제목
Product failure prediction with missing data
저자
Kang, Seokho; Kim, Eunji; Shim, Jaewoong; Chang, Wonsang; Cho, Sungzoon
DOI
10.1080/00207543.2017.1407883
발행일
2018
유형
Article
저널명
International Journal of Production Research
권
56
호
14
페이지
4849 ~ 4859