상세 보기
유튜브 악플 탐지를 위한 기계학습: 스태킹 앙상블 모델의 적용을 중심으로
초록
This study examined the classification performance of machine learning models that automatically detect malicious comments, which easily spread on personalized media platforms such as YouTube, and applied a stacking ensemble model as a way to improve them. Though web-crawling, 59,999 comments were collected from popular videos of representative “cyber wrecker” channels that deliver social or entertainment issues involving certain celebrities and stimualte hatred of them. In doing so, we prepared a data set of 5,702 comments, of which 2,851 were labeled as malicious by human coders and 2,851 were randomly extracted from the remaining non-malicious comments to balance the data. As classification algorithms, we selected the logistic regression, Naïve Bayes, random forest, and support vector machine models. As a result, the performance of the models based on a single algorithm was clearly dependent on evaluation metrics. In particular, the random forest model, which showed the best performance on the basis of accuracy and score, fell behind other models in terms of recall. On the other hand, the stacking ensemble model improved the performance of classifying malicious and non-malicious comments by generating a meta-learning classifier that gained from single algorithm classifiers and complementing their weaknesses. Finally, we demonstrated how stacking ensemble models, as opposed to ensemble learning methods of bagging or boosting, can be used to create a better classifier of malicious comments by combining multiple algorithms with different advantages and addressing the imbalances shown in the classification evaluation metrics.
키워드
- 제목
- 유튜브 악플 탐지를 위한 기계학습: 스태킹 앙상블 모델의 적용을 중심으로
- 제목 (타언어)
- Machine Learning for Detecting Malicious Comments on YouTube: Focusing on the Application of Stacking Ensemble Model
- 저자
- 이신행
- 발행일
- 2022
- 권
- 24
- 호
- 4
- 페이지
- 1583 ~ 1598
- 언어
- KOR
- 출판사
- 한국자료분석학회
- 발행국가
- 대한민국
- 분량
- 16 페이지
- ISSN
- E 2733-9173
P 1229-2354