상세 보기
Area-wise relational knowledge distillation
- Cho Sungchul;
- Park Sangje;
- Lim Changwon
WEB OF SCIENCE
0SCOPUS
1초록
Knowledge distillation (KD) refers to extracting knowledge from a large and complex model (teacher) and transferring it to a relatively small model (student). This can be done by training the teacher model to obtain the activation function values of the hidden or the output layers and then retraining the student model using the same training data with the obtained values. Recently, relational KD (RKD) has been proposed to extract knowledge about relative differences in training data. This method improved the performance of the student model compared to conventional KDs. In this paper, we propose a new method for RKD by introducing a new loss function for RKD. The proposed loss function is defined using the area difference between the teacher model and the student model in a specific hidden layer, and it is shown that the model can be successfully compressed, and the generalization performance of the model can be improved. We demonstrate that the accuracy of the model applying the method proposed in the study of model compression of audio data is up to 1.8% higher than that of the existing method. For the study of model generalization, we demonstrate that the model has up to 0.5% better performance in accuracy when introducing the RKD method to self-KD using image data.
키워드
- 제목
- Area-wise relational knowledge distillation
- 저자
- Cho Sungchul; Park Sangje; Lim Changwon
- 발행일
- 2023-09
- 유형
- Article
- 권
- 30
- 호
- 5
- 페이지
- 501 ~ 516
- 언어
- ENG
- 출판사
- 한국통계학회
- 발행국가
- 대한민국
- 분량
- 16 페이지
- ISSN
- E 2383-4757
P 2287-7843