Area-wise relational knowledge distillation

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Knowledge distillation (KD) refers to extracting knowledge from a large and complex model (teacher) and transferring it to a relatively small model (student). This can be done by training the teacher model to obtain the activation function values of the hidden or the output layers and then retraining the student model using the same training data with the obtained values. Recently, relational KD (RKD) has been proposed to extract knowledge about relative differences in training data. This method improved the performance of the student model compared to conventional KDs. In this paper, we propose a new method for RKD by introducing a new loss function for RKD. The proposed loss function is defined using the area difference between the teacher model and the student model in a specific hidden layer, and it is shown that the model can be successfully compressed, and the generalization performance of the model can be improved. We demonstrate that the accuracy of the model applying the method proposed in the study of model compression of audio data is up to 1.8% higher than that of the existing method. For the study of model generalization, we demonstrate that the model has up to 0.5% better performance in accuracy when introducing the RKD method to self-KD using image data.

키워드

deep learning; model compression; knowledge distillation; audio signal processing; image processing
제목
Area-wise relational knowledge distillation
저자
Cho Sungchul; Park Sangje; Lim Changwon
DOI
10.29220/CSAM.2023.30.5.501
발행일
2023-09
유형
Article
저널명
Communications for Statistical Applications and Methods
권
30
호
5
페이지
501 ~ 516