Loading...
Multi-Teacher Knowledge Distillation for Improving Student Model Generalization
Saedi, Behnam | 2025
24
Viewed
- Type of Document: M.Sc. Thesis
- Language: Farsi
- Document No: 58765 (19)
- University: Sharif University of Technology
- Department: Computer Engineering
- Advisor(s): Beigy, Hamid
- Abstract:
- In recent years, significant advances have been achieved in deep learning and novel architectures such as convolutional neural networks, vision transformers, and multilayer perceptrons. Despite the high performance of these large-scale models, deploying them in resource-constrained environments remains challenging. An effective approach to reducing model complexity while maintaining their capabilities is knowledge distillation. This study addresses the problem of cross-architecture multi-teacher knowledge distillation in image classification, which is a problem that has not yet been systematically explored. In this approach, the student models learn from multiple heterogeneous teachers, including convolutional neural networks, a vision transformer, and a multilayer perceptron, in order to benefit from the inherent advantages of each architecture. The proposed method seeks to overcome the main challenges of learning from multiple heterogeneous teachers by designing a comprehensive loss function that incorporates centralized kernel alignment and the sliced Gromov-Wasserstein distance, along with logit standardization. Furthermore, three weighting strategies consisting of fixed weights, learnable weighting, and sample-based adaptive weighting, were examined and compared, with results showing the superiority of dynamic weighting approaches over the fixed method. Experimental evaluation on the CIFAR-100 dataset demonstrates that the proposed method not only achieves competitive performance in the single-teacher setting but also outperforms previous methods in multi-teacher scenarios. Specifically, the proposed method was able to achieve an accuracy of 76.82% by training the student model using three teachers with heterogeneous architectures, which is an improvement of 1.06% over the previous best method. Thus, this research makes a significant step toward developing an architecture-agnostic framework for multi-teacher knowledge distillation, capable of meaningfully improving the accuracy and efficiency cross-architecture knowledge distillation in image classification tasks
- Keywords:
- Knowledge Distillation ; Teacher-Student Architecture ; Cross-Architecture Training ; Multi-Teacher Method ; Improving Generalization
-
محتواي کتاب
- view
