Loading...

Multi-Teacher Knowledge Distillation for Improving Student Model Generalization

Saedi, Behnam | 2025

24 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 58765 (19)
  4. University: Sharif University of Technology
  5. Department: Computer Engineering
  6. Advisor(s): Beigy, Hamid
  7. Abstract:
  8. In recent years, significant advances have been achieved in deep learning and novel architectures such as convolutional neural networks, vision transformers, and multilayer perceptrons. Despite the high performance of these large-scale models, deploying them in resource-constrained environments remains challenging. An effective approach to reducing model complexity while maintaining their capabilities is knowledge distillation. This study addresses the problem of cross-architecture multi-teacher knowledge distillation in image classification, which is a problem that has not yet been systematically explored. In this approach, the student models learn from multiple heterogeneous teachers, including convolutional neural networks, a vision transformer, and a multilayer perceptron, in order to benefit from the inherent advantages of each architecture. The proposed method seeks to overcome the main challenges of learning from multiple heterogeneous teachers by designing a comprehensive loss function that incorporates centralized kernel alignment and the sliced Gromov-Wasserstein distance, along with logit standardization. Furthermore, three weighting strategies consisting of fixed weights, learnable weighting, and sample-based adaptive weighting, were examined and compared, with results showing the superiority of dynamic weighting approaches over the fixed method. Experimental evaluation on the CIFAR-100 dataset demonstrates that the proposed method not only achieves competitive performance in the single-teacher setting but also outperforms previous methods in multi-teacher scenarios. Specifically, the proposed method was able to achieve an accuracy of 76.82% by training the student model using three teachers with heterogeneous architectures, which is an improvement of 1.06% over the previous best method. Thus, this research makes a significant step toward developing an architecture-agnostic framework for multi-teacher knowledge distillation, capable of meaningfully improving the accuracy and efficiency cross-architecture knowledge distillation in image classification tasks
  9. Keywords:
  10. Knowledge Distillation ; Teacher-Student Architecture ; Cross-Architecture Training ; Multi-Teacher Method ; Improving Generalization

 Digital Object List

 Bookmark

...see more