Knowledge Distillation Techniques in Deep Learning

Authors

  • Clara Garcia Author
  • Helena Costa Author
  • Erik Schmidt Author

DOI:

https://doi.org/10.5281/zenodo.19614611

Keywords:

knowledge distillation; model compression; teacher-student learning; soft labels; feature alignment; self-distillation; data-free distillation; dark knowledge

Abstract

Knowledge distillation transfers learned representations from a large teacher model to a compact student model, enabling deployment of high-quality inference on resource-constrained hardware without the computational cost of running the teacher directly. This study presents a controlled evaluation of six distillation strategies -- response-based (soft label KD), feature-based (FitNets), attention-based (AT), relational (CRD), self-distillation, and data-free distillation - across four tasks: image classification (CIFAR-100, ImageNet), object detection (COCO), NLP classification (GLUE), and speech recognition (LibriSpeech). Teacher-student pairs spanned compression ratios from 2x to 50x (ResNet-110 to ResNet-20; BERT-base to TinyBERT-4L; Wav2Vec2-Large to Wav2Vec2-Small). A total of 2,040 experiments were conducted under standardised training protocols. Response-based KD recovered a mean of 68.4 +- 2.8% of the teacher-student accuracy gap across all tasks -- a strong baseline that more complex methods improved upon only marginally. Feature-based distillation (FitNets) achieved the best recovery on image classification (74.2 +- 2.4%) by aligning intermediate representations. Contrastive relational distillation (CRD) excelled on NLP tasks (76.8 +- 2.2% recovery) by preserving inter-sample relationships. Self-distillation improved teacher models themselves by 1.2-2.4% without any student, confirming that distillation captures regularisation benefits beyond compression. Data-free distillation achieved 82.4% of data-available distillation quality without access to training data, enabling distillation when the original training dataset is proprietary. The teacher-student capacity ratio was the strongest predictor of distillation success (partial R2 = 0.54): students smaller than 10% of teacher capacity showed sharp performance cliffs. A practical distillation pipeline selection guide is proposed.

Downloads

Published

2026-08-19

How to Cite

Knowledge Distillation Techniques in Deep Learning. (2026). Bio-QI  Journal, 1(4), 174-182. https://doi.org/10.5281/zenodo.19614611