RetroKD: Leveraging past states for regularizing targets in teacher-student learning

Surgan Jandial, Yash Khasbage, Arghya Pal, Balaji Krishnamurthy, Vineeth N. Balasubramanian

Research output: Chapter in Book/Report/Conference proceedingConference PaperResearch


Several recent works show that higher accuracy models may not be better teachers for every student, and hence, refer this problem as student-teacher "knowledge gap". Further, they propose techniques, which, in this paper, we discuss are constrained to certain pre-conditions: 1). Access to Teacher Model/Architecture 2). Retraining Teacher Model 3). Models in Addition to Teacher Model. Being well known that for a lot of settings, these conditions may not hold true challenges the applicability of such approaches. In this work, we propose RetroKD, which smoothes out the logits of a student network by leveraging students' past state logits with the ones from the teacher. By doing so, we hypothesize that the present target will no longer be as hard as the teacher target and not as more uncomplicated as the past student target. Such regularization on learning the parameters alleviates the needs as required by other methods. Our extensive set of experiments comparing against the baselines for CIFAR 10, CIFAR 100, and TinyImageNet datasets and a theoretical study further help in supporting our claim. We performed crucial ablation studies such as hyperparameter sensitivity, the generalization study by showing the flatness on loss landscape and feature similarly with teacher network.

Original languageEnglish
Title of host publication2023 Proceedings of the 6th Joint International Conference on Data Science & Management of Data (10th ACM IKDD CODS and 28th COMAD)
EditorsAbhinandan S P, Charu Sharma
Place of PublicationNew York NY USA
PublisherAssociation for Computing Machinery (ACM)
Number of pages9
ISBN (Electronic)9781450397988
Publication statusPublished - 2023
Externally publishedYes
EventACM India Joint International Conference on Data Science and Management of Data 2023 - Mumbai, India
Duration: 4 Jan 20237 Jan 2023
Conference number: 6th (Proceedings) (Website)


ConferenceACM India Joint International Conference on Data Science and Management of Data 2023
Abbreviated titleCODS-COMAD 2023
Internet address


  • Knowledge Distillation
  • Past States
  • Regularization

Cite this