arxivcs.DCcs.AIcs.LG2026-06-26
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
Adrian P. Dieguez, Victor Conchello Vendrell, Alex Batlle, Vinnam Kim, Jordi Ros-Giralt, Harris Teague
Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL library implements it. Yet, TRL treats both models symmetrically, missing opportunities to exploit their pronounced asymmetry in memory foot…