arxivcs.CV2026-07-01
Condensing Large-Scale Datasets Directly with Minimal Information Loss
Xinyi Shang, Peng Sun, Bei Shi, Zixuan Wang, Tao Lin
Recent advancements in scaling dataset distillation rely heavily on decoupled information extraction pipelines, comprising SQUEEZE, RECOVER, and RELABEL stages. Despite their scalability to large-scale datasets, these methods suffer from prohibitive computational overhead and poo…