arxivcs.CLcs.AI2026-07-09
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
Xinlong Zhao, Dongsheng Liu, Hengyu Zhao, Zixuan Fu, Zheng Wang, Jie Cai, et al.
As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Language Models (LLMs) now depends less on data expansion and more on higher-quality data utilization. However, in the context of large-scale co…