arxivcs.LGcs.AI2026-06-28
On the Nonlinearity of Learning Rate Scaling for LLM Training
Zaiwen Yang, Huaqing Zhang, Jing Xu, Jingzhao Zhang
Learning-rate transfer can reduce the cost of training large language models: instead of sweeping learning rates at target scale, practitioners extrapolate from smaller runs. Existing approaches often assume that the optimal learning rate follows a log-linear scaling law in data…