arxivcs.LGcs.AIstat.ML2026-06-28
Reducing Per-Sample Harm in Stochastic Optimization
Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. While highly effective, aggregating across the batch and incorporating this history can produce parameter updates that increase the loss of in…