arxivstat.MLcs.LG2026-07-18
Backpropagation-Free Trunk Training via the Split Forward Gradients
Backpropagation makes training deep networks memory intensive because it must store intermediate activations. Forward-mode methods avoid this cost, but their gradient estimates become increasingly noisy as the number of trained parameters grows. We introduce Split Forward Gradien…