arxivcs.LG2026-07-08
Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts
Andrei Cristian Popescu, Haitz Sáez de Ocáriz Borde, Pietro Liò
Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block. Learned halting objectives in looped Transformers typically use a single exit distribution both as the inference-time stopping rule and as the training-time weighting of per-depth…