arxivcs.AIcs.CLcs.LG2026-07-08
Length Penalties Make Chain-of-Thought Less Monitorable
Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even though the models' chains of thoug…