arxivcs.LG2026-07-09
How are linear representations learned? Exact solutions to the dynamics of abstraction
William W. Yang, Andrew M. Saxe, Peter E. Latham
In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear p…