The correlation geometry of learned representations: a mean-field baseline with a weak anomalous excess, and its cross-domain universality
The hidden states of large language models carry a power-law correlation along the token axis, C(r) ∝ r^{−α}, whose exponent is stable across model scale and across architectures that share little else. We ask what this apparent “attraction” between nearby positions in representation space actually is. We argue that it is neither a mechanism imposed from outside nor an artefact of a particular training choice, but a near-inevitable consequence of representing structured data compactly — a compression signature. We read the measured exponent as a sum, α = 1/2 + Δ: a mean-field, structureless baseline of one half, plus a weak excess Δ that co-varies with higher-order structure in the data. We then observe that this same “mean-field baseline plus weak anomalous correction” shape is not particular to neural networks. It recurs, with anomalous pieces that others have measured to high precision, in critical phenomena — the anomalous dimension η — and in fully developed turbulence — the intermittency correction δζ₂ — where the anomalous term is universally small, a few percent. We give a sharp criterion separating this structure from a trivial “leading term plus correction”, state carefully what the convergence does and does not establish, and set out the derivation that would be needed to connect the empirical decomposition to an effective-field-theory description of lossy compression.