arxivcs.LGcs.AI2026-06-26
A Gravitational Interpretation of Fine-Tuning Reversion
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge, latent traits can transfer through apparently unrelated supervision, and related post-alignment frag…