arxivcs.CLcs.LG2026-07-10
Complexity-Guided Component-wise Initialization for Language Model Pretraining
Konstantin Garbers, Nicholas Oh
Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-style langua…