CORTEXA
← Browse

Subha Raut

2 papers indexed

arxivcs.LG2026-07-21

Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention

Subham Singh, Ashutosh Mishra, Subha Raut

Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not. This is a robust empirical regularity, and the explanations offered for it so far are chie…

View free PDFSource page
arxivcs.LG2026-07-14

Same Loss, Same Noise, Opposite Schedules: Noise Structure and Optimizer Normalization Jointly Determine Whether Learning-Rate Cooldown Helps

Subham Singh, Ashutosh Mishra, Subha Raut

The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others. We give a provable account of which case obtains, and it turns on two properties togethe…

View free PDFSource page