arxivcs.LGcs.AIcs.CL2026-07-02
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safety constraints must e…