arxivcs.LGcs.AI2026-07-03
Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning
Eric Lei, Hsiang Hsu, Chun-Fu Chen
Inference-time alignment methods, such as Best-of-$N$, offer a flexible alternative to training-based alignment by using reward models to select high-quality responses generated by a reference LLM. However, the efficacy of these methods is inherently limited by the response quali…