arxivstat.MLcs.LG2026-07-31
The Greedy Advantage in Finite-Horizon Bandits
Kai Zhou, Michael Lingzhi Li, Kai Wang
Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algorithms with strong asymptotic regret guarantees, many practical applications operate over finite and externally imposed horizons. Mo…