arxivcs.LG2026-07-15
Optimal and Efficient Contextual Combinatorial Semi-bandits with General Function Approximation
We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and receives the reward of each selected arm; the goal…