arxivcs.LGmath.STstat.ML2026-07-09
Stochastic Linear Bandits with Partially Observed Actions
Gautam Dasarathy, Vineet Gattani, Lalit Jain
The stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making. We study a partially observed variant of this problem in which the learning agent only sees a random subset of coordinates for each…