CORTEXA
← Browse
arxivcs.LG2026-07-06

Non-Convex Sparse Reinforcement Learning via Non-Monotone Inclusions

Kyohei Suzuki, Konstantinos Slavakis

This work delivers two key contributions: one to efficient feature selection in reinforcement learning (RL), the other to the theory of non-monotone inclusions. On the RL side, the estimation bias inherent in conventional regularization schemes is addressed by augmenting classical least-squares temporal-difference (LSTD) policy evaluation with the sparsity-inducing, non-convex projected minimax concave (PMC) penalty. Because the PMC penalty is weakly convex, the resulting fixed-point problem is no longer monotone; instead, it falls under a broader class of non-monotone inclusions involving the sum of a monotone Lipschitz operator and a hypomonotone operator. On the theory side, novel convergence conditions are developed for the forward-reflected-backward splitting (FRBS) method applied to this broader class of non-monotone inclusion problems. Under mild conditions, Lyapunov stability and the existence of a limit point of the sequence of FRBS iterates are established; alternatively, under the weak Minty variational inequality assumption, exact convergence is guaranteed. Numerical tests on benchmark datasets show that the proposed FRBS iterates, applied to the non-convexly regularized LSTD problem, substantially outperform state-of-the-art feature-selection methods, especially when many noisy features are present.

View free PDFSource page

Related papers

arxivcs.LGcs.AI2026-07-14

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

Yassine Chemingui, Chenhua Fan, Honghao Wei, Janardhan Rao Doppa

Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate, a boundary-aware distributional safety certificat…

View free PDFSource page
arxivstat.MLcs.LGeess.SP2026-07-12

Demixing Sparse Signals from Nonlinear Observations using Generalized Non-convex Regularization

Raziyeh Takbiri

We consider the recovery of a pair of sparse vectors from a limited number of nonlinear observations of their superposition: $y_i=g(\inner{\ba_i}{\bPhi\bw^\ast+\bPsi\bz^\ast})+e_i$, $i=1,\dots,m$, with $m\ll n$, incoherent orthonormal bases $\bPhi,\bPsi$, a scalar link $g$, and n…

View free PDFSource page
arxivcs.LGcs.CRstat.ML2026-07-01

Unveiling the Non-Monotonic Effect of Privacy on Generalization under Byzantine Robustness

Thomas Boudou, Batiste Le Bars, Nirupam Gupta, Aurélien Bellet

Recent work has established a fundamental trilemma between Byzantine robustness, local differential privacy (LDP), and optimization error in distributed learning. We show that this trilemma does not universally extend to generalization error, but instead depends critically on the…

View free PDFSource page
arxivcs.AIcs.LG2026-07-13

Learning Residual Kinematic Corrections for Continuous Neural Decoding via Reinforcement Learning

Jiamian Li, Niall McShane, Attila Korik, Naomi du Bois, Karl McCreadie, Leen Jabban, et al.

Decoding continuous three-dimensional (3D) motor imagery (MI) using non-invasive electroencephalography (EEG)-based brain--computer interfaces (BCIs) remains challenging due to signal variability and residual decoding errors. Deep learning architectures such as convolutional neur…

View free PDFSource page
arxivquant-phcs.LG2026-06-29

Staged Hybridisation for Visual Quantum Reinforcement Learning via Knowledge Distillation

Javier Lazaro, Juan-Ignacio Vazquez, Pablo Garcia-Bringas

Visual environments are a demanding setting for quantum reinforcement learning (QRL): high-dimensional observations, unstable RL optimisation, and constrained variational quantum circuits (VQCs) are difficult to train jointly. This paper studies knowledge distillation (KD) as a s…

View free PDFSource page