arxivcs.CV2026-06-29
H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning
Eric Peh, Debaditya Roy, Basura Fernando
Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining "black boxes", yet they remain prone to hallucination or rely on superficial shortcuts. In this work, we propose a framework designed to enhance both performance and interpretability throug…