arxivcs.CV2026-07-22
Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning
Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible. We introduc…