As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerged as a key research priority. However, existing…
Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward…
The yield assessment process during maize harvesting is a necessary means to ensure farmers’ economic benefits and stable agricultural production. Predicting the mass of maize kernels is an important condition for yield detection. This study proposes a maize kernel mass predictio…