CORTEXA
← Browse

Tat-Seng Chua

2 papers indexed

arxivcs.CV2026-06-30

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

Chih-Ting Liao, Fei Shen, Xin Cao, Tat-Seng Chua

The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-language model (VLM) actually grounds in the image. We show this on spatial reasoning, where the error is invisible to both probin…

View free PDFSource page
arxivcs.CVcs.CLcs.LG2026-06-25

DanceOPD: On-Policy Generative Field Distillation

Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, et al.

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance,…

View free PDFSource page