CORTEXA
← Browse

Zhi Wang

7 papers indexed

arxivcs.AI2026-07-04

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

Zican Hu, Xuyang Hu, Yiming Liu, Zuwei Long, Wei Liu, Yunzhuo Hao, et al.

Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Existing approaches apply RL exclusively to text steps, relegatin…

View free PDFSource page
arxivcs.RO2026-07-04

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models

Xinyi Xie, Zican Hu, Zhanyu Liu, Yicheng Dong, Wenhao Wu, Zhenhong Sun, et al.

Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLMs. The prevailing remedy, post-training via supervised fine-tuning or reinforcement learning, improv…

View free PDFSource page
arxivcs.AI2026-07-02

Atomic Task Graph: A Unified Framework for Agentic Planning and Execution

Yue Zhang, Sihan Chen, Ziwen Huang, Hanyun Cui, Kangye Ji, Zhi Wang

LLM-based agents have shown strong potential for solving complex multi-step tasks, yet existing performance improvements often rely on either scaling to larger backbone models or task-specific fine-tuning. The former incurs substantial computational costs, while the latter typica…

View free PDFSource page
crossrefEngineering Reports2026-07-01

Comparative Evaluation of Statistical, Machine‐Learning, and Deep‐Learning Models for Construction Sand and Gravel Price Forecasting: A Synthetic‐Data, Simulation‐Based Benchmark

You Wu, Peng Li, Jing Zhang, Fengsheng Guo, Danshu Hu, Wei He, et al.

ABSTRACT Forecasting construction sand and gravel prices is critical for infrastructure cost control, yet reliable comparisons among model families in the small‐sample, multidriver setting typical of regional markets are lacking. This study benchmarks eight algorithms—Ridge regre…

View free PDFSource page
arxivcs.AIcs.CLcs.LO2026-06-30

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization

Ke Zhang, Patricio Gallardo Candela, Sudhir Murthy, Yi Xie, Zhi Wang, Maziar Raissi

Theorem-proving benchmarks evaluate proof search against fixed formal statements, but natural-language-to-Lean formalization must generate the formal statement itself. In this setting, compilation is only a validity check: a Lean declaration may type-check while omitting hypothes…

View free PDFSource page
arxivcs.AIcs.CV2026-06-25

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang, Zhi Wang, Hailan Ma, et al.

As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where…

View free PDFSource page
crossrefPhotonics2021-03-15Cited by 10

Optical Machine Learning Using Time-Lens Deep Neural NetWorks

Luhe Zhang, Caiyun Li, Jiangyong He, Yange Liu, Jian Zhao, Huiyi Guo, et al.

As a high-throughput data analysis technique, photon time stretching (PTS) is widely used in the monitoring of rare events such as cancer cells, rough waves, and the study of electronic and optical transient dynamics. The PTS technology relies on high-speed data collection, and t…

View free PDFSource page