CORTEXA
← Browse

Tianyi Zhang

11 papers indexed

arxivcs.CVcs.AI2026-07-21

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, et al.

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology benchmarks evaluate models on pre-cropped patches or pre-extracted slide features,…

View free PDFSource page
arxivcs.RO2026-07-20

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, et al.

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared wi…

View free PDFSource page
arxivcs.ROcs.AI2026-07-08

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation

Shaopeng Zhai, Qi Zhang, Tianyi Zhang, Haoran Zhang, Fuxian Huang, Zhanhui Lin, et al.

When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post-training are often required to progressively address policy weaknesses. In this report, we focus on maximizing human efficiency during this iterative process, measured by policy improve…

View free PDFSource page
arxivcs.HC2026-07-07

Head, Gaze, or Finger? Comparing Object Selection Techniques in Augmented Reality for People with Low Vision

Ruijia Chen, Tianyi Zhang, Sanbrita Mondal, Yukang Yan, Yuhang Zhao

Augmented reality (AR) can enhance visual perception for people with low vision (PLV) by overlaying multimodal information. Selection-based augmentation further allows users to flexibly choose and augment relevant information while reducing distraction and visual clutter. However…

View free PDFSource page
arxivcs.RO2026-07-05

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

ACE-Brain Team, :, Ziyang Gong, Haoming Gu, Zehang Luo, Tianyi Zhang, et al.

Embodied AI is moving from isolated perception or action modules toward physical agents that understand, plan under goals, act through robot bodies, monitor progress, and improve from experience. Existing systems address this loop only in parts: end-to-end policies generate actio…

View free PDFSource page
arxivcs.CLcs.AI2026-07-02

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

Tianyi Zhang, Mousumi Das, Abrar Anwar, Jesse Thomason, David Traum

Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns require tailored strategies rather than a one-size-fits-all approach. To address this challenge, we focus on a fire-rescue scenario in which an ope…

View free PDFSource page
arxivcs.CVcs.AIcs.GRcs.RO2026-07-02

NeoMap: Training-free Novel-View Synthesis from Single Images and Videos

Jinxi Li, Tianyi Zhang, Yafei Yang, Zihui Zhang, Peng Huang, Koon Wing Macgyver Lin, et al.

We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate under the assumption that pre-trained video models lack native novel view synthesis capability and enforce view alignment via camera conditioning…

View free PDFSource page
arxivcs.CVcs.AI2026-06-29

LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion

Tianyi Zhang, Wei Shan, Yuan Zong, Tianhua Qi, Wenming Zheng

Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in modern recruitment. Existing approaches often rely on large language models (LLMs) to analyze textual responses of interviewees in AVI. However, u…

View free PDFSource page
arxivcs.CVcs.AI2026-06-28

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models

Xiaohao Xu, Feng Xue, Xiang Li, Haowei Li, Shusheng Yang, Tianyi Zhang, et al.

A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces. Monocular depth estimation, however, reduces this structure to one scalar depth per pixel. Transparent scenes make this…

View free PDFSource page