CORTEXA
← Browse

Yuming Yang

3 papers indexed

arxivcs.AI2026-07-09

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

Yuming Yang, Xiao Sun, Yuanwei Zou, Zhengxiao Wu, Yun Chen, Jiang Zhong, et al.

Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual eva…

View free PDFSource page
arxivcs.AI2026-07-06

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, et al.

Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate agents in simplified, idealized settings. They t…

View free PDFSource page
arxivcs.AI2026-06-30

FARS: A Fully Automated Research System Deployed at Scale

Qiong Tang, Tianxiang Sun, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao, et al.

Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Autom…

View free PDFSource page