CORTEXA
← Browse

Baochang Zhang

6 papers indexed

arxivcs.ROcs.LG2026-07-21

End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation

Jingzheng Li, Yufei Ge, Zhijun Chen, Qianren Mao, Zizhe Wang, Binhang Qi, et al.

Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offer…

View free PDFSource page
arxivcs.CVcs.AI2026-07-09

LDFE: Laplacian Decoupled Feature Enhancement Block for Dual-Stream CNN-based RGB-IR Object Detection

Wenhao Dong, Xiaoyan Luo, Linlin Yang, Haodong Zhu, Xiaorong Shi, Guodong Guo, et al.

The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Existing methods prefer dual-stream CNN backbones built upon YOLO for feature extraction and focus on the design of feature fusion. In this pap…

View free PDFSource page
arxivcs.CV2026-07-07

MAC-XA: Multi-view Anatomy-Correspondence Fusion for Coronary Stenosis Reporting from X-ray Angiography

Chen Jia, Baochang Zhang, Fatia Kusuma Dewi, Amir Yousefi, Heribert Schunkert, Reza Ghotbi, et al.

Multi-view reasoning in coronary X-ray angiography is inherently a cross-projection geometric problem, yet automated report generation in this setting remains largely unexplored. The 3D vascular topology leads to projection-dependent branch overlap and foreshortening, rendering s…

View free PDFSource page
arxivcs.CV2026-07-04

InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection

Zichao Feng, Haodong Zhu, Jingying Yang, Sheng Xu, Yangyang Ren, Yuguang Yang, et al.

Robust object detection under adverse visual conditions remains a long-standing challenge for multi-modal perception systems. Existing fusion-based methods typically require both RGB and infrared (IR) inputs, and treat them equally during both training and inference, which compro…

View free PDFSource page
arxivcs.CV2026-07-02

Teaching Vision-Language-Action Models What to See and Where to Look

Yuguang Yang, Canyu Chen, Zhewen Tan, Yizhi Wang, Zichao Feng, Chunyang Liu, et al.

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought reasoning data, which emphasizes linguistic reasoning rather…

View free PDFSource page
arxivcs.CV2026-06-26

Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading

Hong Li, Minqi Meng, Yanjun Liang, Chongjie Ye, Houyuan Chen, Weiqing Xiao, et al.

Reconstructing high-fidelity, relightable 3D avatars from a single in-the-wild image is a challenging ill-posed problem, primarily hindered by the scarcity of high-quality PBR data and the complexity of disentangling illumination from intrinsic materials. In this paper, we presen…

View free PDFSource page