CORTEXA
← Browse

Peng Zhang

17 papers indexed

arxivcs.CVcs.AIcs.LGcs.MMeess.IV2026-07-21

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, et al.

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed compo…

View free PDFSource page
arxivcs.CVcs.GRcs.LG2026-07-20

SciForma: Structure-Faithful Generation of Scientific Diagrams

Yuxuan Luo, Peng Zhang, Xinjie Zhang, Xun Guo, Zhouhui Lian, Yan Lu

Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can inva…

View free PDFSource page
arxivcs.CV2026-07-15

MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

Zhongyi Zhang, Guangyuan Wang, Li Hu, Wenbo Zhou, Peng Zhang, Tianyi Wei, et al.

Recent advances in generative models and technological innovations have significantly addressed the fundamental challenges of character image animation. However, existing approaches predominantly focus on character animation from a single reference image, substantially limiting t…

View free PDFSource page
arxivcs.CV2026-07-14

WanToFight: Real-Time Generative Game Engine for Multi-Player Combat Interaction

Li Hu, Guangyuan Wang, Peng Zhang, Bang Zhang

We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Prior generative game engines target either single-player first-person settings or non-real-time cooperative scenarios; multi-play…

View free PDFSource page
crossrefSustainability2026-07-14

Can Digital Infrastructure Predict Regional Innovation Capacity? Evidence Based on Machine Learning and the SHAP Explanatory Framework

Shasha Xie, Shusheng Xu, Wei Xu, Guo Yu, Yunli Li, Jianqiu Wu, et al.

Digital infrastructure is a critical foundation for promoting regional innovation and achieving sustainable development. However, existing studies have primarily focused on its impact effects, while paying limited attention to whether digital infrastructure can effectively identi…

View free PDFSource page
arxivcs.CVcs.SD2026-07-10

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Ruoshi Zhang, Yi Lu, et al.

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether t…

View free PDFSource page
arxiveess.SP2026-07-06

Unexpected Far-Near-Far Transition in Mobile Near Field Terahertz Communications

Peng Zhang, Hanmei Yuan, Zhe Wang, Vitaly Petrov, Emil Björnson

At THz frequencies, the radiative near-field distance can be sufficiently large to matter in real deployments. Existing near-field formulas are often understood in a simple way: as the link distance decreases, the propagation regime is expected to change only once, i.e., from far…

View free PDFSource page
arxivcs.CV2026-07-06

Fully Rotation-Equivariant Spectral-Spatial Learning for Multispectral Object Detection

Peng Zhang, Tingfa Xu, Shuaihao Han, Jianan Li

Existing multispectral detectors are limited by discrete spectral processing, a scale-dependent shift in the relative reliability of spectral and spatial cues across pyramid levels, and the lack of explicit rotation-equivariant geometric priors for arbitrarily oriented objects. T…

View free PDFSource page
arxivcs.CVcs.AIcs.GRcs.LG2026-07-05

Wan-Streamer v0.2: Higher Resolution, Same Latency

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Junjie He, et al.

We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises the interactive output stream from 192x336 to 640x368 while preserving approximately 200 ms model-si…

View free PDFSource page
arxivcs.CV2026-07-04

Reward Lightning: Fast Video Generation via Homologous Preference Distillation

Jiaxiang Cheng, Bing Ma, Xuhua Ren, Kai Yu, Peng Zhang, Tianxiang Zheng, et al.

Achieving simultaneous preference alignment and distillation acceleration in video diffusion models remains an open challenge. Existing methods optimize the two objectives over mismatched representation spaces, where improving one objective often compromises the other. To overcom…

View free PDFSource page
arxivcs.DBcs.AIcs.CLcs.LG2026-07-02

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Zhaoyan Sun, Shan Zhong, Daizhou Wen, Jiaxing Han, Guoliang Li, Ying Yan, et al.

Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process is essential to reducing labor-intensive efforts for data scientists and enabling scalable data-dri…

View free PDFSource page
arxivcs.LG2026-06-30

Visualizing High-Dimensional Graph Embeddings via Informed Multi-View Projections

Ya Ji, Xuefeng Li, Timo Brand, Jacob Miller, Peng Zhang, Stephen Kobourov, et al.

Graphs are commonly visualized in 2D, where humans readily interpret spatial relationships, yet such layouts often distort higher-dimensional structure. We propose to embed graphs in high-dimensional space and search for informative 2D viewpoints that optimize aesthetic and reada…

View free PDFSource page
crossrefDiagnostics2025-07-08Cited by 4

Machine Learning and Deep Learning Hybrid Approach Based on Muscle Imaging Features for Diagnosis of Esophageal Cancer

Yuan Hong, Hanlin Wang, Qi Zhang, Peng Zhang, Kang Cheng, Guodong Cao, et al.

Background: The rapid advancement of radiomics and artificial intelligence (AI) technology has provided novel tools for the diagnosis of esophageal cancer. This study innovatively combines muscle imaging features with conventional esophageal imaging features to construct deep lea…

View free PDFSource page
crossrefApplied Sciences2024-10-02Cited by 2

A Deep Learning Inversion Method for Airborne Time-Domain Electromagnetic Data Using Convolutional Neural Network

Xiaodong Yu, Peng Zhang, Xi Yu

Due to the high detection efficiency of the airborne time-domain electromagnetic method, it can quickly collect electromagnetic response data for large area-wide regions, but it also brings great challenges to the inversion interpretation of the data because there are numerous su…

View free PDFSource page
crossrefBuildings2024-08-08Cited by 4

Research on Prediction of Excavation Parameters for Deep Buried Tunnel Boring Machine Based on Convolutional Neural Network-Long Short-Term Memory Model

Yunfu Jia, Chengyuan Pei, Mingjian Dai, Xuan Che, Peng Zhang

Hard rock tunnel boring machines (TBMs) are increasingly widely used in tunnel construction today; however, TBMs are deeply buried underground and have a low perception of the underground surrounding rock conditions and excavation parameters. In order to ensure the safety of TBM…

View free PDFSource page
crossrefRemote Sensing2023-10-10Cited by 9

PM2.5 Estimation in Day/Night-Time from Himawari-8 Infrared Bands via a Deep Learning Neural Network

Junwei Wang, Kun Gao, Xiuqing Hu, Xiaodian Zhang, Hong Wang, Zibo Hu, et al.

Satellite-based PM2.5 estimation is an effective means to achieve large-scale and long-term PM2.5 monitoring and investigation. Currently, most of methods retrieve PM2.5 from satellite-derived aerosol optical depth (AOD) or top-of-atmosphere reflectance (TOAR) during daytime. A f…

View free PDFSource page