CORTEXA
← Browse

Chen Li

7 papers indexed

arxivcs.CVcs.AI2026-07-21

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

Jialong Zuo, Haotong Zuo, Shiwei Zhang, Xiang Wang, Chen Li, Nong Sang, et al.

Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at short, single-scene clips within narrow temp…

View free PDFSource page
arxivcs.CV2026-07-17

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen, Hao Liu, Mohan Zhang, et al.

Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-$p$ routing creates uneven per-head workloads under multi-GPU seq…

View free PDFSource page
arxivcs.CLcs.LG2026-07-12

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan, Liang Ding

On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In this paper, we systematically investigate this pathology and i…

View free PDFSource page
arxivcs.CVcs.AIcs.CLcs.LG2026-07-09

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

Feng Wang, Canmiao Fu, Zhipeng Huang, Chen Li, Jing Lyu, Ge Li

Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a shared context window, limiting long-horizon multimodal dialog…

View free PDFSource page
arxivcs.CV2026-06-26

EMOSH: Expressive Motion and Shape Disentanglement for Human Animation

Dongbin Zhang, Hao Liu, Binquan Dai, Kangjie Chen, Chuming Wang, Chen Li, et al.

High-fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods face a dilemma between expressiveness and disentanglement. Mainstream 2D pose-conditioned approaches suffer from "motion-shape en…

View free PDFSource page
arxivcs.CV2026-06-25

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation

Xiaomin Li, Qian Liang, Yinan Li, Ying Zhang, Chen Li, Jing Lyu, et al.

Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods often favor superficial aesthetics, such as over-saturated colors, leaving critical flaws like AI artifacts and biological implau…

View free PDFSource page
crossrefFoods2025-06-05Cited by 3

Efficient and Non-Invasive Grading of Chinese Mitten Crab Based on Fatness Estimated by Combing Machine Vision and Deep Learning

Jiangtao Li, Hongbao Ye, Chengquan Zhou, Xiaolian Yang, Zhuo Li, Qiquan Wei, et al.

The Chinese mitten crab (Eriocheir sinensis) is a high-value seafood. Efficient quality-grading methods are needed to meet rapid increases in demand. The current grading system for crabs primarily relies on manual observations and weights; it is thus inefficient, requires large a…

View free PDFSource page