CORTEXA
← Browse

Zhiyong Wang

4 papers indexed

arxivcs.AI2026-07-20

FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models

Bing Tian, Haikun Liu, Xiaocheng Zhong, Zhuohui Duan, Zhaokai Luo, Huayi Jin, et al.

Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly serial. Prior work has attempted to unlock inter-block parallelism through post-training methods, bu…

View free PDFSource page
arxivcs.IRcs.AIcs.LG2026-07-19

WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture

Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang, Velvin Fu, Maggie Zhuang, et al.

As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features…

View free PDFSource page
arxivcs.CLcs.HCcs.IR2026-07-13

PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing

Keshen Zhou, Lintao Wang, Suqin Yuan, Zhuqiang Lu, Yu Luo, Zhiyong Wang

Researchers organize the papers they collect into personal folder hierarchies in reference managers, and route each new paper into the folder where it belongs. This task differs from standard hierarchical text classification. A user's folder hierarchy is not a fixed, shared taxon…

View free PDFSource page
arxivcs.AI2026-07-07

Akashic: A Low-Overhead LLM Inference Service with MemAttention

Yang Liu, Zhaokai Luo, Huayi Jin, Ruozhou He, Chenchen Hong, Zhiyong Wang, et al.

Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts increase prefill cost, may exceed context limits,…

View free PDFSource page