CORTEXA
← Browse

Yunfan Shao

4 papers indexed

arxivcs.AI2026-06-30

FARS: A Fully Automated Research System Deployed at Scale

Qiong Tang, Tianxiang Sun, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao, et al.

Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Autom…

View free PDFSource page
arxivcs.CLcs.AI2026-06-26

Position Bias Correction is Insufficient for One-Pass Attention Sorting

Qiong Tang, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao

Long-context language models suffer from position bias, where information in middle positions is underutilized. Attention Sorting addresses this by iteratively reordering documents based on attention patterns, but its multiple sort-and-generate cycles increase deployment cost. We…

View free PDFSource page
arxivcs.CLcs.AI2026-06-26

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation

Qiong Tang, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao

Hybrid attention models that mix full and sliding-window attention across layers offer a promising approach to efficient long-context inference, but the critical question of \emph{which layers} should retain full attention remains unsolved. Existing methods use either fixed perio…

View free PDFSource page
arxivcs.CLcs.AI2026-06-26

Output-Space Allocation Costs for Calibration-Guided LLM Compression: An Empirical Study

Qiong Tang, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao

Training-free compression methods for large language models (LLMs) often use calibration data to guide compression decisions. ROCKET, a recent method combining sparse-dictionary factorization with multi-choice knapsack problem (MCKP) allocation, derives its per-layer factorizatio…

View free PDFSource page