CORTEXA
← Browse

Zexin Zhuang

3 papers indexed

arxivcs.SEcs.LG2026-07-01

The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

Zhichao Fan, Yanhang Li, Zexin Zhuang, Xian Sun, Yingshuo Wang

Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpoints of one release line as if the model behind it had not shifted. We test whether it has shifted by auditing four open-source re…

View free PDFSource page
arxivcs.CRcs.LG2026-06-30

Probe Choice Changes Canary-Memorization Verdicts: Three Post-Hoc Disagreement Case Studies in a Text-Dominant LoRA-Tuned Autoregressive Testbed

Zhichao Fan, Zexin Zhuang, Yanhang Li

We audit a fixed prefix-window mean-NLL memorization probe (K=20) on a Qwen2.5-VL-7B canary testbed and report three post-hoc cases where it disagrees with full-span secret NLL or greedy exact-recall. C3 (false negative, window truncation): damage lands on hex tokens outside K=20…

View free PDFSource page