CORTEXA
← Browse

Yibo Yang

5 papers indexed

arxivcs.AIcs.CLcs.LG2026-07-14

Self-Improvements in Modern Agentic Systems: A Survey

Zhe Ren, Yimeng Chen, Dandan Guo, Guowei Rong, Tonghui Li, R. B. Xiong, et al.

Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as adaptive systems that conve…

View free PDFSource page
arxivcs.LGcs.CL2026-07-13

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey, Bang An, Bernard Ghanem, Yibo Yang

Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification,…

View free PDFSource page
arxivcs.LG2026-07-05

SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity

Liyang Yuan, Yibo Yang, Dandan Guo, Peter Richtarik, Zhouchen Lin

Federated Learning (FL) is fundamentally challenged by statistical heterogeneity, where non-identically distributed (non-IID) data induces client drift that severely hampers global convergence. While existing approaches attempt to mitigate this drift through spatial-domain gradie…

View free PDFSource page
arxivcs.CRcs.AI2026-06-29

Defending Against Harmful Supervision Hidden in Benign Samples

Bang An, Yibo Yang, Dandan Guo, Ebtisam Alshehri, Carlos Hinojosa, Bernard Ghanem

Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign tasks. We propose Embedded Attack, where harmful QA pairs are embedded within benign training samples,…

View free PDFSource page