CORTEXA
← Browse
openalexCityU Scholars2026-08-01Cited by 0

Exploring Recommender System Evaluation:A Multi-Modal LLM Agent Framework for A/B Testing

Wenlin Zhang, X J Li, Qiyuan Ge, Kuicai Dong, Pengyue; id_orcid 0000-0003-4712-3676 Jia, X J Li, Zijian Zhang, Maolin; id_orcid 0000-0002-0073-0172 Wang, Yue Wang, Huifeng Guo, Ruiming Tang, Xiangyu; id_orcid 0000-0003-2926-4416 Zhao

diningIn recommender systems, online A/B testing is a crucial method for evaluating the performance of different models. However, conducting online A/B testing often presents significant challenges, including substantial economic costs, user experience degradation, and considerable time requirements. With the Large Language Models' powerful capacity, LLM-based agent shows great potential to replace traditional online A/B testing. Nonetheless, current agents fail to simulate the perception process and interaction patterns, due to the lack of real environments and visual perception capability. To address these challenges, we introduce a multi-modal user agent for A/B testing (A/B Agent). Specifically, we construct a recommendation sandbox environment for A/B testing, enabling multimodal and multi-page interactions that align with real user behavior on online platforms. The designed agent leverages multimodal information perception, fine-grained user preferences, and integrates profiles, action memory retrieval, and a fatigue system to simulate complex human decision-making. We validated the potential of the agent as an alternative to traditional A/B testing from three perspectives: model, data, and features. Furthermore, we found that the data generated by A/B Agent can effectively enhance the capabilities of recommendation models. Our code is publicly available at https://github.com/Applied-Machine-Learning-Lab/ABAgent. © 2026 Owner/Author.

View free PDFSource page

Related papers

openalexCityU Scholars2026-08-01

Medusa:Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation

Yingjia Shang, Yi; id_orcid 0000-0002-0811-6150 Liu, H Wang, Feng Li, Wenfang Sun, Chengyu Wu, et al.

With the rapid advancement of retrieval-augmented vision-language models, multimodal medical retrieval-augmented generation (MMed-RAG) systems are increasingly adopted in clinical decision support. These systems enhance medical applications by performing cross-modal retrieval to…

View free PDFSource page