arxivcs.CLcs.AI2026-07-08
Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System
Yiming Gai, Junde Lu, Xuefei Huang
The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabili…