arxivcs.CLcs.SE2026-07-30
Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
Burak Payzun, İrem Demirtaş, Simona Scala, Elena Ferretti, Seçil Arslan
Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monitoring, and human escalation. Yet evaluation often remains model-centric: benchmark scores, task accuracy, or one-off qualitative…