arxivcs.CLcs.AI2026-07-07
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exp…