arxivcs.AIcs.CLcs.LG2026-07-09
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis be…