arxivcs.DBcs.AIcs.CLcs.LG2026-07-24
DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng
LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a running database); observation-space…