arxivcs.CLcs.CV2026-07-02
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines. Existing art-focused evaluations rely on synthetic questions and rarely report item-level properties. This paper intro…