arxivcs.CLcs.AIcs.LG2026-07-20
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The test set entirely consists of composite problems:…