Relay-Bench evaluates LLMs on multi-domain reasoning chains; GPT-5.5 (xHigh) scores 43.3% on composite problems across domains
Read the original at arxiv.org→arXiv:2607.18438v1 Announce Type: new Abstract: Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single...
Original headline: "Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains"