OrchestraBench evaluates failure, recovery, and decomposition in multi-agent orchestration with a seed-reproducible failure-injection framework over templated enterprise workflows
Read the original at arxiv.org→arXiv:2608.05263v1 Announce Type: new Abstract: Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where...
Original headline: "OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality"
Coverage timeline
- Aug 7, 04:00 UTC arXiv cs.AI lead source OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality