XiangqiBench provides closed-loop evaluation of LLM agents by requiring mate delivery in 119 tactical endgames against an engine defender
Read the original at arxiv.org→arXiv:2610.02425v1 Announce Type: new Abstract: Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We...
Original headline: "Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents"
Coverage timeline
- Oct 5, 04:00 UTC arXiv cs.CL lead source Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents