AgentHorizon evaluates agentic judges for long-horizon computer-use tasks
Read the original at arxiv.org→arXiv:2610.11050v1 Announce Type: new Abstract: Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without...
Original headline: "AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks"
Coverage timeline
- Oct 9, 04:00 UTC arXiv cs.AI lead source AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks