What makes a terminal-bench task hard? Separating genuine hardness from fake-hardness on an adjudicated agentic corpus
Read the original at arxiv.org→arXiv:2609.26826v1 Announce Type: new Abstract: Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come...
Original headline: "What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus"
Coverage timeline
- Sep 24, 04:00 UTC arXiv cs.LG lead source What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus