Reasoning models fail to ration test-time compute across questions
Read the original at arxiv.org→arXiv:2608.07968v1 Announce Type: new Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet...
Original headline: "Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions"
Coverage timeline
- Aug 11, 04:00 UTC arXiv cs.CL lead source Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions