Sample count is not enough: candidate-generation strategy shapes the energy and performance of LLM test-time scaling
Read the original at arxiv.org→arXiv:2609.19499v1 Announce Type: new Abstract: Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget...
Original headline: "Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling"
Coverage timeline
- Sep 18, 04:00 UTC arXiv cs.LG lead source Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling