Learning from the gap between Pass@K and Pass@1
Read the original at arxiv.org→arXiv:2609.35793v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR). An exact verifier can also support test-time scaling...
Original headline: "Learning from the Gap Between Pass@K and Pass@1"
Coverage timeline
- Sep 30, 04:00 UTC arXiv cs.LG lead source Learning from the Gap Between Pass@K and Pass@1