Information-Gain rewards over diversity-pruned tests: GT-anchored verifier co-training for reliable code generation
Read the original at arxiv.org→arXiv:2609.21208v1 Announce Type: new Abstract: Self-play methods that co-train a single language model as both coder and test author promise to move code-generation RL beyond fixed test suites, but they suffer from...
Original headline: "Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation"
Coverage timeline
- Sep 21, 04:00 UTC arXiv cs.AI lead source Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation