Stabilized best-of-K training for neural combinatorial optimization improves Leader Reward with a stabilized rank signal for POMO on TSP-100; achieves 7.7662 under 100-start, 8-augmentation greedy decoding
Read the original at arxiv.org→arXiv:2608.00296v1 Announce Type: new Abstract: Leader Reward modifies POMO training to emphasize the best trajectory produced by repeated inference. We test a narrow extension: replace its binary leader/non-leader...
Original headline: "Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization"
Coverage timeline
- Aug 4, 04:00 UTC arXiv cs.LG lead source Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization