Multilingual verifier bias in RLVR: benchmark, rollout diagnosis, and the cross-lingual selection bottleneck
Read the original at arxiv.org→arXiv:2608.20362v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier...
Original headline: "Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck"
Coverage timeline
- Aug 24, 04:00 UTC arXiv cs.CL lead source Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck