Data artifact isn’t a shortcut; causal auditing of synthetic RLVR corpora reveals potential leakage between correctness and provenance in GooseReason-0.7M
Read the original at arxiv.org→arXiv:2610.00202v1 Announce Type: new Abstract: Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The...
Original headline: "When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora"
Coverage timeline
- Oct 2, 04:00 UTC arXiv cs.CL lead source When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora