Convergence is likely learned during pretraining, with semantic convergence observed from the first alignment stage (instruction-tuning); output homogeneity traces back to base models
Read the original at arxiv.org→arXiv:2608.11426v1 Announce Type: new Abstract: The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue...
Original headline: "Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models"
Coverage timeline
- Aug 13, 04:00 UTC arXiv cs.CL lead source Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models